New Parser in Python
From:
"andrew cooke" <andrew@...>
Date:
Mon, 12 Jan 2009 00:54:55 -0300 (CLST)
Not ready for release yet, but I've just got a new parser, written in
Python, to the point where it's useful.
It includes full backtracing (parse forests etc), the (untested) ability
to 'automatically' control resource use (think 'maximum backtrace stack')
and enough syntactic sugar to rot your teeth :o)
This test:
from logging import basicConfig, DEBUG
from unittest import TestCase
from lepl.match import *
from lepl.node import Node
class NodeTest(TestCase):
def test_node(self):
basicConfig(level=DEBUG)
class Term(Node): pass
class Factor(Node): pass
class Expression(Node): pass
expression = Delayed()
number = Digit()[1:,...] > 'number'
term = (number | '(' / expression / ')') > Term
muldiv = Any('*/') > 'operator'
factor = (term / (muldiv / term)[0:]) > Factor
addsub = Any('+-') > 'operator'
expression += (factor / (addsub / factor)[0:]) > Expression
(ast, _) = next(expression.match_string('1 + 2 * (3 + 4 - 5)'))
print(ast[0])
(ast, _) = next(expression('1 + 2 * (3 + 4 - 5)'))
print(ast[0])
Prints the following (twice):
Expression
+- Factor
| +- Term
| | `- number=1
| `- ' '
+- operator=+
+- ' '
`- Factor
+- Term
| `- number=2
+- ' '
+- operator=*
+- ' '
`- Term
+- '('
+- Expression
| +- Factor
| | +- Term
| | | `- number=3
| | `- ' '
| +- operator=+
| +- ' '
| +- Factor
| | +- Term
| | | `- number=4
| | `- ' '
| +- operator=-
| +- ' '
| `- Factor
| `- Term
| `- number=5
`- ')'
Andrew
Syntax
From:
"andrew cooke" <andrew@...>
Date:
Mon, 12 Jan 2009 01:11:39 -0300 (CLST)
A quick explanation of the syntax:
This allows forward references to 'expression', which will be defined later.
expression = Delayed()
This defines 'number' as one or more digits, specified via '[1:]', and
combines the digits into a single string, specified via '[...]'. The
result is then associated with the tag 'number'.
number = Digit()[1:,...] > 'number'
This defines term as either 'number' or (with backtracing) a bracketed
expression. The strings are automatically promoted to literal matches and
the '/' indicate that there are optional spaces between the matchers ('//'
for required space). The result is used to construct a Term instance,
which is a subclass of Node (and which will automatically construct
attributes for the contents).
term = (number | '(' / expression / ')') > Term
Define 'muldiv' to be either '*' or '/' and tag the result.
muldiv = Any('*/') > 'operator'
Hopefully this is becoming obvious. The '[0:]' here means '0 or more'
instances of 'muldiv', optional space, and 'term'.
factor = (term / (muldiv / term)[0:]) > Factor
Nothing new here.
addsub = Any('+-') > 'operator'
This defines the 'Delayed' matcher introduced earlier (it was introduced
so that we could reference it in 'term', even though we cannot define it
until later).
expression += (factor / (addsub / factor)[0:]) > Expression
Not sure if it's obvious, but one major aim has been to try to combine the
best of both OO and functional programming, in what I feel is a very
'Pythonic' way.
Andrew
With Bactracking
From:
"andrew cooke" <andrew@...>
Date:
Mon, 12 Jan 2009 01:25:59 -0300 (CLST)
Changing the spec slightly to:
expression = Delayed()
number = Digit()[1:,...] > 'number'
term = (number | '(' / expression / ')') > Term
muldiv = Any('*/') > 'operator'
factor = (term / (muldiv / term)[0:]) > Factor
addsub = Any('+-') > 'operator'
expression += Drop(Any()[0:]) & \
(factor / (addsub / factor)[0:]) > Expression
And using:
for (ast, _) in expression('1 + 2 * (3 + 4 - 5)'):
print(ast[0])
Gives:
Expression
`- Factor
`- Term
`- number '5'
Expression
+- Factor
| +- Term
| | `- number '4'
| `- ' '
+- operator '-'
+- ' '
`- Factor
`- Term
`- number '5'
Expression
`- Factor
+- Term
| `- number '4'
`- ' '
Expression
+- Factor
| `- Term
| `- number '4'
+- ' '
+- operator '-'
+- ' '
`- Factor
`- Term
`- number '5'
Expression
+- Factor
| `- Term
| `- number '4'
`- ' '
Expression
`- Factor
`- Term
`- number '4'
Expression
+- Factor
| +- Term
| | `- number '3'
| `- ' '
+- operator '+'
+- ' '
+- Factor
| +- Term
| | `- number '4'
| `- ' '
+- operator '-'
+- ' '
`- Factor
`- Term
`- number '5'
Expression
+- Factor
| +- Term
| | `- number '3'
| `- ' '
+- operator '+'
+- ' '
`- Factor
+- Term
| `- number '4'
`- ' '
Expression
+- Factor
| +- Term
| | `- number '3'
| `- ' '
+- operator '+'
+- ' '
`- Factor
`- Term
`- number '4'
Expression
`- Factor
+- Term
| `- number '3'
`- ' '
Expression
+- Factor
| `- Term
| `- number '3'
+- ' '
+- operator '+'
+- ' '
+- Factor
| +- Term
| | `- number '4'
| `- ' '
+- operator '-'
+- ' '
`- Factor
`- Term
`- number '5'
Expression
+- Factor
| `- Term
| `- number '3'
+- ' '
+- operator '+'
+- ' '
`- Factor
+- Term
| `- number '4'
`- ' '
Expression
+- Factor
| `- Term
| `- number '3'
+- ' '
+- operator '+'
+- ' '
`- Factor
`- Term
`- number '4'
Expression
+- Factor
| `- Term
| `- number '3'
`- ' '
Expression
`- Factor
`- Term
`- number '3'
Expression
`- Factor
`- Term
+- '('
+- Expression
| `- Factor
| `- Term
| `- number '5'
`- ')'
Expression
`- Factor
`- Term
+- '('
+- Expression
| +- Factor
| | +- Term
| | | `- number '4'
| | `- ' '
| +- operator '-'
| +- ' '
| `- Factor
| `- Term
| `- number '5'
`- ')'
Expression
`- Factor
`- Term
+- '('
+- Expression
| +- Factor
| | `- Term
| | `- number '4'
| +- ' '
| +- operator '-'
| +- ' '
| `- Factor
| `- Term
| `- number '5'
`- ')'
Expression
`- Factor
`- Term
+- '('
+- Expression
| +- Factor
| | +- Term
| | | `- number '3'
| | `- ' '
| +- operator '+'
| +- ' '
| +- Factor
| | +- Term
| | | `- number '4'
| | `- ' '
| +- operator '-'
| +- ' '
| `- Factor
| `- Term
| `- number '5'
`- ')'
Expression
`- Factor
`- Term
+- '('
+- Expression
| +- Factor
| | `- Term
| | `- number '3'
| +- ' '
| +- operator '+'
| +- ' '
| +- Factor
| | +- Term
| | | `- number '4'
| | `- ' '
| +- operator '-'
| +- ' '
| `- Factor
| `- Term
| `- number '5'
`- ')'
Expression
`- Factor
+- Term
| `- number '2'
+- ' '
+- operator '*'
+- ' '
`- Term
+- '('
+- Expression
| `- Factor
| `- Term
| `- number '5'
`- ')'
Expression
`- Factor
+- Term
| `- number '2'
+- ' '
+- operator '*'
+- ' '
`- Term
+- '('
+- Expression
| +- Factor
| | +- Term
| | | `- number '4'
| | `- ' '
| +- operator '-'
| +- ' '
| `- Factor
| `- Term
| `- number '5'
`- ')'
Expression
`- Factor
+- Term
| `- number '2'
+- ' '
+- operator '*'
+- ' '
`- Term
+- '('
+- Expression
| +- Factor
| | `- Term
| | `- number '4'
| +- ' '
| +- operator '-'
| +- ' '
| `- Factor
| `- Term
| `- number '5'
`- ')'
Expression
`- Factor
+- Term
| `- number '2'
+- ' '
+- operator '*'
+- ' '
`- Term
+- '('
+- Expression
| +- Factor
| | +- Term
| | | `- number '3'
| | `- ' '
| +- operator '+'
| +- ' '
| +- Factor
| | +- Term
| | | `- number '4'
| | `- ' '
| +- operator '-'
| +- ' '
| `- Factor
| `- Term
| `- number '5'
`- ')'
Expression
`- Factor
+- Term
| `- number '2'
+- ' '
+- operator '*'
+- ' '
`- Term
+- '('
+- Expression
| +- Factor
| | `- Term
| | `- number '3'
| +- ' '
| +- operator '+'
| +- ' '
| +- Factor
| | +- Term
| | | `- number '4'
| | `- ' '
| +- operator '-'
| +- ' '
| `- Factor
| `- Term
| `- number '5'
`- ')'
Expression
`- Factor
+- Term
| `- number '2'
`- ' '
Expression
+- Factor
| `- Term
| `- number '2'
`- ' '
Expression
`- Factor
`- Term
`- number '2'
Expression
+- Factor
| +- Term
| | `- number '1'
| `- ' '
+- operator '+'
+- ' '
`- Factor
+- Term
| `- number '2'
+- ' '
+- operator '*'
+- ' '
`- Term
+- '('
+- Expression
| `- Factor
| `- Term
| `- number '5'
`- ')'
Expression
+- Factor
| +- Term
| | `- number '1'
| `- ' '
+- operator '+'
+- ' '
`- Factor
+- Term
| `- number '2'
+- ' '
+- operator '*'
+- ' '
`- Term
+- '('
+- Expression
| +- Factor
| | +- Term
| | | `- number '4'
| | `- ' '
| +- operator '-'
| +- ' '
| `- Factor
| `- Term
| `- number '5'
`- ')'
Expression
+- Factor
| +- Term
| | `- number '1'
| `- ' '
+- operator '+'
+- ' '
`- Factor
+- Term
| `- number '2'
+- ' '
+- operator '*'
+- ' '
`- Term
+- '('
+- Expression
| +- Factor
| | `- Term
| | `- number '4'
| +- ' '
| +- operator '-'
| +- ' '
| `- Factor
| `- Term
| `- number '5'
`- ')'
Expression
+- Factor
| +- Term
| | `- number '1'
| `- ' '
+- operator '+'
+- ' '
`- Factor
+- Term
| `- number '2'
+- ' '
+- operator '*'
+- ' '
`- Term
+- '('
+- Expression
| +- Factor
| | +- Term
| | | `- number '3'
| | `- ' '
| +- operator '+'
| +- ' '
| +- Factor
| | +- Term
| | | `- number '4'
| | `- ' '
| +- operator '-'
| +- ' '
| `- Factor
| `- Term
| `- number '5'
`- ')'
Expression
+- Factor
| +- Term
| | `- number '1'
| `- ' '
+- operator '+'
+- ' '
`- Factor
+- Term
| `- number '2'
+- ' '
+- operator '*'
+- ' '
`- Term
+- '('
+- Expression
| +- Factor
| | `- Term
| | `- number '3'
| +- ' '
| +- operator '+'
| +- ' '
| +- Factor
| | +- Term
| | | `- number '4'
| | `- ' '
| +- operator '-'
| +- ' '
| `- Factor
| `- Term
| `- number '5'
`- ')'
Expression
+- Factor
| +- Term
| | `- number '1'
| `- ' '
+- operator '+'
+- ' '
`- Factor
+- Term
| `- number '2'
`- ' '
Expression
+- Factor
| +- Term
| | `- number '1'
| `- ' '
+- operator '+'
+- ' '
`- Factor
`- Term
`- number '2'
Expression
`- Factor
+- Term
| `- number '1'
`- ' '
Expression
+- Factor
| `- Term
| `- number '1'
+- ' '
+- operator '+'
+- ' '
`- Factor
+- Term
| `- number '2'
+- ' '
+- operator '*'
+- ' '
`- Term
+- '('
+- Expression
| `- Factor
| `- Term
| `- number '5'
`- ')'
Expression
+- Factor
| `- Term
| `- number '1'
+- ' '
+- operator '+'
+- ' '
`- Factor
+- Term
| `- number '2'
+- ' '
+- operator '*'
+- ' '
`- Term
+- '('
+- Expression
| +- Factor
| | +- Term
| | | `- number '4'
| | `- ' '
| +- operator '-'
| +- ' '
| `- Factor
| `- Term
| `- number '5'
`- ')'
Expression
+- Factor
| `- Term
| `- number '1'
+- ' '
+- operator '+'
+- ' '
`- Factor
+- Term
| `- number '2'
+- ' '
+- operator '*'
+- ' '
`- Term
+- '('
+- Expression
| +- Factor
| | `- Term
| | `- number '4'
| +- ' '
| +- operator '-'
| +- ' '
| `- Factor
| `- Term
| `- number '5'
`- ')'
Expression
+- Factor
| `- Term
| `- number '1'
+- ' '
+- operator '+'
+- ' '
`- Factor
+- Term
| `- number '2'
+- ' '
+- operator '*'
+- ' '
`- Term
+- '('
+- Expression
| +- Factor
| | +- Term
| | | `- number '3'
| | `- ' '
| +- operator '+'
| +- ' '
| +- Factor
| | +- Term
| | | `- number '4'
| | `- ' '
| +- operator '-'
| +- ' '
| `- Factor
| `- Term
| `- number '5'
`- ')'
Expression
+- Factor
| `- Term
| `- number '1'
+- ' '
+- operator '+'
+- ' '
`- Factor
+- Term
| `- number '2'
+- ' '
+- operator '*'
+- ' '
`- Term
+- '('
+- Expression
| +- Factor
| | `- Term
| | `- number '3'
| +- ' '
| +- operator '+'
| +- ' '
| +- Factor
| | +- Term
| | | `- number '4'
| | `- ' '
| +- operator '-'
| +- ' '
| `- Factor
| `- Term
| `- number '5'
`- ')'
Expression
+- Factor
| `- Term
| `- number '1'
+- ' '
+- operator '+'
+- ' '
`- Factor
+- Term
| `- number '2'
`- ' '
Expression
+- Factor
| `- Term
| `- number '1'
+- ' '
+- operator '+'
+- ' '
`- Factor
`- Term
`- number '2'
Expression
+- Factor
| `- Term
| `- number '1'
`- ' '
Expression
`- Factor
`- Term
`- number '1'
Comment on this post