Regular expressions as siteswap generator filters

The Juggling Lab siteswap generator allows regular expressions to be used as output filters. This page is not intended to be a comprehensive discussion of regular expressions; for this read one of the several good tutorials available on the web.

Juggling Lab uses standard regular expression syntax, with some important differences.

Difference 1: Swapped metacharacters and literals

In standard regular expressions, the characters []()| act as metacharacters with special non-literal meaning. Doing a literal match of one of these characters requires a preceding backslash '\', for example the regex \[ matches the string [. In siteswap notation the characters []()| have special meaning, so relative to standard regular expressions we swap the roles of [ and \[. So within Juggling Lab the regex [ is a literal match for [, and \[ and \] are used to define character classes (see below).

Difference 2: Unanchored matching by default

Regular expression filters in Juggling Lab match anywhere within a pattern by default (unanchored). Thus an include filter of 4 will match a 4 throw anywhere in the pattern. The boundary matchers ^ and $ match to the beginning or end of the pattern respectively. For example, the filter ^4 matches patterns starting with a 4 throw, and 4$ matches patterns ending with a 4 throw.

Juggling Lab regular expression summary

Characters

   Char          Matches any identical character

Character Classes

   \[abc\]       Simple character class
   \[a-zA-Z\]    Character class with ranges
   \[^abc\]      Negated character class

Predefined Classes

   .             Matches any character other than newline
   \d            Matches a digit character
   \D            Matches a non-digit character

Boundary Matchers

   ^             Matches only at the beginning of a pattern
   $             Matches at the end of a pattern, or throw (see note above)

Greedy Closures

   A*            Matches A 0 or more times (greedy)
   A+            Matches A 1 or more times (greedy)
   A?            Matches A 1 or 0 times (greedy)
   A{n}          Matches A exactly n times (greedy)
   A{n,}         Matches A at least n times (greedy)
   A{n,m}        Matches A at least n but not more than m times (greedy)

Reluctant Closures

   A*?           Matches A 0 or more times (reluctant)
   A+?           Matches A 1 or more times (reluctant)
   A??           Matches A 0 or 1 times (reluctant)

Logical Operators

   AB            Matches A followed by B
   A\|B          Matches either A or B
   \(A\)         Used for subexpression grouping
   \(?:A\)       Used for subexpression clustering (just like grouping but no backrefs)

Backreferences

   \1            Backreference to 1st parenthesized subexpression
   \2            Backreference to 2nd parenthesized subexpression
   \3            Backreference to 3rd parenthesized subexpression
   \4            Backreference to 4th parenthesized subexpression
   \5            Backreference to 5th parenthesized subexpression
   \6            Backreference to 6th parenthesized subexpression
   \7            Backreference to 7th parenthesized subexpression
   \8            Backreference to 8th parenthesized subexpression
   \9            Backreference to 9th parenthesized subexpression

You can refer to the contents of a parenthesized expression within a regular expression itself. This is called a 'backreference'. The first backreference in a regular expression is denoted by \1, the second by \2 and so on. So the expression:

\(\[0-9\]+\)=\1

will match any string of the form n=n (like 0=0 or 2=2).

All closure operators (+, *, ?, {m,n}) are greedy by default, meaning that they match as many elements of the string as possible without causing the overall match to fail. If you want a closure to be reluctant (non-greedy), you can simply follow it with a '?'. A reluctant closure will match as few elements of the string as possible when finding matches. {m,n} closures don't currently support reluctancy.

Examples

The following examples demonstrate how to use include (-i) and exclude (-x) regular expression filters with jlab gen from the command line.

Include throws anywhere

   jlab gen 5 7 4 -i 2

Includes only patterns that contain a 2 throw anywhere in the pattern.

Exclude throws anywhere

   jlab gen 5 7 4 -x 2

Excludes all patterns that contain a 2 throw anywhere in the pattern.

Anchor to the beginning of the pattern

   jlab gen 5 7 4 -i '^7'

Includes only patterns that begin with a 7 throw.

   jlab gen 5 7 4 -s -g -i '^(4x'

Includes only synchronous patterns that begin with a (4x throw. Note that unescaped parentheses () match literal siteswap parentheses.

Anchor to the end of the pattern

   jlab gen 5 7 4 -i '4$'

Includes only patterns that end with a 4 throw.

   jlab gen 5 7 4 -s -g -i '2)$'

Includes only synchronous patterns whose final throw ends with 2).

Character classes and literal brackets

   jlab gen 5 7 4 -s -g -i '2\[^x\]'

Matches patterns containing a 2 throw not followed by x (i.e., uncrossed). In Juggling Lab regex syntax, \[ and \] define character classes.

   jlab gen 5 7 4 -m 2 -f -x '['

Excludes all patterns containing multiplex throws. Because unescaped [ matches literal brackets, this matches any multiplex notation.

Combining multiple filters

   jlab gen 5 7 4 -s -g -i 2 -x 2x

Includes patterns that contain 2 throws, but excludes patterns containing crossed 2x throws.

   jlab gen 5 7 4 -s -g -i 2x -x '2\(?!x\)'

Includes patterns that contain 2x throws, but excludes any pattern that also contains an uncrossed 2 using a negative lookahead assertion (where \( and \) represent regex grouping parentheses).