The advanced_regex generator
Use it when you need everything regex does, but the shape of the
string itself carries a statistical distribution of variants — say, exactly 70%
of the codes start with RU, 20% with US, 10% with DE.
Plain regex picks an alternation (RU|US|DE) randomly, so the
split only comes out right on average. advanced_regex adds weighted choice,
which lays the variants out in exact counts.
<gen type="advanced_regex" value="(?%{70:RU;20:US;10:DE})-[0-9]{6}"/>
advanced_regex is a superset of regex: the ordinary
type="regex" generator stays stable and TDC-neutral, while
advanced_regex layers TDC-specific power on top of the same finite, portable
engine. Today that extra power is exactly one construct — weighted choice — and
everything else on this page is inherited unchanged from
Regex.
Why you need it
Without weighted choice, "70% RU, 20% US, 10% DE, each followed by six random
digits" would take several sequences or a <mix>
to express. With advanced_regex it collapses into a single generator. It's a
natural fit for:
- codes for countries, branches, regions, or customer tiers
- test identifiers that must vary in structure but hold fixed proportions
- document numbers where one part of the string has to repeat another
- synthetic data where the shares matter as much as the values
- AI-generated
.tdcfiles, where an agent finds it easier to emit one compact pattern than a tree of sequences
Example outputs below are illustrative — exact rows depend on the seed and can shift slightly between core versions — but the counts a distribution promises are exact.
Weighted choice
The construct is:
(?%{PERCENT:BRANCH;PERCENT:BRANCH;...})
Piece by piece:
(?%{ 70:RU ; 20:US ; 10:DE })
│ │ │ │ │ │ │ │
│ │ branch │ branch │ branch│
│ percent percent percent │
└── start of weighted choice ┘
The rules:
- percentages must be numbers, and they must be non-negative
- the percentages must sum to 100
- a branch may be empty
- a branch is itself a full
advanced_regexexpression, so branches nest ;,}, and:are the control characters — escape them if a branch needs them literally
Exact proportions, by default
This is the headline: weighted choice promises exact percentages (exactly 70 of 100, not "about 70"), and TDC delivers them out of the box — you switch nothing on.
To measure the split exactly, the engine builds the whole column at once and
hands out branches with the Hamilton (largest-remainder) method. It does this
automatically the moment a (?%{…}) appears in the pattern — even under the default
disk mode. All you supply is the pattern; the shares come out exact.
<env count="100" seed="countries">
<sequence name="CountryCode">
<gen type="advanced_regex" value="(?%{70:RU;20:US;10:DE})-[0-9]{2}"/>
</sequence>
</env>
Count the prefixes across all 100 rows — exactly 70/20/10:
RU 70 US 20 DE 10
To measure the split exactly, a weighted-choice column is built entirely in RAM.
Fine for small and medium sets. If you need exact shares while streaming (O(1)
memory, any output size), <mix percent>
and <gen type="text" percent="…"> give
exact proportions in a stream, without holding the whole column. And if you
manually force the pure streaming engine (mode="stream" on <env>, an attribute-only
legacy alias, or --engine 2 on the command line — --mode stream is not a thing), TDC
won't silently ruin the percentages — it can't count them one row at a time, so it
refuses with a clear error. Drop the override and it's exact again.
The weights really move the distribution
The percentages are not decoration — they are the actual makeup of the column. Take
one pattern, (?%{…})-[0-9]{2} at count="1000", change only the weights, and
count the prefixes:
<gen type="advanced_regex" value="(?%{70:RU;20:US;10:DE})-[0-9]{2}"/> <!-- variant 1 -->
<gen type="advanced_regex" value="(?%{34:RU;33:US;33:DE})-[0-9]{2}"/> <!-- variant 2 -->
<gen type="advanced_regex" value="(?%{10:RU;10:US;80:DE})-[0-9]{2}"/> <!-- variant 3 -->
weights RU US DE 70 / 20 / 10 700 200 100 34 / 33 / 33 340 330 330 10 / 10 / 80 100 100 800
Over 1000 rows the counts reproduce the weights exactly. Change the weights and the makeup of the column changes with them.
regex vs advanced_regex
| Capability | regex | advanced_regex |
|---|---|---|
| Generate a string from a pattern | yes | yes |
| Character classes | yes | yes |
| Named Unicode alphabets | yes | yes |
| Groups and backreferences | yes | yes |
Length cap via regex_max_length | yes | yes |
| Exact percentages inside the pattern | no | yes |
| Nest weighted variants | no | yes |
| Weighted choice in the output block | no | no |
The difference is easiest to see side by side. Plain regex picks each
character freely — the shape is fixed, the shares are not:
<gen type="regex" value="[A-Z]{2}[0-9]{6}"/>
FZ399441 YH481897 LR586083 YA900972 WT831899
advanced_regex keeps the same kind of code but pins the prefix shares — 70% RU,
20% US, 10% DE over the full run:
<gen type="advanced_regex" value="(?%{70:RU;20:US;10:DE})-[0-9]{6}"/>
RU-441627 RU-476822 RU-948319 US-450875 RU-398584 RU-131212 RU-418648 RU-830959
In short: use regex for a string of a given form; use advanced_regex when the
form itself carries a statistical distribution.
It inherits the whole regex language
Every finite construct from regex works here too — literals, escapes,
character classes, BMP ranges, named alphabets \a{…}, \d/\w/\s and their
inverses, ., alternation, groups, backreferences, bounded quantifiers, and the
regex_max_length cap. A simple weighted split over
plain Latin codes:
<gen type="advanced_regex" value="(?%{70:[A-Z]{2};30:[A-Z]{3}})-[0-9]{4}"/>
QY-3500 ZT-3381 GSK-1914 VO-5921 DW-7570 SO-1660 MSE-2247
Unicode demo. Because the branches accept named alphabets, you can hold exact shares across scripts. Here 7 of 10 codes take a Cyrillic prefix and 3 take a Latin one — a deliberate Unicode/localization example, showing that the exact-percentage machinery is script-agnostic:
<gen type="advanced_regex" value="(?%{70:\a{cyrillic.ru.upper}{2};30:\a{latin.upper}{2}})-[0-9]{4}"/>
ЭЗ-2477 WJ-0170 ЧП-8026 СЦ-1020 ЫЦ-2747 FJ-7879 РЛ-6827 ЩЕ-4485 ПВ-0297 UD-1550
Nested weighted choice
Because branches are full expressions, weighted choices nest — and the inner split is computed inside the subset that reached the outer branch:
<gen type="advanced_regex" value="(?%{50:A(?%{80:X;20:Y});50:B})"/>
At count="100":
AX 40 AY 10 B 50
80% of the 50 A rows gives 40 AX; the other 20% gives 10 AY. This
"percentages within a subset" behavior matches TDC's
sequence-hierarchy philosophy exactly.
Several weighted choices in one pattern
<gen type="advanced_regex" value="(?%{60:M;40:F})-(?%{25:00;75:99})"/>
Each distribution is worked out exactly, over the same set of rows. At count="100":
M-99 44 F-99 31 M-00 16 F-00 9
Add up each part and both are exact: M = 44 + 16 = 60 and F = 31 + 9 = 40
(60/40); 99 = 44 + 31 = 75 and 00 = 16 + 9 = 25 (75/25). Two independent picks,
each laid out by the same exact-percentage method.
With a parent filter
advanced_regex lives inside the normal sequence-dependency model. If the sequence
is filtered by parent, the
percentages count only within the filtered subset:
<sequence name="Gender">
<gen type="text" value="M,F" percent="50,50"/>
</sequence>
<sequence name="MaleCode" parent="Gender.M">
<gen type="advanced_regex" value="M-(?%{40:A;60:B})-[0-9]{2}"/>
</sequence>
At count="100" (50 men):
F 50 (MaleCode empty) M-A 20 M-B 30
The 40/60 split is measured against the 50 filtered rows, not the full 100.
Captures and backreferences
A backreference repeats a group that was already generated. This works exactly as in
plain regex — the first three digits are echoed at the end:
<gen type="advanced_regex" value="([0-9]{3})-[A-Z]{2}-\1"/>
299-YZ-299 929-UE-929 462-VR-462 905-BC-905 876-JF-876
A weighted branch can be captured and repeated with \1 — the captured part is
echoed verbatim and the percentages still hold. At count="40":
<gen type="advanced_regex" value="((?%{25:AB;75:CD}))-\1"/>
AB-AB 10 CD-CD 30
The captured pair is repeated literally, and the 25/75 split survives.
A capture made before a weighted choice can be used inside a branch. Here half
the rows repeat the captured two letters, half print the fixed XX (count="8"):
<gen type="advanced_regex" value="([A-W]{2})-(?%{50:\1;50:XX})"/>
TV-XX GR-GR RN-XX OU-OU WM-WM SS-XX CL-XX QG-QG
A backreference can live inside a branch. Where the (A[0-9]) branch was taken,
\1 repeats its capture; where the B branch was taken, there is no capture, so
\1 comes out empty (count="20"):
<gen type="advanced_regex" value="(?%{40:(A[0-9]);60:B})-\1"/>
A2-A2 B- B- A1-A1 B- A8-A8 B-
It isn't a full if, but it's already a useful logical link: one part of a string
can depend on an already-generated group.
Escaping inside a weighted choice
;, } and : are the control characters of a weighted choice. To use them as
literal text in a branch, escape them (count="6"):
<gen type="advanced_regex" value="(?%{50:A\;\}\:;50:B})"/>
A;}: A;}: B B A;}: B
The branch A\;\}\: prints the literal A;}:, and the branch B prints just B.
Where weighted choice works
Weighted choice sets an exact percentage split, so the runtime has to know how many
rows fall into each branch. That count is known everywhere a <gen> can appear:
- inside a
<sequence>— fromcount, or from aparentsubset - inside a
<case>of a<mix>— from the size of that case
Both are valid:
<sequence name="CountryCode">
<gen type="advanced_regex" value="(?%{70:RU;20:US;10:DE})-[0-9]{2}"/>
</sequence>
<mix name="Country" percent="50,50">
<case><gen type="advanced_regex" value="(?%{70:RU;30:US})"/></case>
<case><data>-</data></case>
</mix>
There are no generators in the output block — a
<line> only formats — so weighted choice
never lands there. To print one, declare it as a sequence and interpolate it with
${{Name}}:
<tdc>
<env count="100" seed="demo" inject="${{%}}">
<sequence name="Code">
<gen type="advanced_regex" value="(?%{70:A;30:B})-[0-9]{4}"/>
</sequence>
</env>
<block>
<line><data>code=${{Code}}</data></line>
</block>
</tdc>
The first few rows, with the run splitting exactly 70 A / 30 B:
code=A-8870 code=B-2495 code=B-1961 code=A-8865 code=A-9221 code=A-3234
The row order is shuffled deterministically by seed; the totals stay exact.
Practical examples
Customer code by segment — count="1000", count the prefixes:
<gen type="advanced_regex" value="(?%{80:REG;15:VIP;5:TEST})-[A-Z]{2}[0-9]{4}"/>
REG 800 VIP 150 TEST 50
Document with a repeating block — the first three digits and the last three
digits always match, while the middle splits 60% A / 40% B (count="100"):
<gen type="advanced_regex" value="([0-9]{3})-(?%{60:A;40:B})-\1"/>
924-B-924 419-B-419 788-A-788 692-B-692
Short vs long technical codes — 85% short, 15% long (count="100"):
<gen type="advanced_regex" value="(?%{85:[A-Z]{2}[0-9]{2};15:[A-Z]{4}[0-9]{8}})"/>
length 4 85 (AB42) length 12 15 (ABCD12345678)
<mix> or advanced_regex?
Both do exact percentages, for different jobs.
Use <mix> when the branches have
different structure — each with its own generators and its own literal text (and
<mix percent> gives exact shares while streaming, without holding the column in
memory):
<mix name="Kind" percent="70,30">
<case><data>{"type":"regular"}</data></case>
<case><data>{"type":"vip","bonus":true}</data></case>
</mix>
Use advanced_regex when the whole thing fits in one pattern (exact by default; the
column is built in RAM — for very large outputs prefer <mix percent>):
<gen type="advanced_regex" value="(?%{70:REG;30:VIP})-[0-9]{6}"/>
Invalid patterns
<gen type="advanced_regex" value="(?%{70:A;20:B})"/> <!-- sums to 90, needs 100 -->
<gen type="advanced_regex" value="[a-z]+"/> <!-- unbounded, as in plain regex -->
Both are rejected before generation. Here's the error for the second one:
error: invalid advanced_regex generator pattern: unbounded "+"
quantifier is not allowed; use "{1,n}"Naming a group — (?<name>…)
A group can carry a name as well as a number:
(?<sex>(?%{50:male;50:female}))
That is an ordinary capturing group — \1 still reads it — with a label. The label
exists so that something LATER in the pattern can ask what this group produced,
which is what the next section is for.
\k<name> repeats the group by that name, exactly as in regex, and it
is the same group \1 repeats:
<gen type="advanced_regex" value="(?<c>(?%{50:AB;50:CD}))-\k<c>"/>
The name must belong to a group that has already closed — the rule \1 and
(?if{…}) both follow, since a group further along has produced nothing to repeat.
Names start with a letter or _ and hold letters, digits and _. Two groups
cannot share one name: (?if{sex=…}) would then be a coin toss between them. That
holds when one of the two sits inside the other, as well as side by side.
Making one part follow another — (?if{…})
Use it when two parts of the same value have to agree. Everything else in this
generator decides a value from randomness alone, which is why a pattern could
describe a postcode or an identifier but never a title that matches a sex chosen
two characters earlier. That used to mean abandoning advanced_regex and
rebuilding the column as a <switch>.
<gen type="advanced_regex" value="(?<sex>(?%{50:male;50:female}))/(?if{sex=male:Mr;sex=female:Ms})"/>
female/Ms male/Mr male/Mr female/Ms male/Mr female/Ms female/Ms male/Mr
Every row agrees with itself. male is never followed by Ms.
How it reads
(?if{name=value:branch;name=value:branch}). Branches are separated by ;, and
each is what to test then : then what to produce. They are tried in the
order written and the first match wins.
A branch is a full advanced_regex expression, so weighted choices and further
conditionals nest inside it:
<gen type="advanced_regex" value="(?<country>(?%{60:RU;30:US;10:DE}))-(?if{country=RU:[0-9]{3};country=US:[A-Z]{3};*:[A-Z]{2}[0-9]})"/>
RU-683 US-NGS DE-ZQ5 US-VNS RU-867 RU-722 RU-372 RU-890 US-SEA RU-589
* is the branch that matches everything else — the DE rows here. Without
one, a row that matches no branch produces nothing at all for that part:
(?<c>(?%{50:a;50:b}))-(?if{c=zzz:NEVER}) → a- b- a- b- …
That is deliberate. The pattern said nothing about what an a row should produce,
and quietly falling back to the first branch would pair the wrong things in a file
that otherwise looks right. Write a * branch to say what you meant.
What it will not do
- Read a group declared later. The pattern is built left to right, so a group further along has produced nothing to compare against and the branch could never be taken. Refused, rather than left to produce empty halves.
- Read anything but a named group — not another column, not
${{…}}. For cross-column logic use<switch>or aparentfilter; this is for parts of ONE value. - Compare anything but plain text. The test is
name=value, an exact match against what the group produced. A value cannot contain:— that is where the branch begins.
It leaves the quota exact
The conditional reads the weighted choice; it does not disturb it. Over 200 rows
(?<c>(?%{70:RU;20:US;10:DE})) still gives exactly 140, 40 and 20 — and each row's
second half follows from its own first half.
See also
- Regex — the finite constructs this page inherits.
- Symbol — the named alphabets
(
\a{name}). regex_max_lengthin the attribute reference.<mix>— exact percentages across structurally different branches.