Generators
A <gen> produces the values of a sequence. Its
type attribute chooses which generator to use, and every other attribute is a
parameter of that generator:
<sequence name="Status">
<gen type="text" value="new,active,closed"/>
</sequence>
active new closed active closed
Example outputs on this page are illustrative — the exact values depend on the seed and can shift between core versions. What stays fixed is the shape of the result: the format, the counts, and the distribution.
Where a generator can live
A <gen> lives where data is declared:
- inside a
<sequence>— simple, composed, compound, or conditional — where it fillscountvalues (or as many as the filtered subset holds, if the sequence has aparent). "An array ofcountvalues" is the model to reason with; the default engine produces them one row at a time as the file streams, without ever holding the array; - inside a
<case>of a<mix>— one branch of a percentage split.
Several <gen>s can share one sequence body, and name decides what each becomes.
Leave a generator unnamed and its value is concatenated into the sequence's own
value, along with any <data> literal beside it — a
composed sequence. Give it a
name and it is a field of its own, read as ${{Sequence.Field}} — a
compound sequence. The two mix
freely in one body.
It is not allowed directly in the output block. A <gen> placed as a child of
<line> is error TDC131: the block only
formats text; it doesn't generate anything. To put a generated value in the output,
declare a named sequence and reference it with ${{Name}} — see
Output & formatting.
Common attributes
These work on every generator; the rest depend on type.
| Attribute | Required | What it does |
|---|---|---|
type | yes | Which generator to use (see the table below) |
name | no | Makes this generator a field of its sequence, read as ${{Sequence.Field}}. Without it the value joins the sequence's own value |
if | no | Branch condition inside a conditional sequence — the first true <gen> wins |
comment | no | Free-form comment, ignored by the engine |
The generators
Each type has its own page, with every parameter and worked examples.
type | Produces |
|---|---|
text | A value from a set — uniform, or by exact percent |
number | An integer in a range, or a fixed-width digit string |
template | Built-in realistic data and technical IDs |
file | Values read from your own files and CSV columns |
date | A date or date-time in a range and format |
symbol | A string of characters from a set or named alphabet |
regex | A string matching a finite regular expression |
advanced_regex | Regex plus weighted choice between alternatives |
increment / decrement | Rising and falling counters |
timeseries | A time series — trend + seasonality + noise |
pattern | A distribution shaped like a drawn curve |
http | Values answered by your own service, batch by batch |
running | A total that carries down the column — a balance, a high-water mark |
stat | One number for the whole run — an average, a total, a largest |
formula | Arithmetic over the other columns of the same row |
On presets. The old type="preset" no longer exists. Algorithmic identifiers —
UUIDs, IBANs, credit-card numbers, git SHAs, national IDs — are now
template paths: global ones under the common. prefix (for example
common.id.uuid), country ones under their country (for example usa.docs.ssn). The
full catalog lives on the template and
generators reference pages.
Declared shares, or a draw from a source
Two kinds of generator sit side by side in that table, and they answer "how often does each value appear?" differently. It is worth knowing which one you are holding.
You declared the shares — you get them exactly. Where the values are written out in
the config, TDC lays the quota across the rows and then shuffles it. percent="30,70"
is 30 and 70, not "about". Left unset, the shares are equal, and equal is exact too:
text: 100 100 100 100 100 100 100 100 100 100
That covers text and <mix>, and
number when its percent splits length groups — length="2,3" percent="70,30" over a thousand rows is 700 and 300 exactly.
missing= on the same generator changes the countsThe quota is laid over the whole column first, and missing= then blanks cells without
regard to which value they hold. So percent="90,10" missing="0.5" over a thousand rows
gives about 450 / 50 / 500 blank: the RATIO of the surviving values is still 90:10, and
the absolute counts are not what percent alone would give.
That is not a rounding slip, and no ordering fixes it — the two requests are inconsistent.
Exactly 100 fail rows AND half the file blank would make fail 100 of the 500 surviving
values, which is 20%, not the 10% asked for. If you need an exact number of a value in the
finished file, keep missing= off that generator.
A plain numeric range is the other kind. value="1..10" draws, and over a thousand rows
the ten values come out 97 84 106 112 107 102 90 95 86 121 — the spread of a draw, not
a quota. The rule is what the config wrote down: shares written out are honoured exactly,
a range is reached into.
You pointed at a source — you get a draw. A file or a pool is a set you reach into, once per row, independently. Over the same 1000 rows the counts land where chance puts them:
file: 81 88 93 97 98 102 103 105 111 122 pool: 90 92 95 97 100 102 104 105 106 109
This is not a weaker version of the first. Nobody declared a proportion, so there is none to honour — a source behaves the way drawing from a hat behaves, which is what makes it look like real usage rather than a rota.
When a source needs proportions, it takes them from the data. A CSV that knows how
often each item sells says so in a column, and
weight="sales" makes the draw follow it — exactly, like
percent. That is the honest place for the numbers: a catalog of 3,000 items has its
frequencies in the file, not in your config.
Formatting on any generator
A handful of attributes work on any type. The value is generated as usual,
then reshaped on its way out. The generator itself is unaffected by which of these
you attach.
case= / mask= — letter case and display masks
Use it when the raw value is correct but should look a certain way: a column that has to be all uppercase, or a plain number that should read like a formatted ID.
case= changes letter case; mask= splits and rearranges characters into a fixed
template. Both wrap the whole generator. The example below feeds the same four
US cities through three sequences — the raw value, then the same list with
case="lower" and case="upper". order="sequential" keeps the cities in step so
the columns line up.
<sequence name="Raw"><gen type="text" value="New York,Chicago,Denver,Austin" order="sequential"/></sequence>
<sequence name="Low"><gen type="text" value="New York,Chicago,Denver,Austin" order="sequential" case="lower"/></sequence>
<sequence name="Up"><gen type="text" value="New York,Chicago,Denver,Austin" order="sequential" case="upper"/></sequence>
...
<data>${{Raw}} -> lower: ${{Low}} | upper: ${{Up}}</data>
New York -> lower: new york | upper: NEW YORK Chicago -> lower: chicago | upper: CHICAGO Denver -> lower: denver | upper: DENVER Austin -> lower: austin | upper: AUSTIN
mask= does the same for identifiers that should read as formatted: a bare
37898432363 with mask="xxx-xxx-xxx xx" comes out as 378-984-323 63. Every
mask slot (x, w, *), every case mode, and multi-step filter chains are covered
in full on Masks & case.
order= / cycle= — the order of values
Use it when values must come out in a fixed sequence rather than at random — month names in calendar order, a lookup list walked top to bottom, or two columns that have to stay aligned (as in the example above).
By default order="random". Set order="sequential" and row i takes the i-th
value in order, cycling back to the start when the list runs out. cycle="false"
turns that wrap-around into a clear error instead — useful when running out of
values should be a failure, not a silent repeat.
Both are read by the three generators that have an order to walk: text (its
comma-separated list), file (its lines or CSV column) and date (its range=, walked
in step= units). On any other type there is nothing to walk — a draw never runs out —
so the engine refuses them (TDC015) rather than accepting a request it cannot honour.
<sequence name="Rand"><gen type="text" value="Jan,Feb,Mar"/></sequence>
<sequence name="Seq"><gen type="text" value="Jan,Feb,Mar" order="sequential"/></sequence>
...
<data>random=${{Rand}} sequential=${{Seq}}</data>
random=Feb sequential=Jan random=Feb sequential=Feb random=Mar sequential=Mar random=Jan sequential=Jan random=Feb sequential=Feb random=Mar sequential=Mar random=Jan sequential=Jan
The same applies to files: <gen type="file" src="cities.txt" order="sequential"/>
walks the file line by line. Full details are on
Masks & case.
Several values per row — repeat="N"
Add a fixed repeat and the row takes N values
off the walk instead of one. The walk carries on across rows, so element k of row r
is source value r×N+k:
<sequence name="Step"><gen type="text" value="created,paid,shipped,delivered" repeat="4" order="sequential"/></sequence>
created,paid,shipped,delivered created,paid,shipped,delivered created,paid,shipped,delivered
Four steps over a list of four is the whole list on every row — which is how you write a
record's whole lifecycle in one sequence. Over a
shorter list the rows differ: value="a,b,c" repeat="2" gives a,b then c,a then
b,c.
Carrying on rather than restarting is what makes repeat="1" mean exactly what
order="sequential" alone means — the same column, not a special case.
Two shapes are refused rather than guessed at. A ranged repeat="2..5" (TDC254): a
walk advances by a fixed number of values per row, and a row whose length comes from the
length quota has no such number. A walked date
(TDC254): it carries an instant beside its text, and a row holding several dates has no
single one to give of= and plus=. distinct="true" is refused too (TDC307) — a
walked row draws nothing, so there is no draw to make without replacement.
missing= / anomaly= — blanks and outliers
Use it when you need data that looks like the real world — where some fields are
empty and a few values sit far outside the normal range. missing= injects blank
cells (missing-completely-at-random gaps); anomaly= injects outliers so a
downstream pipeline or model has something abnormal to cope with. Both attach to the
generator like case= and are applied after the value is produced.
Both attach to any generator and change what the column looks like as a whole — some
cells empty, a few values far out of range — rather than how the underlying value is
drawn. They differ in what they can act on: missing= blanks a cell whatever was in it,
while anomaly= multiplies, so it only bites on values that read as numbers. A
numeric string from a text list is multiplied; a name beside it is passed through
unchanged, and a list with no numbers at all is refused outright. Full rules on
Anomalies & missing values.