This is the crash course on Yuneta’s persistence layer. At the end you know the difference between timeranger2 (the append-only time-series log) and treedb (the graph database on top), how schemas are declared, how nodes link to each other, and which rules will ruin your day if you ignore them.
Conceptual frame. This document describes the information plane of Yuneta’s typed-graph model. The behavior plane is in
GOBJ.md. The claim that both planes share one set of primitives —topic/gclass,node/gobj,hook/subscription — is laid out in The Typed-Graph Model. Read that first if you want to know why treedb and gobj look so similar before diving into either one.
Companion to GOBJ.md. Sibling to YUNO_LIFECYCLE.md
(which uses these topics to store realms, yunos, binaries and
configurations), YUNO_AUTH.md (which uses them for users, roles
and audit), and REALMS.md (the realm hooks lifecycle).
1. Mental model¶
┌───────────────────────────────────────┐
│ your gclass calls gobj_*node() │
└───────────────────┬───────────────────┘
│
▼
┌───────────────────────────────────────┐
│ c_treedb / c_node │ gobj wrappers
│ (graph operations, in-memory hooks) │
└───────────────────┬───────────────────┘
│
▼
┌───────────────────────────────────────┐
│ tr_treedb.c │ graph layer
│ topics, nodes, hooks, fkeys, schema │
└───────────────────┬───────────────────┘
│
▼
┌───────────────────────────────────────┐
│ timeranger2.c │ append-only log
│ per-key files + md2 binary index │
└───────────────────┬───────────────────┘
│
▼
filesystem
(one directory per topic,
one subdir per key,
one .json + .md2 per day)Two distinct things:
| Layer | What it is |
|---|---|
| timeranger2 | An append-only time-series log with a key index. Stores records keyed by a primary key, time-partitioned, with a 32-byte binary metadata index for fast lookup by rowid, time, or pkey. Knows nothing about graphs. |
| treedb | A graph database that uses timeranger2 as its persistent store. Adds the notion of topics with schemas, typed columns, hooks (parent→children in-memory pointers), and fkeys (child→parent persistent references). |
If you want raw time-series, you go straight to timeranger2. If you want
a graph of typed nodes, you use treedb. The agent uses treedb for
everything: realms, yunos, binaries, configurations, users and roles. The
logcenter yuno uses raw timeranger2 to write records.
2. timeranger2¶
2.1 The on-disk layout¶
For each opened database (a top-level directory):
The same layout in text:
<database>/
__timeranger2__.json ← metadata + master lock
<topic_1>/
topic_desc.json ← {topic_name, pkey, tkey, system_flag}
topic_cols.json ← persisted cols schema ⚠ versioning trap
topic_var.json ← user-mutable per-topic flags
keys/
<key_value_a>/
2026-05-22.json ← appended JSON records, one per line
2026-05-22.md2 ← 32-byte binary index, one per record
2026-05-23.json
2026-05-23.md2
…
<key_value_b>/
…
disks/ ← non-master / cross-yuno hardlink slots
<rt_id>/
<key_value_a>/ ← hardlinks to the keys/ files
<key_value_b>/
…
<topic_2>/
…Path-building lives in kernel/c/timeranger2/src/timeranger2.c.
The data filename mask is "%Y-%m-%d" by default — each appended
record lands in the file whose mask matches its __t__. Big topics
naturally rotate every day.
2.2 Records and the md2 index¶
Each .md2 file is an array of fixed 32-byte records in big-endian
order. The struct (timeranger2.c, in-memory shape
md2_record_ex_t at timeranger2.h):
typedef struct {
uint64_t __t__; // storage timestamp + high-16-bit user flags
uint64_t __tm__; // creation timestamp + high-16-bit system flags
uint64_t __offset__; // byte offset of the record in the paired .json
uint64_t __size__; // byte size of the record
// (in memory only:)
uint16_t system_flag;
uint16_t user_flag;
uint64_t rowid;
} md2_record_ex_t;The high 16 bits of __t__ and __tm__ are reserved for flags. Macros
at timeranger2.c extract and pack them. Lookup by rowid is
O(1) — multiply by 32, seek the .md2, read offset+size, seek the
.json. Lookup by time range is O(N) over .md2 records, which is
still fast, because each record is 32 bytes.
2.3 g_rowid vs i_rowid — the rule¶
Two rowids per record, both maintained only by timeranger2:
| Name | Meaning |
|---|---|
g_rowid | Global rowid for that key — cumulative across all files, never reset |
i_rowid | Rowid within the current .md2 file — (offset / sizeof(md2_record_t)) + 1 |
tranger2_append_record (timeranger2.c:2332) computes both and
returns them in md_record_ex->rowid (timeranger2.c). Callers
never set them. For topics with sf_rowid_key, timeranger2 also
asserts g_rowid == i_rowid (timeranger2.c) — a mismatch is
a data-corruption indicator.
If you write test fixtures and you fill g_rowid by hand, stop. That is
the work of the framework.
2.4 __t__ vs __tm__¶
Both timestamps, but semantically distinct:
| Field | What it means | When it is set |
|---|---|---|
__t__ | When timeranger2 wrote the record to disk | At append time. Defaults to “now”. |
__tm__ | When the underlying event happened (from the record’s tkey field) | Caller-controlled via tkey config. |
__t__ partitions files. __tm__ is the event-time for your queries.
For records that are events as they happen, the two are usually
identical (within milliseconds). For batch imports of historical data
the two diverge — __tm__ is the original event, __t__ is “now I
imported it”.
2.5 Topic declaration¶
When you create a topic you provide a topic_desc_t
(timeranger2.h):
typedef struct {
const char *topic_name;
const char *pkey; // primary-key field name
const system_flag2_t system_flag;
const char *tkey; // time-key field name
const json_desc_t *jn_cols; // column schema
const json_desc_t *jn_topic_ext;
} topic_desc_t;system_flag bits (timeranger2.h):
| Flag | Meaning |
|---|---|
sf_string_key | pkey is a string. Directory names use it verbatim. |
sf_int_key | pkey is a uint64. Directory names zero-padded. |
sf_rowid_key | pkey is auto-generated rowid. g_rowid == i_rowid enforced. |
sf_t_ms | __t__ in milliseconds (default: seconds). |
sf_tm_ms | __tm__ in milliseconds. |
sf_zip_record | .json records are zlib-compressed. |
sf_cipher_record | .json records are encrypted. |
Persisted in topic_desc.json at create time (timeranger2.c)
and loaded on open.
2.6 Public API in 12 calls¶
timeranger2.h. Grouped by purpose:
// lifecycle
json_t *tranger2_startup (hgobj, json_t *jn_tranger, yev_loop_h);
int tranger2_stop (json_t *tranger);
int tranger2_shutdown (json_t *tranger);
json_t *tranger2_create_topic(json_t *tranger, const char *topic_name,
const char *pkey, const char *tkey,
json_t *jn_topic_ext, system_flag2_t system_flag,
json_t *jn_cols, json_t *jn_var);
json_t *tranger2_open_topic (json_t *tranger, const char *topic_name, BOOL verbose);
int tranger2_close_topic (json_t *tranger, const char *topic_name);
// append
int tranger2_append_record(json_t *tranger, const char *topic_name,
uint64_t __t__, uint16_t user_flag,
md2_record_ex_t *md_record_ex, json_t *jn_record);
// read
json_t *tranger2_open_iterator (json_t *tranger, const char *topic_name, const char *key,
json_t *match_cond, tranger2_load_record_callback_t,
const char *iterator_id, hgobj creator, json_t *data, json_t *extra);
json_t *tranger2_iterator_get_page(json_t *tranger, json_t *iterator,
uint64_t from_rowid, int limit, BOOL backward);
int tranger2_close_iterator (json_t *tranger, json_t *iterator);
// realtime
json_t *tranger2_open_rt_mem (…); // master-side realtime (writes pushed via callback)
json_t *tranger2_open_rt_disk(…); // non-master realtime (watches hardlinks)tranger2_open_rt_disk is the workhorse for cross-yuno reads —
see §4.5.
2.6b The two time axes (t and tm)¶
Every record carries two timestamps, and they are independent:
| Axis | Meaning | Its source |
|---|---|---|
t | Persistence time — when the record was appended | the __t__ argument of tranger2_append_record (now, if 0) |
tm | Message time — when the event it carries happened | the record’s tkey field (usually tm), set by the producer |
They diverge whenever data is backfilled or a device uploads a buffer late.
Both are in the topic’s unit: seconds, or milliseconds when the topic
sets sf_t_ms / sf_tm_ms (read system_flag from the topic desc — over the
wire, topics expanded=1).
The match_cond of an iterator takes a range on each axis (from_t and
to_t, from_tm and to_tm), the from_rowid and to_rowid pair, and
the user_flag conditions,
and ANDs them. Every condition is honored per record: a filtered paging
iterator builds its row index when it opens, so tranger2_iterator_size(),
pages and the pages themselves count only matching records — and
get_page’s from_rowid is then a position among the matching rows, not a
global rowid. An unfiltered iterator builds no index (its open stays cheap
regardless of key size) and its positions are the global rowids.
list-keys reports, per key, records plus the key’s span on both axes
(fr_t/to_t, fr_tm/to_tm), read from the topic’s in-memory cache totals —
so a client can bound a time picker to the content of the key, and it reads
no record.
Note (in the md2 record, times carry flags). On disk the 16 high bits of
__t__hold theuser_flagand those of__tm__thesystem_flag. Always read them throughget_time_t()orget_time_tm(). The raw field gives you a timestamp that still contains the flags.
2.7 Master / non-master¶
tranger2_startup (timeranger2.c:330) attempts an exclusive
lock on __timeranger2__.json. Whoever gets it is the master:
The master can read AND write. Only the master can call
tranger2_append_record,tranger2_delete_topicand the other write functions.Non-masters can only read. They are expected to use
tranger2_open_rt_diskso the master can push updates to them via hardlinks in thedisks/<rt_id>/directory.
The lock is held for the lifetime of the process. If a master crashes without releasing, the OS releases the flock on exit and the next yuno that opens the database becomes master.
2.8 Snapshots¶
The current timeranger2 API does not expose a snapshot primitive
named tranger2_*_snap* — those calls live one layer up at the treedb
level (§3.7). The closest underlying mechanism is the disks/<rt_id>/
hardlink trick that gives non-masters a consistent view at the point
the directory was wired.
2.9 The delete-record story¶
Two granularities, both implemented in v7 as of 2026-05-26.
Whole record (= a primary key + every instance under it).
tranger2_delete_key()(renamed fromtranger2_delete_recordon 2026-05-25. A#defineintimeranger2.hkeeps the legacy alias). Removeskeys/<key>/and drops the key fromtopic_cache. Irrecoverable. Used today bytreedb_delete_node.One instance (one row in the
.md2file).tranger2_delete_instance(tranger, topic, key, __t__, rowid, zero_payload)mutates the.md2row in place withsf_deleted_instance = 0x0400(back insystem_flag2_t, inherited side of the mask sort_by_diskfollowers see the tombstone). Optionalzero_payloadoverwrites the matching__size__bytes in the data.jsonfor sensitive-data wipes. Read paths (tranger2_open_iteratorhistory,tranger2_iterator_get_page,publish_new_rt_disk_records) skip dead rows. Master-only, idempotent. Slot ids do NOT renumber —iterator_size/total_rowskeep counting slots, not live rows. Treedb is NOT a consumer:treedb_delete_instance()is per-pkey2-index in-memory cleanup only.
Propagation to subscribers (2026-05-26)¶
tranger2_delete_key() now notifies every subscriber tracking
the deleted key. Two paths:
In-process (rt_mem, rt_disk in the same yuno as the master, open_iterator): a registered
tranger2_key_deleted_callback_tfires for each handle whosekeyfilter matches (""= any).Across-process (
rt_by_diskfollowers): the masterrmrdirstopic/disks/<rt_id>/<key>/BEFORE the livekeys/<key>/so the follower’s inotify watcher catches it asFS_SUBDIR_DELETED_TYPE, which fires the follower’skey_deleted_callback.fire_key_deleted_locally()is split by transport (fs_followersflag): the master in-process call serves non-watcher subscribers, the inotify branch serves the fs-watcher followers — each subscriber fires exactly once. No new IPC channel, no new file convention. (The inotify branch firing the fs-watcher callbacks was completed 2026-05-28. Before that date the shared distribution skipped them, and live deletes were dropped with no message. See the CHANGELOG.)
Register with:
tranger2_set_rt_key_deleted_callback(handle, cb, user_data);…on any handle returned by tranger2_open_rt_mem,
tranger2_open_rt_disk or tranger2_open_iterator. Pre-2026-05-26
followers that polled their cache on a timer can drop the timer.
Memory: project_tranger2_delete_record_deferred.
2.10 Durability¶
tranger2_append_record performs the write but does not fsync
(timeranger2.c). Durability is whatever the OS gives you —
on EXT4 with the default journal, that is “data on disk within the
journal commit interval, usually 5 s”. If you need stronger guarantees,
add an explicit fsync in the wrapping code, but understand the
throughput cost.
3. treedb¶
3.1 The graph model¶
A treedb sits inside a tranger. Topics become entity types, nodes
become records keyed by id, hooks are in-memory pointers from parent
nodes to their children, fkeys are persistent references from child
nodes to their parent. Schema is JSON.
The schemas already documented in this repo’s docs cover the canonical examples:
YUNO_LIFECYCLE.md§2.1-2.3 —binaries,configurations,yunos.REALMS.md§2 —realms.YUNO_AUTH.md§4.1 —users,roles,users_accesses.
Read those for the operational shape. This section explains how the schema works.
3.2 Topic schema JSON¶
A real, minimal example (yuno_agent schema, paraphrased):
{
"id": "yunos",
"schema_version": 1,
"topic_version": 19,
"pkey": "id",
"pkey2s": "yuno_release",
"tkey": "",
"system_flag": "sf_string_key",
"cols": {
"id": { "type": "string", "flag": ["persistent", "required"] },
"realm_id": { "type": "string", "flag": ["fkey"],
"fkey": { "realms": "yunos" } },
"yuno_role": { "type": "string", "flag": ["persistent", "required"] },
"configurations": { "type": "object", "flag": ["hook"],
"hook": { "configurations": "yunos" } }
}
}Six things to notice:
pkey— column name that serves as the primary key. Maps totopic_desc_t.pkey.pkey2s— optional secondary key (composite). Allows multiple records per primary key, for example several versions of a binary.treedb_get_instance(),treedb_list_instances()and the agent’sinstancescommand query them. Invariant (since dbf532ec9): the pkey2 secondary index shares the SAME node object as the primary index, andtreedb_save_node()points it again on every runtime save. Before that correction it held a separate object that only the disk-load filled. A runtimeupdate-nodewas therefore invisible throughlist_instancesuntil the next reload. That was the bug behindlist-binaries, which showed a stale binary immediately afterupdate-binary.schema_versionandtopic_version— these are different. Schema is the overall layout. Topic is per-topic. Raisetopic_versionevery time you changecols— §3.5.colsdeclares typed columns. Type + flag list (next section).fkeyfield on the child points at (parent topic, hook name). Persisted.hookfield on the parent points at (child topic, child fkey name). Rebuilt in-memory at load time.
3.3 Column types and flags¶
Column types live in the JSON spec, parsed by tr_treedb.c. Common
ones: string, integer, boolean, real, array, object,
blob, enum, wild. Plus semantic decorations: email, url,
password, time.
Flags (parsed by kw_has_word throughout tr_treedb.c):
| Flag | Effect |
|---|---|
persistent | Written through to timeranger2 on save. |
required | Cannot be null at creation. |
notnull | Cannot be null ever. |
hook | Parent → children link. In-memory only (rebuilt on load from children’s fkeys). |
fkey | Child → parent reference. Persisted. Encoded as topic^parent_id^hook_name. |
pkey | Marks the primary-key column. |
pkey2 | Marks a secondary key. |
tkey | Marks the time-key column. |
password | Treated as opaque secret on inspection. |
email/url/enum/wild | Semantic types, mostly informational. |
inherit | Inherits a value from a related node. |
Absence of persistent + absence of hook/fkey means volatile —
in-memory only.
3.4 The __md_treedb__ metadata block¶
Every loaded node carries a metadata sidecar (tr_treedb.c,
attached at tr_treedb.c):
"__md_treedb__": {
"treedb_name": "treedb_yuneta_agent",
"topic_name": "yunos",
"g_rowid": 14,
"i_rowid": 14,
"t": 1737499200,
"tm": 1737499200,
"tag": 0,
"pure_node": true
}g_rowid,i_rowid— see §2.3. Never set them yourself.t,tm— the timeranger2 timestamps, surfaced to the node level.tag— user_flag from md2, used for snapshots (§3.7).immutable— present only when set (omitted on ordinary nodes).truemeans the record carries thesf_immutable_recordmd2 bit and cannot be deleted. See §3.10.pure_node— true for ordinary nodes. This is the metadata that you read.
A node that appears in multiple places in a JSON dump (once under the
topic’s id index, once nested inside its parent’s hook) carries the
same __md_treedb__ everywhere. Same record, multiple views.
3.5 The topic_cols.json versioning trap¶
Memory
feedback_treedb_schema_versioning:
Any
colschange needs a highertopic_version. If you do not raise it, the persistedtopic_cols.jsoncontinues to mask the new schema. Deletestore/when you reproduce the problem.
What happens: treedb_open_db() (tr_treedb.c:485) reads the
persisted topic_cols.json and compares its topic_version against
the schema in code. If they match, the persisted file wins. If you
edited the schema in code but forgot to bump topic_version, your
running yuno sees the old schema and silently ignores any new
columns you added.
The fix:
Bump
topic_versionin the schema JSON every time you changecols.While debugging schema problems, wipe the topic’s directory in
store/to force a clean load.
3.6 Node CRUD: the public API¶
Two layers — treedb_* (the low-level graph API) and gobj_*node (the
gobj wrappers most user code uses).
Low-level (tr_treedb.h):
json_t *treedb_create_node(json_t *tranger, const char *treedb_name,
const char *topic_name, json_t *kw);
json_t *treedb_update_node(json_t *tranger, json_t *node, json_t *kw, BOOL save);
int treedb_delete_node(json_t *tranger, json_t *node, json_t *jn_options);
json_t *treedb_get_node (json_t *tranger, const char *treedb_name,
const char *topic_name, const char *id);
json_t *treedb_list_nodes (json_t *tranger, const char *treedb_name,
const char *topic_name, json_t *jn_filter,
BOOL (*match_fn)(json_t *node, json_t *jn_filter));
// links (graph operations)
int treedb_link_nodes (json_t *tranger, const char *hook_name,
json_t *parent_node, json_t *child_node);
int treedb_unlink_nodes(json_t *tranger, const char *hook_name,
json_t *parent_node, json_t *child_node);gobj-level wrappers (gobj.h):
json_t *gobj_create_node(hgobj, const char *topic, json_t *kw, json_t *opt, hgobj src);
json_t *gobj_update_node(hgobj, const char *topic, json_t *kw, json_t *opt, hgobj src);
int gobj_delete_node(hgobj, const char *topic, json_t *kw, json_t *opt, hgobj src);
json_t *gobj_list_nodes (hgobj, const char *topic, json_t *filter, json_t *opt, hgobj src);
int gobj_link_nodes (hgobj, const char *hook,
const char *parent_topic, json_t *parent_rec,
const char *child_topic, json_t *child_rec, hgobj src);
int gobj_unlink_nodes(hgobj, const char *hook,
const char *parent_topic, json_t *parent_rec,
const char *child_topic, json_t *child_rec, hgobj src);Most production code calls gobj_*node. Those functions route to the right
treedb from the priv of the gobj, and they integrate the authzs and the
traces.
What a write validates (since 7.13.0 — before it, less than this):
| create | update | |
|---|---|---|
| Type of each field, per the topic’s cols | yes | yes |
required on a missing field | yes | n/a (the node already has one) |
notnull | yes | yes |
enum membership | yes | yes |
An update used to store whatever it was handed: no type, no notnull, no
enum. And enum was checked only when a schema was parsed, never when a
node was written, so the list a column declares did not survive the first
write — on either path. Both now run the same normalization, and an update
validates every incoming field before touching the node, so a refusal
leaves nothing half-applied.
A refused write returns NULL (-1 for links), and cmd_create_node /
cmd_update_node answer result: -1 with the cause. Until 7.13.0
mt_update_node dropped the return of treedb_update_node and answered the
collapsed view of the unchanged node — a refused update read as a success.
3.7 The link/unlink-saves-child rule¶
CLAUDE.md hard rule, reproduced verbatim from tr_treedb.c:
PUBLIC int treedb_link_nodes(...) {
_link_nodes(gobj, tranger, hook_name, parent_node, child_node, FALSE);
/*---Save persistent: Only children are saved---*/
return treedb_save_node(tranger, child_node); // ← only child
}
PUBLIC int treedb_unlink_nodes(...) {
_unlink_nodes(gobj, tranger, hook_name, parent_node, child_node, FALSE);
/*---Save persistent: Only children are saved---*/
return treedb_save_node(tranger, child_node); // ← only child
}The rule: link/unlink writes the child to disk, never the parent.
Why: the persistent reference lives on the child (the fkey field).
The parent’s hook field is in-memory and gets rebuilt on the next
load by scanning all children for fkey == parent.id.
Two consequences:
After
treedb_link_nodes, the child’sg_rowidadvances by 1 (one new record appended). The parent’sg_rowiddoes not change.If you write tooling that snapshots state by reading rowids, the parent’s rowid is a bad signal of “has anything happened to this node’s relationships” — you have to look at the children too.
3.8 Cross-yuno reads: the rt_by_disk pattern¶
When a non-master yuno needs to read another yuno’s store, it opens
the master’s database in read-only mode and registers an
rt_by_disk watcher. The master, on every change, writes hardlinks
into disks/<rt_id>/ for that subscriber. The subscriber’s
filesystem watcher fires, and it re-reads the hardlinks.
Memory
feedback_cross_yuno_via_store_not_command:
in wattyzer (and by extension other multi-yuno SPAs), cross-yuno
queries from the SPA go through db_history_wz reading B+ yunos’
stores non-master via this pattern. cmd_command_yuno does not
work for B+ yunos, because they do not publish their service through
__top_side__. The store path is the correct one.
Code: tranger2_open_rt_disk at timeranger2.h. The
mechanism is purely filesystem-mediated — no socket between the master
and the watchers.
3.9 Snapshots (treedb-level)¶
Snapshots tag a point in time across the treedb. APIs at
tr_treedb.h:
int treedb_shoot_snap (json_t *tranger, const char *treedb_name,
const char *snap_name, const char *description);
int treedb_activate_snap(json_t *tranger, const char *treedb_name,
const char *snap_name);
json_t *treedb_list_snaps (json_t *tranger, const char *treedb_name,
json_t *jn_filter);gobj_list_snaps(gobj, filter, src) is the gobj-level wrapper.
Snapshots are how the agent picks which binary version to run when
multiple are stored — see YUNO_LIFECYCLE.md §4.3. The
binary resolver tries the active snapshot first
(gobj_list_snaps, c_agent.c). If that fails, it does a
direct (role, role_version) lookup.
3.10 Immutable nodes and non-deletable topics¶
Some records must never be deleted by CRUD (the seed root role and
yuneta user — see YUNO_AUTH.md §4.2), and some topics
must never be dropped (the __system__ treedb’s structural topics, and
every treedb’s __snaps__ / __graphs__). The protection is metadata,
never a data column — it does not touch the user schema and never bumps
topic_version. Design write-up:
DESIGN-immutable-topics-records.md.
Record level rides a free md2 system_flag bit, sf_immutable_record
(0x0800, inherited band) — the same metadata channel as the snapshot
tag, persisted on disk and decoded on every load:
Set it with
treedb_set_node_immutable(tranger, node, set), which rewrites the node’s current primary record in place (no new record) via the gatedtranger2_set_system_flag(), and flips__md_treedb__immutable` in memory.treedb_save_node()re-stamps the bit after every update (the re-append inherits only the topic-defaultsystem_flag, so the bit is re-applied exactly liketag).treedb_delete_node()andtreedb_delete_instance()refuse an immutable record, andforcedoes NOT override (stronger than the snapshot-tag guard).tranger2_delete_instance()carries the same refusal as a backstop.Because the mark is not a JSON field, a client cannot inject it via
create-node/update-node— only an in-processmt-level caller can set it. No strip boundary needed.
Topic level rides system_topic: true in the topic’s topic_var.json
(additive, no topic_version bump). Declare it in the schema next to
topic_version, or pass system_topic=TRUE to treedb_create_topic().
treedb_delete_topic() (and tranger2_delete_topic() as a backstop) refuse
it. A system topic’s records stay deletable — only the topic is frozen.
Out of scope on purpose: delete-treedb / a whole-store rm -rf. This
protects against CRUD/control-plane deletion, not against an operator wiping
the realm — “only a full store wipe removes them”. Regression coverage:
tests/c/tr_treedb_immutable.
3.11 The __system__ treedb: a schema stored as data¶
A schema has two homes. The one you write is the C literal
(treedb_schema_*.c), persisted as
<tranger_dir>/<treedb_name>.treedb_schema.json on first open (§3.5). The
other is the __system__ treedb, which every C_TREEDB service builds
next to the treedbs it manages, at <path>/__system__. There the same
schema is stored as ordinary treedb data:
treedbs ── id, schema_version, c_schema_version,
system_schema_version ──hook topics──▶
topics ── id (rowid), value, pkey, pkey2s, system_flag, tkey,
topic_version, system_topic ──hook cols──▶
cols ── id (rowid), value, header, fillspace, type, placeholder,
flag, enum, template, hook, pkey2s, default,
description, propertiesIts schema is treedb_system_schema.c, and it is the reason a schema can be
read, listed and edited at runtime with the same nodes / create-node /
update-node commands as any other data — no new command surface.
topics and cols are keyed by rowid, and the name lives in value. A
name is unique only inside its parent: two topics with an id column would
collide on a single cols topic keyed by name, and two treedbs with a users
topic would collide on a single topics topic keyed by name — users is a
topic of authzs, mqtt_broker and controlcenter alike. So addressing
either one costs its rowid, not its name (fetch_node needs id; a pkey2
only refines a primary lookup). This is also why the
descriptor used to validate a user column is derived, not copied, from that
topic: _treedb_create_topic_cols_desc() renames value back to id and
drops the storage-only fields (id, topics, _geometry). Add a field for
user columns to the cols topic; add a storage-only field there and to
that skip list.
Who fills it, and who wins. C_TREEDB’s open-treedb projects the C
literal into __system__ the first time it sees a treedb, and afterwards only
when the literal is strictly newer, exactly as schema_version already
decides between the literal and the persisted schema file (§3.5). Raising the
version is how either side publishes a change.
The treedbs node carries three numbers, and they are not
interchangeable:
| Written by | Means | |
|---|---|---|
schema_version | whoever edits the schema (an editor raises it on save) | what this schema is worth to treedb_open_db |
c_schema_version | only the projection | which version of the C literal this projection came from |
system_schema_version | only the projection | which version of the meta-schema produced it |
The third one exists because a projection is a function of two things: the
literal, and the meta-schema that says how a schema is stored and projected.
Comparing only the literal froze a projection made by an older SDK forever —
and an older SDK is exactly the one whose projection may be missing what it did
not know how to store yet. So raising the meta-schema’s own schema_version is
the lever that re-projects every store on the next start, and it moves when
the projector changes even if no field does.
Reconciliation compares the literal against c_schema_version. Sharing one
counter would mean that the first edit made here — which has to raise
schema_version to reach the treedb at all — silently outranks every later
release of the literal, and nothing would say so. A re-projection also
publishes under max(stored, literal) + 1, or the persisted schema file,
sitting at the edited number, would keep masking it. Stores projected before
c_schema_version existed fall back to schema_version.
Reconciling is an upsert — nothing is ever deleted. A column’s id is a
rowid handed out from the topic size, so re-creating columns renumbers all of
them and can hand a retired number to a different column; and a delete is the
one destructive primitive of the store, which drops the schema’s own history
(the reason to keep a schema in a treedb at all) and refuses a snapshot-tagged
node. An update appends a new version instead, so what a column used to declare
stays readable with instances. What exists in __system__ and not in the
incoming schema is left alone: it is indistinguishable from an operator
addition, and removing a topic or a column is a deliberate action, never a side
effect of an upgrade.
A treedb opens from its projection, always. There used to be a flag
(use_internal_schema) to open from the literal instead, and with it an edit
made in __system__ reached nothing until every yuno’s config was changed one
by one. It distinguishes nothing now: the projection is seeded from the literal
and re-made whenever the literal or the projector moves ahead, so opening from
it is opening from the literal until somebody edits it — which is the point.
The literal remains the fallback, for a projection that cannot be rebuilt into
a valid schema.
The schema file still has the last word. Whichever home supplies the
schema, treedb_open_db compares its schema_version against the persisted
<treedb_name>.treedb_schema.json and the file wins on ties (§3.5). Same
rule again, one layer down. So a change reaches a running treedb only if
schema_version moved on the treedbs node and topic_version on each
topic touched — the second is what regenerates topic_cols.json, and without
it the new columns exist in the schema and not in the topic.
You do not raise them: the write does. A change that forgets either does
nothing and says nothing, so leaving the rule to whoever writes means every
editor, script and console has to carry it — and it is easy to get wrong even
while looking at it. Writing a cols or topics node of __system__
therefore raises the versions that publish it, walking up the fkeys to the
column’s topic and its treedb. The projector sets them itself and marks the
tranger while it works (__schema_publishing__), which is also what stops the
rule from answering its own writes.
A write here is a schema change, so it answers to the rules of a schema. On top of the ordinary validation of §3.6, writes to these topics are refused when they could not produce a working schema — at the point of writing, because none of these is loud later:
a column is checked against the descriptor a user column answers to, the same one
parse_schema_cols()applies when a schema is opened. Stored unchecked, the column breaks the treedb at its next open, far from whoever wrote it;pkeymust beidandsystem_flagmust besf_string_key—treedb_open_db()silently drops a topic that disagrees;pkey,tkeyandsystem_flagcannot change once the topic exists:topic_desc.jsonis written at creation and never rewritten, so the change would be stored here, shown by every reader, and ignored by the topic for good;two columns with the same name in one topic are refused when the column is linked to it, which is when the clash becomes real. The name is the key a schema is rebuilt by, so a duplicate drops one of the two definitions on the next read.
Applying an edit: pause-yuno + play-yuno, never close-treedb. An
edited schema reaches a running treedb only when the treedb is reopened, and
the reopen has to be driven by the yuno that opened it. close-treedb
destroys the treedb’s C_NODE and its C_TRANGER, and an owner typically
keeps raw handles that no framework cleanup can reach — the service pointer,
the tranger json_t read from it, copies of both on a hot path, and whatever
else it opened on that same tranger (db_history_co opens its msg2db_alarms
there). Called from outside on a playing yuno, the next record processed
writes into released memory. Every in-tree consumer therefore closes only from
mt_pause and reopens in mt_play, which re-acquires every handle; from
outside, that pair is pause-yuno + play-yuno, and it does not restart
the process. cmd_close_treedb refuses while the yuno plays (force=1 for a
caller that holds nothing of the treedb). Note pause stops the yuno’s other
services too, so its gate goes down for the cycle.
Round-trip coverage:
tests/c/c_treedb_system_schema.
Known gap: delete-treedb (delete_client_treedb_schema()) removes the
parent node before its children and passes collapsed views where pure nodes
are required. It does not work, and it is unrelated to the data of the client
treedb, which it never touches.
4. Sharp edges¶
4.1 g_rowid and i_rowid are read-only to user code¶
(§2.3, §3.4.) Never set them in test fixtures, code that calls
treedb_create_node, or anywhere else. timeranger2 computes them and
shows them in __md_treedb__ for inspection only.
4.2 link/unlink saves the child, not the parent¶
(§3.7.) If you read g_rowid on the parent after a link operation and it
did not change, that is correct. Read the g_rowid of the child instead.
4.3 Schema changes need a higher topic_version¶
(§3.5.) A stale topic_cols.json overrides new code, and it gives no
message. The trap is worse because the yuno still works. For treedb
the new columns do not exist. Always raise the version.
4.4 Master-only writes¶
(§2.7.) tranger2_append_record does nothing on a non-master and returns
-1. If you write in a yuno that is the non-master, you have a deployment
bug: two yunos opened the same store.
4.5 timeranger2 is append-only — with two scoped deletes¶
(§2.9.) Nothing ever rewrites the .json data log itself. Appends go
to the end, and nothing else changes. What is mutable is the .md2 index, and
two delete primitives operate on it:
tranger2_delete_key()removes a key’s directory wholesale (every instance with it) and propagates the deletion to in-process and cross-process subscribers via inotify + callback fan-out.tranger2_delete_instance()tombstones one row of the.md2index in place (bitsf_deleted_instance = 0x0400). Readers skip it, and rowids do not renumber. Opt-inzero_payloadoverwrites the matching bytes in the.jsonfor GDPR-style wipes.
Both are master-only and irrecoverable. The append-only contract still holds at the data-log level — only the index is mutated.
4.6 No fsync after append¶
(§2.10.) Durability is what the OS gives you. For audit logs where
a crash window of a few seconds is unacceptable, add an explicit
fsync — but understand the throughput cost.
4.7 Do not open the same store twice in the same process¶
tranger2_startup caches by path. Two starts of the same path return
the same tranger handle, but two distinct yunos in the same process
trying to coexist on the same store is unsupported.
4.8 The deprecated range_ports/last_port columns on realms¶
(See REALMS.md §7.1.) Same class of trap as §3.5:
columns that the schema still declares but the runtime ignores. Reading
them returns stale data. Trust the agent’s own attrs, not the schema
column.
4.9 Multiple node occurrences in dumps share one g_rowid¶
A node listed under topic.id_index[id] and also nested inside a
parent’s hook array is the same record. They share the
__md_treedb__.g_rowid. Do not count it twice when you compute stats from
a dump.
4.10 Hooks rebuild on load — only fkeys persist¶
(§3.7.) In the database on disk you find the fkeys of the children but not the hooks of the parents. Hooks are in-memory pointers only, and treedb builds them again when it scans the children. This is why a corrupt fkey on a child makes the hook of its parent look short. Read the child first.
4.11 No raw malloc / free for treedb-allocated json_t¶
CLAUDE.md hard rule. gbmem_* everywhere. Jansson is routed through
gbmem_*, so all json_* APIs are safe. Never free() a json_t
yourself.
4.12 Do not cache a json_t * from treedb_get_node across a¶
restart
The pointer is valid for the life of the loaded tranger. After a
tranger2_stop and tranger2_startup cycle the pointer is stale. If
you keep references across stops, the framework does not detect it. Your
crash does.
4.13 Link events are OFF by default — and turning them on REMOVES an event¶
C_NODE publishes EV_TREEDB_NODE_LINKED / EV_TREEDB_NODE_UNLINKED
only when its with_link_events attr is set (SDF_RD, default
false). Two things bite here:
It is an either/or, not additive. With the flag ON, a link/unlink publishes the link event and stops publishing the backward-compatible
EV_TREEDB_NODE_UPDATEDof the parent. So enabling it on a treedb that also serves an older consumer changes what that consumer receives. Check every subscriber before flipping it.The compat event names the wrong node for edge tracking. An edge is a fkey of the child (§4.2, link-saves-child), but the compat path announces the parent — whose fkeys did not change. A consumer that derives edges from fkeys therefore sees “a node was updated” and correctly concludes there is nothing to redraw, so its graph shows stale edges. That is the reason for the dedicated link events. Their kw is the relationship, not a node:
{hook_name, parent_topic_name, child_topic_name, parent_id, child_id, treedb_name}— note there is notopic_name, so a per-topic subscription filter matches nothing (filter bytreedb_name).
5. Recipes¶
5.1 Browse a topic from the CLI¶
yutils/c/ylist/ ships ylist for this. Without it, raw find + jq:
# every record in the realms topic (date-partitioned)
cat /yuneta/store/agent/treedb_yuneta_agent/realms/keys/*/*.json | jq .
# specific node
cat /yuneta/store/agent/treedb_yuneta_agent/yunos/keys/<id>/*.json | jq .For machine-friendly access, prefer ycommand against the agent
(list-yunos, list-realms, list-binaries, list-configs) —
those go through the treedb’s in-memory state and apply schema
correctly.
5.2 Add a new column to an existing topic¶
"cols": {
…
+ "my_new_field": { "type": "string", "flag": ["persistent"] }
},
- "topic_version": 19
+ "topic_version": 20Without the topic_version bump the field will be silently ignored
on load. With it, treedb migrates: every existing node gets the
column with its default value on first save.
For a hot rollout in which you cannot restart the yunos:
Update the schema file in source. Bump
topic_version.Build + redeploy (see
YUNO_LIFECYCLE.md§6.2).Verify the new field shows up:
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=list-nodes topic=<topic>'
5.3 Create a node and link it to a parent¶
In C, inside an action or command handler:
json_t *node = gobj_create_node(
gobj,
"users",
json_pack("{s:s, s:b}", "id", "alice", "disabled", 0),
NULL,
src
);
json_t *parent = gobj_get_node(gobj, "roles",
json_pack("{s:s}", "id", "operator"), NULL, src);
gobj_link_nodes(gobj, "users",
"roles", parent,
"users", node,
src);
// note: only `node` has been saved (the child with the fkey).
// `parent` is unchanged on disk.5.4 Inspect snapshots¶
ycommand -c 'command-yuno id=<yuno> service=__yuno__ command=list-snaps'Snapshots are global to a treedb. You see one entry per “tag”.
5.5 Recover from a botched schema change¶
# 1. stop the yuno that owns the store
ycommand -c 'kill-yuno id=<yuno>'
# 2. wipe the topic's data (do NOT do this in production — this is
# for fresh-checkout / dev-loop recovery)
sudo rm -rf /yuneta/store/<realm>/<yuno>/treedb_<name>/<topic>/
# 3. restart — the topic is recreated from the schema
ycommand -c 'run-yuno id=<yuno>'For production, do this against a backup. Never rm -rf a live store.
5.6 Read another yuno’s topic non-master (rt_by_disk)¶
Pseudocode in a different yuno that does NOT own the store:
json_t *tranger = tranger2_startup(gobj, json_pack(
"{s:s, s:b}",
"path", "/yuneta/store/<other_yuno>",
"master", false
), yev_loop);
tranger2_open_rt_disk(
tranger,
"events",
"*", // every key
NULL, // no extra filter
my_on_record_callback,
"my_unique_rt_id", // mandatory unique id
gobj,
NULL
);The master will detect your disks/my_unique_rt_id/ directory and
start writing hardlinks there on every change. Your callback fires
as soon as the kernel notifies the filesystem watcher. No socket
between the two yunos — pure inode plumbing.
6. Code pointers¶
| What | Where |
|---|---|
| timeranger2 public API | kernel/c/timeranger2/src/timeranger2.h (747 lines) |
| timeranger2 runtime | kernel/c/timeranger2/src/timeranger2.c (~7.8k lines) |
md2_record_t (32-byte index) | timeranger2.c |
md2_record_ex_t (in-memory) | timeranger2.h |
system_flag2_t (sf_string_key, sf_int_key, …) | timeranger2.h |
| Master / non-master lock | timeranger2.c |
tranger2_append_record | timeranger2.c (g_rowid set at 2667, i_rowid at 2634) |
tranger2_open_rt_disk (cross-yuno reads) | timeranger2.h |
| TRACE_FS sites | timeranger2.c (multiple) |
| treedb public API | kernel/c/timeranger2/src/tr_treedb.h (617 lines) |
| treedb runtime | kernel/c/timeranger2/src/tr_treedb.c (~8.9k lines) |
__md_treedb__ builder | tr_treedb.c |
Topic schema loader (topic_cols.json) | tr_treedb.c |
topic_version matching | tr_treedb.c |
treedb_link_nodes / treedb_unlink_nodes | tr_treedb.c (saves child only) |
treedb_create/update/delete/get/list_node[s] | tr_treedb.h, tr_treedb.c |
| Snapshot API | tr_treedb.h |
gobj wrappers (gobj_*node) | gobj.h |
gobj_list_snaps | gobj.h |
| Canonical schemas | yunos/c/yuno_agent/src/treedb_schema_yuneta_agent.c, kernel/c/root-linux/src/treedb_schema_authzs.c |
| Treedb gclass (gobj wrapper) | kernel/c/root-linux/src/c_treedb.c, c_node.c |