Changelog¶
v7.26.4 (2026-10-05)¶
The tests and their tooling, after four rounds of cloud verification of the
suite’s speed work (REV-CLOUD.md). One real defect fixed: 18 tests linked
the kernel archives by a bare name and ran on the old library after a
change of the one they test. No change in the libraries or in the yunos, so
a node running 7.26.x needs no upgrade. The suite passes in parallel on the
dev machine (298/298, ctest -j8) and on wattyzer (298/298, -j8). The
yunetas CLI goes with it: 0.21.2.
The tests that assert a time window of the wall clock run alone under
ctest -j(RUN_SERIAL):test_yevent_timer_once1,_once2,_periodic1,_kept_after_post,test_c_timer,test_c_timer0,test_static_resolv_numeric.test_yevent_timer_once1failed once in the cloud verification’s parallel suite (1.80 s, window 0.9-1.1 s). Andtimeranger2/test_fs_watcher_overflow, which asserts the watcher’s cost per directory on a big tree against a small one (at most 6 times): wattyzer’s parallel suite (-j8, load ~4) failed it with 48 us against 713; it passes alone in 31 s. Andtest_c_controlcenter_scenarios, whose rate of a burst over its 1.5 s tick fails when the load stretches the tick over 50% (the last test of that class, by the fifth check ofREV-CLOUD.md).scripts/check_test_databases.pyno longer erases why a test failed. ctest truncatesTesting/Temporary/LastTest.logon every run, its--show-onlytoo: the check keeps the file and puts it back. And the CLI 0.21.2 runs ctest with--output-on-failure, so the output of a failed test is in the keptbuild/<date>.j<N>.txtas well.scripts/check_md_fences.py: a fenced block whose closing fence has text glued after it does not close, and the rest of the page renders as code;docs/doc.yuneta.io/deploy.shrefuses to build such a page. It happened intest_suite.md(fixed).check_test_links.pyreports by line.18 tests are relinked again when a kernel library changes. The 14
tr_treedb_*tests,tr_dt_unknown,perf_timeranger2,perf_tr_treedbandperf_rotatorynamedlibtimeranger2.a,libyev_loop.a,libytls.aandlibyunetas-gobj.abare intarget_link_libraries(): a-lsearch, not a dependency of the link. After a change in timeranger2 or gobj -- what they test -- they ran on the OLD library, silently. Themake cleanofyunetas testhid it, and then (7.26.1-7.26.2) the reinstall of the root tree. They name them${LIB_DEST_DIR}/lib*.anow: a marker appended togobj.creaches every test binary that holds gobj code (286), where it reached 268. Found by the cloud verification ofREV-CLOUD.md.scripts/check_test_links.pyfails on an archive named bare in aCMakeLists.txtoftests/c,performance/corstress/c.scripts/check_test_databases.py: two tests that can use one directory at the same time are found, ascheck_test_ports.pyfinds two that use one port. It reads the directories of each test from its sources (also the"database"key of aC_NODE/C_TREEDBkw, abuild_path()segment afterpath_root, and the domain ofregister_yuneta_environment()) and itsRESOURCE_LOCK/RUN_SERIAL/DEPENDS/ fixtures from ctest, and exits 1 when two tests name one directory (or one under the other) and ctest can run them at once. The tree passes it today: 185 directories, the four shared ones each under a lock.timeranger2/test_tr2checkruns thetr2checkof its own checkout (utils/c/tr2check/build/tr2check), not/yuneta/bin/tr2check: that one is machine-wide, the last one installed, and a run from a worktree (the A/B against the last tag, a node’s release suite) tested somebody else’s. A missing tool fails as “tr2check not found at ”, not as a tool that answered no json.scripts/check_test_links.pyreads everyCMakeLists.txtoftests,performance,stress,yunos,utilsandmodules, comments left out: alib*.awithout a directory anywhere (also in aset()), and-lfooor a barefoointarget_link_libraries()whenlibfoo.ais an archive ofoutputs/liboroutputs_ext/lib. It fails when a directory cannot be read, and says how many files it checked. It found two more:pkey_to_jwks(libssl.a,libcrypto.a, now${EXT_LIB_DIR}/...) anddba_postgres(libyunetas-c_postgres.a, a name no build makes, now${MODULE_POSTGRES}). The release checklist runs the three checks of the tests before a tag.perf_c_tcp/test4andperf_c_tcps/test4no longer hold theperf_topic_integerlock, which onlytest5needs (both stayRUN_SERIAL).tests/c/README.mdandperformance/c/README.mdsay what a new test must declare to run in parallel, and the three checks.The
yunetasCLI 0.21.1:yunetas testruns ctest in parallel from SDK 7.26.3 on (not 7.26.1),yunetas buildtakes--jobs, and the ctest log isbuild/<date>.j<N>.txt, so serial and parallel timings are not mixed.
v7.26.3 (2026-10-05)¶
A fix of the test suite: 7.26.2’s suite failed on two nodes running it in
parallel (hidraulia, artgins), because the groups of
test_c_treedb_literal_wins exhausted the per-user inotify instances. No
change in the libraries or in the yunos, so a node running 7.26.0 to 7.26.2
needs no upgrade.
test_c_treedb_literal_winsreleases its inotify instances as it goes. A sweep (CR, O4C/O4A/O4X, FP) and the 41 scenarios of the groupearlyran in ONE step of the loop. The watcher of a treedb that closes releases its inotify instance only when the loop completes its cancel, so a group held up to ~1900 instances. Since 7.26.2 runs the groups at once, they reached the per-user limit of 4096: on hidrauliactest -j12failedemailsender/late_server,test_fs_watcher_overflowand the groupearlywith “inotify_init1() FAILED: The user limit on the total number of INOTIFY INSTANCES has been reached”. Each step now runs one scenario or one iteration, and the steps are 10 ms apart, as the test asked:C_TIMER0instead ofC_TIMER, whose steps ran on the yuno’s 1 s tick. The largest group now holds ~290 instances, and the whole suite peaks at ~700 instead of ~3860. The groups at once still take ~34 s. The whole suite in parallel then passed on hidraulia (298/298, 164 s) and artgins (298/298, 534 s).
v7.26.2 (2026-10-05)¶
A second release of the test suite and its build: no change in the
libraries or in the yunos, so a node running 7.26.0 or 7.26.1 needs no
upgrade. On the dev machine (8 cores), yunetas test with no change in the
sources takes 226 s in all, against 470 s at 7.26.1 and ~31 minutes before
7.26.1. A machine with a root build/ from an earlier release stops
building the SDK in it at its next build of that tree (the changed
CMakeLists.txt re-runs its configure).
The root
build/tree builds the tests, not the SDK (ENABLE_SDK, OFF by default). The kernel, modules, utils and yunos are built and installed by each module’s own build dir (yunetas build, and the first step ofyunetas test). The root tree built them too, and installed a second copy of every library intooutputs/libover the first. Each tree replaced the other’s copy, and the new mtime relinked every executable of the other tree on every run, with nothing changed. Now a secondyunetas testwith no change relinks nothing (its root build: 1 s), and a changed library still relinks every test that links it by its full path. (Corrected after 7.26.3: 18 tests named the kernel archives bare and were not relinked; the reinstall of the root tree had hidden it.)-DENABLE_SDK=ONgives the old self-contained tree (the ASan recipe oftest_suite.mduses it).test_c_treedb_literal_winsis seven tests, one per group of scenarios (test_c_treedb_literal_wins/early,/late1,/late2,/dcd,/dcn,/o4,/o4xa; DC is split into its sweep with drafts and its sweep without, O4X into the literal that changes and the one that adds), each in a store of its own (~/tests_yuneta/c_treedb_literal_wins_<group>). In one binary, the whole run took ~215 s and was the floor of a parallel suite. The seven run at the same time in ~34 s (8 cores). The binary takes--group=<name>; with no--group, it runs every scenario as before.A child of
test_c_treedb_literal_winsretries its io_uring ring on ENOMEM/EAGAIN, asyev_loop_create()does (5 attempts, 100 ms doubling). With its groups running at once underctest -j, a child of the DC sweep gotio_uring_queue_init(4096) FAILED: Cannot allocate memorywith memlock unlimited (a transient failure of the kernel to allocate the ring), and the sweep failed. In three rounds of the treedb/timeranger2 tests at-j8afterwards, 4 retries, all recovered.
v7.26.1 (2026-10-04)¶
A release of the test suite and its tools: no change in the libraries or in
the yunos, so a node running 7.26.0 needs no upgrade. yunetas test (CLI
0.21.0) compiles with make -j and no make clean, and runs ctest in
parallel from this SDK on. On the review’s 4-core machine, a run with no
change in the sources went from ~31 minutes to ~7-8.
The tests declare what they share, so
ctest -jruns the suite. The tests that use one directory under~/tests_yunetatake a ctestRESOURCE_LOCK(tr_topic_pkey_integer,tr_msg,tr_delete_instance,perf_topic_integer;inotify_floodfor the two tests that flood the inotify queue).test_topic_pkey_integeris theFIXTURES_SETUPof its six iterators, soctest -R iterator3alone passes now (before, it failed without the data). The benchmarks and that chain areRUN_SERIAL, so the times that the release trend reads stay alone.perf_yev_ping_pong2writes~/tests_yuneta/perf_yev_ping_pong2; before, it wiped the iterators’ database. Before this change,ctest -j4failed five iterators; now the suite passes 292/292. TheyunetasCLI 0.21.0 runs ctest in parallel only on an SDK that carries this change.test_fs_watcher_overflowallows the big tree to cost six times the small one per directory, not two. On a machine with a small L3 cache (artgins, 6 MiB), the 4096 directories of the small pass stay in the cache and the 69632 of the big pass do not. The test failed one run in three there, with correct code. The regression that the test guards against (the index rebuilt per slice) was 14 times.gobj-ui and yunos-js submodules: vite
^8.3.2and maplibre-gl 6.12.0 (gobj-ui: devDependency only, no release; gui_agent 0.29.12, gui_treedb 0.17.78).
v7.26.0 (2026-10-04)¶
What changed after 7.25.22: five changes an operator or a developer has to
know about (each in the upgrade steps below) -- a key delete signalled to the
rt_disk followers with the master’s sequence (a protocol change between
master and followers), with_link_events on by default, the tm markers of
timeranger2 removed (an API removal), write-attr limited to SDF_WR
attributes, and a negative t/tm bound that selects rows again -- the md2 rows
read in blocks and the match condition parsed once per scan, and the defects
of five reviews and of TODO.md, each with a test that fails on the code
before it unless the entry says otherwise.
Performance, against 7.25.22¶
Measured against 7.25.22, each release built from its own tree, the two run
alternately (8 rounds, 24 for the timeranger2 tests); the report is
performance/reports/7.26.0.html. Reading
a timeranger2 history is 16-17% faster (184,338 -> 215,157 records/s;
page by page 166,017 -> 192,697), opening a treedb 20% faster (390 ->
311 us per node), a replica opens 12.6% faster (108.4 -> 94.8 ms, no tm
marker looked for), and a tm query of one minute of a key on a topic 7.25.22
had NOT marked takes 12.8 ms instead of 402 (31x): the md2 rows read in
blocks of 1024 and the match condition parsed once per scan. The prices, each
with its reason: the same query on a topic 7.25.22 HAD marked takes 12.8 ms
instead of 7.5 (+70%: no tm markers, every row read). Found by this A/B
and fixed before the tag: a master that opened a topic wrote its
delete_seq.json durably, which doubled the creation of topics and made the
first open of 40 treedbs 55% longer; written not durably now (it holds 0),
both are back within the noise of 7.25.22 (134 ms for 10 topics, 10.3 s for
the 40 treedbs). Not measured: a key delete on a topic with an rt_disk feed, which by construction
writes the record of its sequence first. Everything else within its noise.
Upgrade steps (operators, read first)¶
Master and followers of a timeranger2 topic upgrade together (key delete signals, below). A follower of 7.25.22 under a 7.26 master takes each
.d<seq>.<key>signal for a key: itskey_deletedcallback gets a key that does not exist, and the real delete is not heard. A 7.26 follower under a 7.25.22 master hears no delete at all, and says so when it opens a feed (“The master of the topic does not signal key deletes as this follower hears them”: there is no<topic>/delete_seq.json, which a 7.26 master makes when it opens the topic). Stop the followers (rt_disk readers: replica yunos,tr2list --follow, C_TRANGER lists of a non-master) before the master, and start them after it.with_link_eventsis on by default (C_NODE,C_TREEDB,C_AUTHZ): a link publishesEV_TREEDB_NODE_LINKED/UNLINKED, not the parent’sEV_TREEDB_NODE_UPDATED. A yuno served by a v1 SPA sets it0in itsC_TREEDBand in itsC_AUTHZ’s kw (estadodelaire and hidraulia do, indb_history); a gobj that subscribes to every event of a treedb service declares the two events or subscribes only what it handles; a GUI that shows a parent’s hooks needs gobj-ui 7.25.26 or later.write-attrwrites onlySDF_WRattributes. A script that set a persistent attribute withoutSDF_WR(max_sessions_per_user,allowed_ips, emailsender’surl, ...) uses the config or the attribute’s own command.tranger2_mark_tm_order()and C_TRANGER’smark-tm-orderare removed. Nothing to run on upgrade. A rollback to 7.25.5..7.25.22 runs that release’smark-tm-order all=1on every topic it marks, or a tm query can miss rows.A negative
from_t/to_t/from_tm/to_tmselects rows again (it matched none since v7): adb_historystarted withfrom_t=-86400now reads the last day at its start -- expect that load.fix:
jtreewithoutrename_hookgave every child twice. The hook kept the refs of the collapsed view ({id, topic_name}) and the children were appended after them, whole:[dev, ops, dev{...}, ops{...}]. The hook now starts empty and holds each child once, whole, asrename_hookalready did. Nothing in the SDK, the projects or the SPAs called it withoutrename_hook. In 7.25.22 (from the port before v7).test_c_node_commands(red on the previous code:dev ops dev ops).tests, emailsender: what the sender rules and skip-email do, and a window that measured the test. New
skip_paused(skip-email works with the service paused),default_from_case(afromthat is the default in other letters IS the default: a not-found refusal is the account’s) andforeign_access_denied(a 5.1.8 that is no not-found is the account’s, whatever the from).url_change_resets_pacingwas widened to 1.8 s without finding why the session connected 1 s late: it did not -- it connects in 1-2 ms; the fake server told the test on its C_TIMER, which is accurate to the second. Withnotify_delay0 it tells at the connection itself, and the window is 800 ms -- under the 1 s oftimeout_retry, with room for a loaded machine (red withoutreset_pacing(): 7 s).tests:
test_fs_watcher_overflowmeasures the watcher against itself. It failed on artgins (a Xeon D-1521) with the code right: 224 us per directory against a fixed 200. Most of it was the test’s own: its loop counted the names told by walking all 69632 of them on every turn, and billed that to the watcher (73 us per directory here; 20 without it). The watcher’s cost on a tree of 69632 directories is now held against the same pass over 4096 on the same machine: at most twice, plus 20 us (14 and 20 us here; with the index of 7.25.10 rebuilt per slice, 31 and 437: red).tests: red tests for fixes that had none.
deactivate-snapanswering-1when its save fails (test_tr_treedb_failed_save, the writes of the snap’s key failed by--wrap=write);save_json_to_file()answering-1on a failedclose()(test_helpers,--wrap=close); a kw that carries a gbuffer through C_NODE’streedbs/links/hooks/node(test_c_node_commands) and C_MQIOGATE’sview-channels(test_c_qiogate_stats);kw_update_missing()with a gbuffer (test_command_binary_kw); and the templates ofyuno-skeleton, built and run (yuno_skeleton/templates: a template timer that is a plain child never reachesac_timeout,MSGSET_INTERNAL_ERRORdoes not compile; it tests the tree it is in, never/yuneta/development/yunetasor the installedyuno-skeleton). Each one red on the code before its fix (checked by putting that code back), green now.fix: C_NODE
parents/childrenwith nooptionslogged an ERROR per option. The options are optional, and were read off a NULL: six ERRORs with a stack (“kw must be list or dict”) for eachparents, one for eachchildren. They are an empty dict now, asnodesalready did. Newtest_c_node_commands: what the read commands of C_NODE answer (treedb-info,node,instances,pkey2s,jtree,parents,children,hooks,links,print-tranger) and what the snaps do to the data (red on the previous code: 18 unexpected ERRORs).BREAKING (protocol between master and followers): a key delete is signalled with the master’s delete sequence.
tranger2_delete_key()removes the feed’sdisks/<rt_id>/<key>/as before, and then makes and removesdisks/<rt_id>/.d<seq>.<key>in the directory of every feed (.h<seq>.<sha256>for a key longer than NAME_MAX - 23 bytes);seqis the topic’s delete sequence, recorded in<topic>/delete_seq.json(durable, before the signal) and growing across restarts of the master. A follower keeps, per key, the highest sequence it applied: a signal above it is a new delete (the key is forgotten), one at or below it a delete another feed heard first (the feed is told, nothing forgotten). It replaces the matching of a delete by its place in each feed’s queue (debts, doubts, the place each feed was watched from), which left two holes: a SECOND delete of a key heard first by a feed opened while the first one was in flight was taken for the first, and a feed opened AFTER another one heard a delete took its signal for a new one -- the key forgotten again (a key written again meanwhile went with it) and, in both, the delete told twice to a feed at its next overflow ([DEL DEL]). Both existed in 7.25.22. The master takes a delete’s sequence BEFORE the key leaveskeys/, and a follower at an overflow listskeys/and reads the sequence AFTER, so every key it finds missing is bounded at or above its delete: the keys told then are noted, and a signal of theirs queued behind the overflow is not told again; a delete not told there (of a key on disk again, or that no feed had heard) is told by its signal. What every feed heard is pruned, so the follower keeps no more than the deletes some feed may still hear. The master makesdelete_seq.json(0) when it opens a topic that has none -- not durably: it holds 0, and one lost to a power cut (or left empty by it) is made again at the next open, while every number above 0 is written durably; a record it cannot read, or cannot WRITE, refuses the key delete (logged) before anything is removed and is left as it is (written late, a follower that read it in between would bound the key one below its delete), and a sequence taken for a delete that could not remove its key stays taken -- a gap, which the followers do not mind (given back, a follower that read it at an overflow would take the next delete for one it was told). A follower that hears a signal below its last one logs that the sequence went back; one that opens a feed on a topic with no record logs an ERROR, and an info when the feed hears its first signal after all. A signal the master made and could not remove (it died in between, or thermdir()failed) is removed at its next open of the topic: the feed hears it then if it was the last delete; an older one, followed by deletes the feed heard, is moved out of the feed’s directory and removed unheard, said (told late, it could take a key written again out of the cache), and so is every leftover while the record cannot be read. Upgrade master and followers of a topic together (see the upgrade steps). The delete of a topic with rt_disk feeds costs two disk flushes more (the record of the sequence); a topic with none, nothing. The hashed form is chosen by the key’s LENGTH and kept apart from every key (no key holds a/), so a key starting with#is a key, and#<sha256 of a long key>never stands for that long key; the follower keeps the key a hash named while the delete is kept, so every feed is told it, not only the first one.test_delete_key_propagation:opened_after_heardandsecond_delete_in_doubt(red on the previous code:[DEL DEL]after the overflow),odd_keys(a#key, a key too long for the name, and#<its hash>, heard by two feeds: each delete forgets its own key, and each feed is told the three),delete_seq_record(the record made at the open, read from inside thermdir()of the key to show it was written BEFORE the key leftkeys/, kept taken by a delete whosermrdir()fails (never given back), a record that cannot be written or read refusing the delete, a left signal heard at the master’s next open and an older one removed unheard, a follower with no record saying so); the white-box checks of the old debts became “nothing of the deletes kept once every feed heard them” (intest_rt_disk_overflow: nothing kept below what every feed heard), andstale_debt_reborn(a debt put there by hand) went with the debts.BREAKING (default):
with_link_eventsis ON by default -- inC_NODE, inC_TREEDB(copied to every treedb it opens) and inC_AUTHZ, which gets the attribute (new) and copies it totreedb_authzs. A link or an unlink now publishesEV_TREEDB_NODE_LINKED/UNLINKED(the relationship:hook_name,parent_topic_name,parent_id,child_topic_name,child_id) and NOT the parent’sEV_TREEDB_NODE_UPDATED; the child’s own update, from its save, is published as before. Why: the parent’s update is the parent collapsed whole -- every hook list, every child id -- on every link, whether anybody listens or not, so filing a child under a parent with thousands of children cost O(children), and a fleet of N new devices O(N²) at the moment it comes on line (measured in a stress test: ~70 ms of cpu per link, 3000 -> ~13 new devices/s). What to do:A yuno served by a v1 SPA (it reads the parent’s update) sets it off:
"with_link_events", 0where it creates itsC_TREEDB, and'kw': {'with_link_events': false}in itsC_AUTHZwhen the SPA editstreedb_authzs(estadodelaire and hidraulia do, indb_history).A gobj that subscribes to EVERY event of a treedb service, or the parent of a
C_NODEhosted as a pure child (subscribed to everything), now receives the two events, and an FSM that does not declare them answers “Event NOT DEFINED in state”. Subscribe the events handled, or declare the two.C_MQTT_BROKERnow subscribes onlyEV_TREEDB_NODE_CREATED/UPDATED/DELETEDof its treedb (wattyzer’sdb_history_wzthe same).A GUI that shows the parent’s hooks needs gobj-ui 7.25.26 (
C_YUI_TREEDB_TOPICSre-reads the parent on a link event; the graph already followed them): gui_agent 0.29.10, gui_treedb 0.17.76. The next gobj-ui re-reads a parent only when its row is loaded, one read at a time per parent: a burst of links to it costs two reads per viewer, not one per link; and a re-read that fails (a parent deleted meanwhile) is a warning, never the app’s error modal.set-link-eventsstill switches an open treedb at run time, and the configured value comes back at the next start.test_c_node_link_eventscreates itsC_NODEwithout the attribute.
timeranger2: a scan parses its match condition once, not per row. The matcher read
match_condwith ~15json_object_get()per row, the scan 3 more from the segment and 1 for its direction, and it wrote its position into the iterator (two json integers) per row. A scan now parses the condition into a C struct when it starts -- afterget_segments(), which resolves a negative t/tm bound and writes it back, and never kept from one scan to the next -- reads a segment’s rows, file and t order once per segment, and writes its position once, when it ends (nothing reads it during the scan). A tm query of one minute on 1 key x 30 files x 20 000 rows (perf_timeranger2, 8 rounds alternated): 0.152 +- 0.003 s -> 0.0145 +- 0.001 s. The open phase is unchanged (it does not go through the matcher).timeranger2: the scans read md2 rows in blocks. An iterator’s load, its index, its pages and a follower’s new rows made an
lseek()and aread()per 32-byte md2 row; they read 1024 rows with onepread()(a block never reaches past the rows the scan’s segment counts, so a file the master appends to is not read half way, and it is read again after any write of md2 rows by the process, so a user_flag rewritten from a callback is seen). A tm query of one minute on 1 key x 30 files x 20 000 rows (perf_timeranger2, 8 rounds alternated): 0.439 +- 0.071 s -> 0.168 +- 0.047 s. The rest is the per-row json of the matcher. One difference: a follower’s scan sees a user_flag the master rewrote in place DURING that scan up to a block later than it did (no ordering was promised between the two processes either way).test_tm_order: ranges across the block boundaries, both ways, on the three read roads, and a key of 2600 rows read whole.Schema editor (gobj-ui 7.25.25, gui_agent 0.29.9): a column with data behind it is warned about before it is deleted. The records keep the values of a deleted column and no reader shows them any more; the editor reads the topic (paged, capped) up to the first record holding a value and says so, or that it could not read them. Not refused: a decision.
BREAKING (API removal): timeranger2’s tm markers are gone. From 7.25.5 the master marked an md2 file whose
__tm__went back (<file>.tm_unordered, and"marks_tm_unordered": truein a new topic’stopic_desc.json), and a tm query trusted the tm range of the files not marked: it skipped files out of the range and ended a file’s scan at its first row past it.tmis the producer’s time and the store keeps the records as they came, so ordering by it is not the store’s job: afrom_tm/to_tmcondition is now a FILTER on every row of the key, leaving no file out and ending no scan -- correct always, rollback-proof. Removed:tranger2_mark_tm_order()(public API) and themark-tm-ordercommand ofC_TRANGER; nothing writes or reads.tm_unorderedormarks_tm_unordered, and a store that has them keeps them, ignored (no migration). Cost: a tm query reads every md2 row of the key -- one minute on 1 key x 30 files x 20 000 rows takes ~12.8 ms with this release’s block reads and once-parsed condition (below), where a topic marked by 7.25.5..7.25.22 took 7.5 ms and an unmarked one ~0.40 s; who needs fast access by tm keeps it themselves (a topic keyed by it, an index in memory). Thefr_tm/to_tmof a file inlist-keysare its first and last rows unless the file was read whole: approximate when its tm goes back. Thetmarker (<file>.unordered, a late__t__) stays. Upgrade: nothing to run. Rollback to 7.25.5..7.25.22: run that release’smark-tm-order all=1, or a topic that says it marks can miss rows in a tm query.A yuno keeps the case of its directory.
register_yuneta_environment()lowercased theroot_diranddomain_dirit was given, silently, while the agent builds a yuno’s directory as its realm and role are written and keeps itsbin/there: a capital in a realm’s owner, name, role or env, or in a yuno’s role or id, gave the yuno a sibling directory for its logs, data andtemp/that the agent’s commands never looked in. They are kept as given. No node had a capital under/yuneta/realms(checked 2026-10-03), so no directory was split.fix:
C_NODEleaked the binary field of a record kw it shared with others. The treedb takes thegbufferof the record it writes (treedb_store_files()removes the key and decrefs it once, because the treedb releases a record withjson_decref()), andC_NODEhanded it the caller’s kw itself. A kw shared bykw_incref()-- an event’s, held by every layer that published it -- lost the key under the other holders, whosekw_decref()then released nothing: one gbuffer per holder. The MQTT broker writes its session from theEV_ON_OPENof the CONNECT, so every CONNECT with a will leaked its payload (~360 bytes), and so did every retained PUBLISH with a payload (retain__store()). A kw that carries a binary field and that others hold now reaches the treedb as a twin of its own (kw_twin()); and thecreate-node/update-nodecommands andEV_TREEDB_UPDATE_NODEput the bytes of thefilecolumns into a record of their own, not into the sender’srecord(which kept a gbuffer the treedb had released). In 7.25.22 (since thefilecolumns).c_mqtt/will_acl(its memory check: two wills and a retained payload not freed before),test_c_assets(the sender’s record keeps no gbuffer).MQTT broker: the will message obeys the ACL.
will__send()published it with no check, so withenable_aclon a client could publish, as its will, to a topic its groups refused. It is asked like any publish when it is sent; a refused will is not published, with a warning (c_mqtt/will_acl: raw MQTT clients, a will refused and one allowed; red without the check: the refused will reached the subscriber). And theenable_acldescription andmqtt_broker.mdsay what the ACL is not: it is keyed by theclient_ida client chooses, and no client is bound to the user that authenticated it, so it separates topics, not users.CLI (yunetas 0.20.4 on PyPI):
yunetas buildwarns when two builds install a yuno of the same role inoutputs/yunos, where the last one overwrites the other (yunovatios’gate_caudalreached hidraulia that way): a red WARNING per role, read from each tree’sinstall_manifest.txt.Tests:
test_secret_attrs’ no-fallocate case no longer passes untested when the seccomp filter cannot be installed (the child exits 2);test_rt_disk_overflow’s reborn case no longer hangs to the ctest timeout when its room was not made (asize_tsubtraction wrapped).timeranger2: a negative
from_t/to_t/from_tm/to_tmis relative to the key’s last record again. Since v7 it reached the matcher raw, compared with an unsigned__t__(-86400 became ~1.8e19), and no record matched, silently: adb_historythat starts withfrom_t=-86400(hidraulia, estadodelaire, wattyzer) never read what had arrived while it was stopped.get_segments()resolves it against the key’s last record -- itst, and itstm(the last row of the file that holds the key’s highestt):fromafterlast - N, that bound excluded;toup tolast - N. (Before v7 it was the TOPIC’s last record, of any key; a v7 topic is read per key.) It is written back into the condition of that key alone: a list of several keys resolves it per key (each iterator gets its own copy of the condition, intranger2_open_list()and C_TRANGER’sopen-list); the realtime half of such a list keeps the bound unresolved -- there is no one last record for all its keys -- so a negativefrombounds nothing there and a negativetotakes no new record. The matcher compares signed, and on a key with no record yet a negativefrombounds nothing and a negativetotakes no row.test_tm_order(the three read roads, both ways, master in memory, reloaded, replica; atmagainst the last record’s, not the highest one, which only the master in memory knew; a list of two keys with different last records) andtest_c_tranger(open-listover two keys).Builds on a glibc older than 2.34. The daemon start closed its inherited files with
close_range(), whose glibc wrapper is 2.34’s: a source build on an older glibc did not compile. It calls the system call directly (syscall(SYS_close_range)), with the loop as before where the kernel lacks it. Kernel headers older than 5.9 do not define its number: it is given (436, the same on every architecture built), or such a build would always take the loop.fs_watcher: a directory watched again brings its subtree. Under
FS_FLAG_RECURSIVE_PATHS, a directory that could not be watched (out of inotify watches) was watched alone when watches came back: a subdirectory made in it meanwhile was never watched, and nothing made in it was heard. Its subdirectories not watched are queued behind it now, and each is watched and handed as created, parent first -- within the batch’s bound: every try of a batch counts against its 64, anENOSPC/ENOMEMends them, and a directory is looked up by its path, not in the whole table of watches (a tree of 100000 keys, at the moment watches ran out, would have blocked the loop). “Directories watched again” is said at the end of the batch where none is left, not while a subtree is still queued. A directory renamed during an overflow and watched again at its new path no longer leaves its old path in the index, where a directory made later there was taken for one watched. Areaddir()that fails half way (EIO,ESTALE) is said, in the re-watch and in the pass after an overflow, not taken for the end. And the first walk of a recursive watch takes hidden directories too, asIN_CREATEand the pass always did: a.namethere before the watch was never heard. ForC_FSand the watch utilities withrecursive: a tree that already holds a.gitor a cache now takes inotify watches, and events, for all of it at the start. The pass after an overflow no longer follows a symlink to a directory (the first walk and the re-watch never did), and a failedfstatat()there is said. The half-of-the-open-files warning has a hysteresis (said again only after the count fell under 40%, not on every swing around the half), and an unparsablemax_queued_eventsis said, as an unreadable one was.C_TCP: nothing of the gobj is touched afterEV_DISCONNECTEDif a host destroyed it there.set_disconnected()published the event and then reset the gobj’s volatile attrs: a host that destroyed a transport on that event would have been a use after free (none in the tree does). The publish is guarded by the liveness marker of 7.25.22.ytls: a subscriber’s errors are not summed into TLS codes; OpenSSL’s write stall is bounded.
flush_clear_data()summed the answers ofon_clear_data_cbinto the number space of -2222 (the session freed) and of the “< -1000: a TLS error” band: more than 1000 records answered -1 in one read became a TLS error and closed the connection, and a sum of exactly -2222 left it hanging. It answers -1 now, however many. OpenSSL’sencrypt_data()looped onWANT_READ/WANT_WRITEwith no bound; it stops after 5 tries in a row with a warning, as mbedTLS does (a write that progresses starts the count again). OpenSSL’sflush_clear_data()checksgbuffer_create(). AndC_TCPreads the -1 the same way on the flush that follows the handshake: a subscriber that answered -1 to the first data, arrived with the handshake, closed the connection as for a TLS error; only a TLS error (below -1000) does now, as on the decrypt path.gbuffer_printf()writes a text that fits exactly.vsnprintf()was given the free bytes only, not the byte every gbuffer keeps for the NUL: a text of exactly the free bytes was refused (“MAXIMUM SPACE REACHED”) when the gbuffer could not grow, or, grown to exactly the text, written one character short with “NOT ENOUGH SPACE” and the NUL counted as data.C_AUTHZwith no users treedb refusesadd-jwk/remove-jwk. Such a yuno (local access only) validates no JWT, and asked nobody whether the caller could change its keys: any valid JWT could add one. The key was never used, and leaked (4.4 KB per key, seen by the test). They are refused now, with “no users treedb in this yuno: local access only, no JWT is validated here”;list-jwkanswers, with the keys of the config (jwks), which such a yuno never uses.A persistent attr the gclass no longer persists is said at load. The load wrote only the
SDF_PERSISTattrs of the file and dropped the rest in silence, while every save kept them in the file: thejwksthatadd-jwkpersisted before 7.25.22 (now the config’s) vanished at the first restart, and a node whose keys came only from there lost its JWT logins with nothing in its log. The load now warns once per attr (“Persistent attrs file holds an attr that is not persistent: NOT loaded”);remove-persistent-attrsdrops it from the file.Agents’ units:
ExecStopPostworks on a v1 or hybrid cgroup host, and says what it kills. It read/sys/fs/cgroup/system.slice/%n/cgroup.procs, a v2 path written by hand: a no-op, unsaid, elsewhere, and its kills left nothing in the journal. It asks systemd for the unit’sControlGroup, reads it under the v2, hybrid or v1 mount,loggers each pid it kills (and each kill that fails), and says when no cgroup was found.Agent: a yuno still alive 10 s after a node bounce gets no second instance.
restart_nodes()waits for the yunos it killed, and after 10 s relaunched them anyway, without asking whether they were alive: a yuno stuck in a disk wait got a second instance. Those are now spared, listed by “yunos killed for the restart still alive after 10 s: each one launched when it is gone”, and each one is launched when it is gone: looked at every second, for 5 minutes, after which the warning says torun-yunothem. Gone is every pid the restart killed gone AND no process left running the yuno (its role and configuration). A pid is gone when it does not exist, is a zombie, or runs with another start time than the one recorded when it was killed -- a pid reused by another process; a task in D state is alive, whether or not it still has its command line (it loses it while it still holds its files), and so is one whose/proc/<pid>/statcannot be read (logged). A pid gone before it was recorded is not waited for. Only those are launched. A spared yuno the operator stops (kill-yuno), disables or launches meanwhile leaves the window; one stopped, disabled or launched in the first 10 s is held apart, not run by the relaunch nor spared (enable-yuno, in those 10 s only, takes back a disable, not a stop or a launch of the same yuno).Agent:
kill-yunoof a yuno found only by the scan says it is not waited for. Such a yuno is signalled and the answer comes at once; arun-yunosent before it is gone finds it alive and does not launch it. The answer now says so.Restarts without root put the agent back in its unit.
/yuneta/bin/restart-yuneta(the certbot hook’s fallback) and logcenter’s defaultrestart_yuneta_command(on a queue alarm) restarted the agent, as useryuneta, withyshutdownandyuneta_agent --start: with the agents’ units (7.25.22) that agent ran OUTSIDE its unit, where systemd does not see it and the next boot does not start it. Now, where the units are,restart-yunetastops the yunos (yshutdown --no-kill-agent, plus its arguments) and restarts the unit withsudo -n systemctl restart yuneta_agent.service; it refuses, and says so, when sudo does not allow it. A node without the units keeps the old way. logcenter’s default runsrestart-yuneta -swhere it exists (logcenter is not stopped), and logs an ERROR when the command fails or is refused (exit code, or the signal that killed it), or when none is configured: only asystem()that could not run was said. The text names no cause it cannot know. It waits for the command as before -- under the units, the agent’s restart. agent22 was never touched byyshutdown, and is not now./etc/init.d/yuneta_agentanswers what the units answered. Under systemd,start,stopandrestartexited 0 whatever the units did (only logged), soservice yuneta_agent startreported success with the agent’s unit failed;statusasked only the units, so an agent running outside its unit read “not running”. Now each exits with the units’ answer, andstatussays “running OUTSIDE its unit (start puts it back)”.C_TIMER0: a stopped timer no longer ends the yuno’s loop. Its callback answered -1 when its gobj was not running, andyev_loopends the loop on a -1: the cancel of a child timer stopped while the yuno runs -- the control center’smt_stopclears and stops itsrates_timer-- stopped the whole yuno whenever the service was stopped alone. It answers 0 now; the loop ends withset_yuno_must_die(), as it always did. And agobj_stop()followed by an arm in one turn now saysEV_STOPPEDwhen the cancel comes back (it said nothing, and left the arm noted on the stopped gobj).--stopchecks its signals, and a stopped daemon is not relaunched.kill()was never checked: onEPERM(an agent of another user)--stopwaited 10 s, printed “killed (SIGKILL)” and exited 0. And the watcher ignored SIGQUIT, so a child that crashed during its orderly stop was relaunched 2 s later; at 10 s the watcher was SIGKILLed and the new child left alive, an orphan. Now the watcher notes SIGQUIT (SA_RESTART) and does not relaunch the child that ends after it, whatever its end;daemon_shutdown()signals the watchers first, then their children, checks everykill()(a refusal is said, not waited for, and--stopexits 1), and before the SIGKILL scans the name again, so a child started meanwhile is killed too. And it takes only the processes of the name that were STARTED as it (the base name ofargv[0], from/proc/<pid>/cmdline): the SysV script/etc/init.d/yuneta_agent, root’s and of the same name, was signalled by the--stopit ran -- a script’sargv[0]is its interpreter. An agent of another user is taken, so itsEPERMmakes--stopexit 1, and so is one whose binary was renamed (*.bak-pre-<version>) or replaced on disk. (A first form of this compared/proc/<pid>/exe, which another user’s process does not let read, and which follows a rename: both left the agent up and exited 0.) A name and anargv[0]are whatever a process’s starter chose, so the list of processes has no fixed size: 64 look-alikes, made first, pushed the real agent out of a list of 64, and it was left up -- run as root too. A zombie of the name is dead and makes noEPERM; acmdlinethat cannot be read for another reason than the process’s end is said, and--stopexits 1 (it is not known whether it was the daemon).daemon_shutdown()returnsint(wasvoid). The agents’ units send SIGQUIT to the watcher ($MAINPID) before the agent, so a stop under systemd does not relaunch it either.C_WEBSOCKET/C_PROT_TCP4H: a frame of exactly the default max no longer hangs. Withmax_payload_size/max_pkt_sizeat 0 the max was the yuno’s max block, but a gbuffer holds one byte less than its block (its NUL): a frame of exactly that length was taken, its buffer could not grow to hold it,istream_consume()dropped the bytes in silence, and the frame waited for its timeout. The max is now what a gbuffer holds (the max block less one byte), and a configured value above it is capped to it; such a frame is refused like any frame too big.istream_consume()checks whatgbuffer_append()answers: a buffer that cannot hold the bytes asked for is an ERROR, and the frame is not completed cut.Deprecated
C_PROT_MQTT: a packet longer than a gbuffer holds is refused. Its payload buffer is capped at the max block, and the remaining length was not: such a PUBLISH, which a peer can send, was delivered cut, and with the check above it never completed and logged an ERROR with a stack per chunk. It is refused at its header now (“Mqtt packet too large”, a warning) and the connection closed, as tcp4h and the websocket do.Tests: the memory check of
c_prot_tcp4h/test1measuredget_cur_system_memory(), which is 0 withoutCONFIG_DEBUG_TRACK_MEMORY-- empty on the nodes. It measuresmallinfo2()now.C_UDP_S/C_UDP: a subscriber’s error does not stop the reading. The read was re-armed only when the publish ofEV_RX_DATAanswered 0, taken as a sign that the gobj lived: one subscriber answering -1 (or an “Event NOT DEFINED”, which the publish sums) left the server deaf, with nothing said -- the defect fixed inC_TCPin 7.25.22. Now, as there, a marker on the stack says whether the gobj lives after the publish, and the read is re-armed whenever it does and the event is idle.rpm: the agents’ SELinux label survives a relabel.
%postlabels the web server’s wrapper and the two agentsbin_twithsemanage fcontext, and fell back tochconwhensemanagewas missing -- which the spec did not require. Achconlabel is lost on an autorelabel or arestorecon -R /yuneta, and then both agent units fail at boot with203/EXEC. The package now requirespolicycoreutils-python-utilswhereselinux-policyis installed (a rich dependency, for%postand%postun), says it when it still has to fall back tochcon, and takes the three rules out of the policy on erase.BREAKING:
write-attrwrites onlySDF_WRattributes.gobj_is_writable_attr()answered TRUE forSDF_WRORSDF_PERSIST, so every persistent attribute was writable at run time bywrite-attr, which only the per-command gate guards (off by default) -- around the checks of the attribute’s own command.C_AUTHZ’smax_sessions_per_usercould be set without theupdatepermissionset-max-sessionsasks for;C_IDP_KEYCLOAK’skc_base_url,C_YUNO’sallowed_ips/denied_ips(not normalised either), emailsender’surl/password, and the rest, were open the same way. NowSDF_PERSISTwithoutSDF_WRis set by the config or by its own command;write-attranswers “attr not writable”. Upgrade: a script that set such an attribute withwrite-attruses the config or the attribute’s command instead; a gclass that wants one writable at run time declares itSDF_WR|SDF_PERSIST.dbsimple: a yuno run as root takes its user from the chain closed from
/down. For a yuno run as root, the persistent attrs file of the yuno’s user is trusted, and that user was the owner of the first closed directory found going UP from the file. A member of theyunetagroup could rename the data directory away (its parent is 02775), make one of their own 0755 -- or a symlink to one -- with their file in it: the walk stopped there, the file was loaded, and the next save wasfchowned to them with the yuno’s secrets. Now the walk goes DOWN from/withopenat(O_NOFOLLOW), and the yuno’s user is the owner of the lowest directory of the chain closed to others, each one owned by root or by that user (/yuneta/realmson a node); it ends at the first directory others can write, or of a third user. A symlink met inside that chain (/yuneta -> /srv/yuneta) is followed -- only root or the chain’s user can have put it there -- and its target is walked from/under the same rules, the user named so far still holding; a..is taken as the parent of the directories walked, without what the directories it leaves had named (/a/X/../Wdoes not trust the owner ofX). (A first form of this stopped at any symlink, and a root yuno under a linked/yunetarefused its own files.)
v7.25.22-3 (2026-10-02)¶
A packaging revision, not a new version: the code is 7.25.22’s, and the
packages are attached to the existing tag 7.25.22.
/etc/init.d/yuneta_agentdrives the two agents’ units. Under systemd the script was redirected toyuneta_agent.servicealone on Debian (by/lib/lsb/init-functions), so agent22 was out of its reach; on RHEL it was not redirected at all, and astartlaunched both agents by hand, outside their units. Now it skips that redirect and drives the units itself, the way the SysV script always treated the pair:startstarts the main agent’s unit and then, once it is up, agent22’s (a broken binary never takes both down);stopstops the main agent alone, agent22 keeps running;restartrestarts the main agent and starts agent22 only if it was not running;statusreports both. An agent running outside its unit (a hand--start) is stopped first and started in it -- found by its executable (/proc/<pid>/exe), because the script is namedyuneta_agenttoo andpgrep -x yuneta_agentfinds the script itself.service yuneta_agent22 ...andsystemctlstill drive one unit alone. Root is needed, assu - yunetaalways required.
v7.25.22-2 (2026-10-02)¶
A packaging revision, not a new version: the code is 7.25.22’s, and the
packages are attached to the existing tag 7.25.22.
rpm: the agents’ units start under SELinux. 7.25.22-1 runs the two agents in native systemd units, and systemd (
init_t) cannot exec adefault_tfile -- everything under/yunetais, outside the policy. On yunovatios-central (Rocky 9, enforcing) the agent’s unit died with203/EXEC, “Failed to locate executable /yuneta/agent/yuneta_agent: Permission denied” (avc: denied { execute } ... scontext=init_t tcontext=default_t), and%post, as designed, left agent22 running outside its unit rather than stopping both.%postnow labels/yuneta/agent/yuneta_agentandyuneta_agent22bin_t(semanage fcontext+restorecon,chconwhere the policy tools are missing), as it already did for the web server’s wrapper. A binary moved in by hand on such a node takes the rule back withrestorecon. The.debis unchanged (Debian runs no SELinux by default).
v7.25.22 (2026-10-02)¶
What changed after 7.25.21: the defects, the nits and the risks of two reviews -- of 7.25.21 itself, and of the fixes made for it -- and the open defects of TODO.md section 1. Each fix has a test that fails on the code before it, except the hook and the entries this list marks “(no red test)”. Two items stay open, in TODO.md: the accounting of key deletes in rt_disk followers (its design is decided: a sequence in the master’s signal, which changes the protocol between processes, so it gets a release of its own), and code of the projects, moved to their own TODO files.
Performance, against 7.25.21¶
Measured against 7.25.21, each release built from its own tree, the two run
alternately (8 rounds, 24 for the timeranger2 tests); the report is
performance/reports/7.25.22.html. A forced
treedb delete takes 39% less time (69.8 -> 42.5 us): a delete no longer
walks every open key. Nothing is slower on a path this release changed. Four
figures moved outside their spread on code that did not change, and are not
claimed: appends -1.9% per second (-1.6% with a live reader) and a treedb
update in memory +5.0% of time, slower; the appends to a store of many keys
-8.5% of time, faster -- placement of the code in a whole-release build.
Upgrade steps (operators, read first)¶
The command attrs run with
system()are set in the config only: the agent’scert_sync_copy_cmdandcert_sync_store_dir, logcenter’srestart_yuneta_command. A value set at run time withwrite-attrand persisted is no longer read: put it in the yuno’s config ("global": {"agent.cert_sync_copy_cmd": "..."}).C_AUTHZ asks the permissions of its users treedb (
treedb_authzs) for every command from a peer, with the per-command gate off too:readto list,create/update/deleteto manage users,updatefor the JWKs,check-user-pwdandset-max-sessions. A gui_agent operator with no such permission now gets-403in the Users workspace.command-yunoforwards the operator’s__username__, so the operator is asked in the target yuno’s OWN treedb: an operator who manages users in the agent may get-403in another yuno. A yuno with no users treedb (local access only: logcenter, auth_bff, sgateway, watchfs, dba_postgres) answers the user commands-1, “no users treedb in this yuno: users, roles and permissions are not managed here”, and the jwk commands as before.A persistent attrs file of another user is refused (not loaded, saves refused) unless it is root’s, or the yuno runs as root and it is its user’s (the owner of the nearest directory above it that nobody else can write:
/yuneta/realmson a node) -- and even then when the group or others can write it. A root file left 0664 by a release before 7.25.19 is refused: make it 0600. A node whose yunos changed of user (a dev box:yunetaafter a reboot, the developer’s own user by hand) gives the files to the yuno’s user.C_AUTHZ’s
jwksanddefault_roleare set in the config only. A value persisted by an older release (write-attr,add-jwk) is no longer read: put the keys inAuthz.jwksand the role inAuthz.default_role.add-jwk/remove-jwkstill change the running set, until the yuno restarts. yunovatios setsdefault_rolewithwrite-attrin its deployment notes: move it to the batch config.A frame before the session is limited (
max_pre_session_frameofC_IEVENT_SRV, 64 KB by default): a client that sends a bigger identity card needs it raised.register-idp-userwith aroleanswers-403to a caller without theupdatepermission oftreedb_authzs(ascreate-userasks to link one). An operator who registers people with a role needs that permission; without a role it is answered as before.kill-yunoof a yuno alive but not connected answers at once: the yuno is signalled and not waited for. Arun-yunosent right after it can find it still exiting and not launch it (“not launched again”): wait for it to be gone (list-yunos, or its pid) before running it.Rebuild every project against the new headers.
fs_event_t(fs_watcher.h) gains ten fields, at its end.An rt_disk follower holds one more descriptor per key directory of each feed (
FS_FLAG_DIR_FDS). A yuno now raises its soft open-files limit to its hard one at its start (C_YUNOlimit_open_files,0by default; up to 7.25.210left it as it came, and a yuno started from a desktop session kept the soft 1024); the packages give the yuneta usernofile unlimited. A follower that is not a yuno raises its own.A stop of a TCP connection no longer waits for ever for a peer that does not take its data: after
timeout_stop_tx(10 s by default, a newC_TCPattr) the connection is aborted. A host that stops a connection expecting a slow peer to drain more than 10 s of data sets it higher, or 0.Two
C_TCP_Sof the new method on one pool of channels are refused: the second leaves the first’s channels to it with an ERROR. That pool was never shared (the second took every channel, silently); a shared pool is forchild_tree_filter.The agent’s audit log redacts more names: one that holds a part of a secret’s name is redacted whatever else it holds (
token_endpointtoo).The agents run as systemd units:
yuneta_agent.serviceandyuneta_agent22.service. Start, stop and restart them withsystemctl; an agent started by hand with--startruns outside its unit, where systemd does not see it and the next boot does not start it. To deploy an agent binary: move it into place, thensudo systemctl restart yuneta_agent(one agent at a time, as always). The package’s upgrade moves both agents into their units (agent22 too, which the SysV script started outside any unit). A restart of an agent’s unit leaves the yunos running (KillMode=process; its journal says “Found left-over process” for them).A direct owner of an fs_watcher handles
FS_WATCHER_GONE_TYPE: the watcher tells it before it goes, already freed when the callback returns, so the owner drops its pointer and never stops it again. timeranger2,C_FSandutils/c/fs_watcherdo; no project owns one directly today.gobj_stats()of aC_IOGATEor aC_QIOGATEanswers the envelope ({"result", "comment", "schema", "data"}) like every othermt_stats; the counters are underdata. A project that read them bare reads 0 now: yunovatios’c_sim_controller.cwas migrated (it reads the queue’smsgs_in_queue/pending_acksattrs). AC_CHANNELread by its parent still answers the bare data.C_YUNO’s
uptimeis in seconds since the yuno started. It was the MACHINE’s uptime in jiffies (/proc/uptime× HZ). A chart or an alarm on it changes of unit; the machine’s uptime is now asked with the new commandinfo-uptime.
Fixes¶
emailsender:
set-email-usersends what was queued while the credentials were missing. The session connects only for a message it holds, and the command only started it: an email queued before the credentials waited for the next one queued, or for a pause and a play. Testemailsender/set_user_queued.emailsender:
skip-emailno longer repeats the “username or password is empty” ERROR. It started the SMTP side after every skip, whatever it was before, so a service playing without credentials said its ERROR once more per skip. It restarts the side only if it had been started. Testemailsender/skip_no_credentials.emailsender:
skip-emailof an email in its mail transaction no longer stalls the queue. The stopped session stays in its transaction state until its C_TCP reports the close, and refused the next email there (“Event NOT DEFINED in state”); the close of a stop tells nothing, so the queue waited for the next email queued.C_SMTP_SESSIONnow takesEV_SEND_MESSAGEin those states too, and the email goes on the next connection. Testemailsender/skip_in_flight.SECURITY: a positional secret written with blanks is not echoed. A required
SDF_SECRETparameter given without its key and with blanks in it (set-password-pos correct horse battery) took the first word, and the answer “command ... with extra parameters” showed the rest in clear. It now shows'<...>', as for the same secret written askey=value. Testsecret_attrs.SECURITY: the
ieventstrace masks a command by its table. It masked by key names only, so a positional secret of a v7 command (in__md_iev__) went in clear; it now masks likeievents2(ievent_kw_masked()). Both traces log an ERROR and trace nothing if the masking fails, instead of a silent failure ofjson_pack(). Testsecret_attrs.yunetas-env.shcan be sourced withset -u. It read two variables that nothing sets ($dir,$pwd_opt), so the cloud SessionStart hook, which runs underset -euo pipefail, stopped there withYUNETAS_HOOK_BUILD_C=1, before the external libraries and the C build.gbuffer_serialize()of aNULLgbuffer answersNULL, logged. It read the gbuffer’s secret flag first: a segfault. Testgbuffer(test_gbuffer_guards).SECURITY: a secret json parameter that does not parse is not logged. A
DTP_JSON/DTP_LIST/DTP_DICTparameter markedSDF_SECRETwas parsed with the verbose parser, which logs the text it could not parse (an ERROR and a dump). It is now refused naming the parameter only (“parameter ‘<name>’ is not a valid json”). Testsecret_attrs.emailsender: a line or a DATA body that cannot be built is not sent half-built. The CRLF and the
.\r\nappended to it, and thejson_pack()of the frame, were not checked: the line is now refused with an ERROR, and the body ends the session as an error of ours. Only out of memory reaches them (no red test).Agent: a configuration file of a yuno with a long name is written. Its temporary file was the name of the file with 8 bytes more, so a name a few bytes under
NAME_MAXfailed withENAMETOOLONGand the yuno was not run. The temporary file is now.config.XXXXXXin the yuno’sbin/. Testhelpers(test_yuno_config_file).Agent: the temporary configuration files that an interrupted write left are removed at the agent’s start, for every yuno. Only the next launch of that yuno removed them, so a yuno not launched again kept them; both names (
.config.XXXXXXand the 7.25.21 one) are recognised. Testhelpers(test_yuno_config_file).fs_watcher: an entry whose
IN_IGNOREDwas lost in an overflow is forgotten. A wd stopped by the pass after an overflow (a directory deleted and created again, or no longer there) kept its entry for good when itsIN_IGNOREDwent with the overflow. It now goes once the stream is past where that event would have come. The directories the pass did not visit are asked to the filesystem, only those: nolstat()per watched directory.fs_event_tgains three fields, at its end. Testtimeranger2(test_fs_watcher_overflow).scripts/check_test_ports.pyfails on a file it cannot read, naming it: it skipped it silently, and its ports were not checked. Its files are closed after reading.SECURITY: a required secret parameter must be the last required one.
gclass_create()refuses, with an ERROR, a command table where a required secret (SDF_SECRET, or a secret’s name) is followed by another required parameter: written with blanks, the secret spilled a piece into the next parameter (shown in the traces and its errors) and the rest into the “extra parameters” answer. No command of the SDK or of the projects declares one. Testcommand_secret_positional.yev_loop: a kernel too old for Yuneta stops the yuno at its start, and says why.
yev_loop_create()probes the io_uring operations the loop uses and tries the cancel of everything with which a loop stops (IORING_ASYNC_CANCEL_ALL|ANY, Linux 5.19); missing one, a CRITICAL (“Linux kernel too old for yunetas: it needs Linux 5.19 or later (or RHEL/Rocky/Alma 9)”, withlacksandkernel) and an abort. On 5.6 to 5.18 a yuno started and could not stop its events. The minimum kernel is now documented (installation, README). (no red test: every kernel at hand has them; the probe checked on 7.0, 6.12 and Rocky 9.7’s 5.14)C_TCP: a connection dropped from inside its own
EV_RX_DATAno longer uses itself after itsEV_STOPPED.EV_STOPPEDis where the host of a volatileC_TCPdestroys it: a clisrv of the legacy method in a channel with noC_TCPof its own (in this tree, sgateway’s input). A host that dropped such a connection while handling its data (a websocket refusing a request that is not HTTP) destroyed it inside its own stack, and theC_TCPthen read its url (set_disconnected()), and over TLS took ytls’s-2222(the session was freed inside the callback) for a TLS error, wrote the cause and stopped the destroyed gobj again. NowEV_STOPPEDis the last thing aC_TCPdoes with itself, and-2222is a return:set_disconnected()ends there (the client’sidle_closedis consumed before it),mt_stop()releases the TLS session and the queues BEFORE the stop of its events, and both TLS backends answer-2222from the encrypt path too, wherewrite_data()used the released gbuffer. The volatile attrs are reset where a stop ends, no longer inmt_stop(). Testc_tcps/test7(red: a SegFault in clear and over TLS); themt_stop()order has no red test (a connected clisrv always has a read in flight, so its stop never ends insidemt_stop()).Agent: an orderly
--stopwith arun-yuno(orkill-yuno,play-yuno,pause-yuno) still waiting logs nothing. Each of them arms a volatileC_COUNTERthat waits for the yunos (30 s); the yuno stops only its direct children, so a counter still waiting reachedgobj_endrunning: “Destroying a RUNNING gobj” and “No subscription found” per counter. The agent’smt_stopstops and destroys them, without answering (their channels are closing). (no red test)SECURITY:
write-attrcan no longer plant a trusted signing key nor hand root to IdP users. C_AUTHZ’sjwksanddefault_rolewereSDF_WR|SDF_PERSIST, sowrite-attr-- guarded only by the per-command gate, off by default -- could add a JWK that the next start trusts, or setdefault_roleto root for every user an IdP provisions. Both areSDF_RD, the config’s. Testcommand_delete_user(red: the role was written).C_AUTHZ with no users treedb: the jwk commands answer again, and the user commands say why they cannot. The always-on permission check of this cycle asked a treedb that a yuno of local access only does not have (logcenter, auth_bff, sgateway, watchfs, dba_postgres): every command from a peer,
list-jwkincluded, got-403. The jwk commands, which never used the treedb, are answered as before; the user commands answer-1naming the cause. (no red test: no test yuno runs a C_AUTHZ without a store)stats=__reset__zeroes the counters of C_TCP, C_TCP_S and C_UDP_S. The reset writes the defaults of theSDF_RSTATSattrs, but these gclasses read their counters from private fields (mt_reading), which nobody zeroed: a reset changed nothing.mt_writingzeroes them now (txBytes,rxBytes,txMsgs,rxMsgs,max_tx_in_progress,refusedConnxs,noChannelConnxs,rxRefusedMsgs). Tested for C_TCP_S and C_UDP_S; C_TCP takes the same code.--pid-filewithout--startis refused at the start, instead of being ignored in silence: there is no watcher whose pid it could hold.A daemon closes every inherited descriptor with
close_range(). The loop up to_SC_OPEN_MAXmissed one inherited above a soft limit lowered after it was opened, and under the limit the agent’s CLI sets (1048576) it made a millionclose()calls. The loop stays for a kernel withoutclose_range(no red test: checked by hand, a yuno started with fd 50 open underulimit -Sn 20keeps only its own).C_YUNO
info-uptimenames the right error. It readerrnoafter the log that may change it.fs_watcher: the half of the open-files limit is the process’s. The warning counted the descriptors of each watcher alone, while
EMFILEis per process: four followers of 400 directories each, under 1024, never said it. It counts every watcher’s now, and says it again only after they fell under the half.fs_watcher: an end said when
FIONREADfails is never short, and aFIONREADthat keeps failing ends the watcher. The end added a read whole to what it could not count: with more queued, the owner did too soon what it left for after those events. It adds all the kernel can hold now (max_queued_events). And the turn that closes that end, whenFIONREADfails again, tells the ownerFS_WATCHER_GONEinstead of waiting for the next batch, for ever on a quiet watcher.C_TIMER0: a timer cleared and armed again in one turn of the loop runs. A pause and a play together (the control center’s rate tick) armed the timer while its cancel was in flight: “cannot start timer: is CANCELING”, and the timer stayed off until the next play. The arm is kept and done when the cancel ends, and that cancel is not published as
EV_STOPPED.ytls: the causes with no reason say it. A peer’s
close_notifyunder mbedTLS ended the connection as “TLS: decrypt failed” alone; data to encrypt before the handshake ended (both backends) and the mbedTLS write that makes no progress gave no reason either. They are said now (“the peer closed the TLS session (close_notify)”, “data to encrypt before the handshake ended”, “the write made no progress”), and C_TCP puts them in itsdisconnect_cause. Tested on both backends (mbedTLS linked by hand: the local build has OpenSSL only).timeranger2: a feed of metadata only is fed a record whose body is lost. A follower that could not read a record’s body (on the by-path fallback, a life already deleted) skipped the record for every feed; the feeds that want the body still skip it, the ones of metadata only get it (no red test: the fallback needs a system without /proc).
timeranger2: on a master, a file missing from a key still on disk is a damaged store, a CRITICAL.
Cannot open file to read/Cannot open md2 filewith ENOENT was a warning (“its key deleted under the reader”) everywhere. A master deletes a key whole, in its own process: when the key’s directory is still there, the reason is now “missing from a key that is on disk: the store is damaged”. A follower, or a key gone, keeps the warning.rmrdir()walks once more a directory filled while it was walked, and logs only a second failure. An entry made by another process after the lastreaddir()failed thermdir()withENOTEMPTY, an ERROR, even when the caller’s retry removed it (a reader closing an rt_disk feed the master was still linking into, which now calls it once).timeranger2: the leftover of a closing feed whose pid was reused is removed.
.closing.<pid>-<start>.<seq>carries the start time of the reader’s process; the master removes one whose pid lives but started at another time. Up to 7.25.21 it stayed while the unrelated process lived.--stopgives a daemon 10 s for its orderly shutdown. Every process of the name gets SIGQUIT at once, and only the ones still alive after 10 s are killed (said on stderr). Each was killed 1 s after its own SIGQUIT, one after the other: the agent had one second to stop its yunos and save, and a--stopalways took two seconds (the watcher, deaf to SIGQUIT, waited its whole second). The packages use--stopto move an agent into its unit. (no red test: checked by hand, a watchfs--stoptakes 109 ms and leaves nothing)fs_watcher: a subdirectory whose watch cannot be made is watched later, and timeranger2 reads its records. Out of watches (
ENOSPCatmax_user_watches) or memory, the directory was never watched: a follower took the key directory for gone and lost every record of the key, with an ERROR per directory. It is tried again at the end of each batch of the watcher, and handed as created again once watched (its owner reads what was made meanwhile); said once when it starts and once when every directory is watched.C_TCP: a subscriber’s error no longer stops the reading, and nothing is touched after a publish that destroyed the gobj. A subscriber of
EV_RX_DATAthat answered an error (-1, its own) stopped the reading: the read was re-armed only on0, used as a sign the gobj lived, and nothing stopped the connection either -- it hung in silence. The liveness is a marker now (a stack variable chained in priv, cleared bymt_destroy), around every publish that may end the connection or the gobj:EV_RX_DATA(clear and TLS), andEV_CONNECTED, after whichset_secure_connected()flushed the session andset_connected()armed the inactivity timer whatever the subscriber had done.start_pending_writes()says when its write ended the connection. And OpenSSL’s flush of the encrypted bytes leaked its gbuffer (and could spin) whenBIO_readfailed. Testc_tcps/test8(new; red: the second message never arrived; theEV_CONNECTEDdrop has no red test).C_PROT_TCP4H no longer reserves the length a peer announces, and a websocket frame of exactly
max_payload_sizeends. tcp4h reserved the whole length of a frame at its 4-byte header, before any payload, up tomax_pkt_size(the max block by default): a peer that wrote only headers reserved it each time -- the reservation C_WEBSOCKET dropped earlier in this cycle. Its buffer starts at 4 KB now and grows with what arrives. And in both, the buffer’s max is what the frame needs + 1 (a gbuffer holds one byte less than its max): C_WEBSOCKET’s wasmax_payload_size, so a frame of exactly that size never ended. Testsc_prot_tcp4h/test1(new; red: 10 MB grown for a header) andc_websocket/test2(the frame at exactly the max; red: it hung).SECURITY:
register-idp-userno longer hands any role to its caller’s pick. With arole, it checked only that the role existed, and the authz plane linked it when the IdP answered: a holder ofregister-idp-usergave root to an email of their own, with no permission on the users treedb. A role from a peer now needs theupdateoftreedb_authzs, ascreate-userasks to link one: C_AUTHZ’shas_roleanswersmay_linkfor the__username__the IdP passes on. Testcommand_delete_user(red: the request was queued).SECURITY: dbsimple no longer reads a trusted file that others can write, nor trusts a data directory’s owner. A root file (or, for a yuno run as root, one of the yuno’s user) was read whatever its mode: one left 0664 by a release before 7.25.19, in the group-writable data directory, could be edited by any member of the group. And the yuno’s user, for a root yuno, was the data directory’s owner, while its parents are 02775: a member of the group could rename it away and make one of their own. Such a file is refused now, and the yuno’s user is the owner of the nearest directory above that nobody else can write. Test
secret_attrs(case 7, red: the group-writable root file loaded); the root-yuno branch has no red test (a directory of another user needs root).C_AUTHZ: a username holding
%sno longer crashes the refusal. The “User not found” warning ofget_user_permissions()had a comma missing, so the username became the FORMAT of the log -- reached on every refused peer command since the always-on permission check, and a username is an email taken from a JWT. Testcommand_delete_user(red: SegFault).Agent: a node bounce relaunches the yunos once the killed ones are gone.
restart_nodes()(deactivate-snap,activate-snap) sent SIGKILL to every yuno and relaunched them at once; a SIGKILL is delivered, not done, so a new instance could meet the old one’s exclusive resources, or the launch found the old one alive and skipped it. It waits now for the pids it killed (every 100 ms, 10 s at most, then relaunches and names the ones left). (no red test; on wattyzer the relaunch began 1 s after the kills)Agent:
kill-yunoreaches a yuno alive but not connected. It selected only the yunos the agent saw running, so one that lost its channel or outlived a restart of the agent answered “Yuno not found or already not running” while it lived. Matching yunos alive in/proc(the search of the launch sweeps) now get the same signal on their watcher and child (signal2kill, SIGKILL withforce=1) and are said at once in the answer with their pids: nothing will say they closed. (no red test: no agent harness; checked on wattyzer with a stopped webstats, 2026-10-02)ytls: a session freed inside ANY callback is no longer read. The owner (C_TCP) frees the session from inside a callback when a write cannot start or a subscriber drops the connection. Only
on_clear_data_cbwas covered: afteron_encrypted_data_cbthe backends went on reading the freed session (the loop of the flush, the log of a failed callback, the state of the handshake), and afteron_handshake_done_cbthe decrypt went on with it. Every call that hands control to the owner now keeps a marker, the markers are chained (the calls nest) and the free clears them all: each call returns-2222on its way out --ytls_do_handshake()andytls_flush()included, which C_TCP now handles at the start of a handshake. Both backends. Testytls/test_free_inside_callback(red: a SegFault on the poisoned session; mbedTLS run linked by hand).The agents are native systemd units, and systemd sees them. The agent was a SysV script that
systemd-sysv-generatorwrapped: an agent started outside it (by hand, by an xscript) leftsystemctl status yuneta_agentsaying inactive while it ran -- an acceptance test on yunovatios central (Rocky 9.7) failed on it -- andyuneta_agent22had no unit at all.yuneta_agent.service(named like the script, so systemd takes it instead of the generated one) andyuneta_agent22.service(independent: either can stop or fail while the other keeps the node reachable) areType=forkingover the agent’s own daemon: a new option of every yuno,--pid-file, makes the process started with--startwrite the WATCHER’s pid before it returns, which is the unit’s main pid. The watcher keeps relaunching a crashed agent (Restart=no: systemd never starts a second one);KillMode=processkeeps the yunos across a stop;ExecStopasks only the unit’s own agent to shut down (not--stop, which ends every process of the name -- and systemd runs ExecStop also when the watcher ended on its own, so a second agent that met the first and left took the first with it). The package restarts both units (the main agent first, agent22 once it runs), stopping first an agent that runs outside its unit; the SysV script stays for another init. Checked by hand on wattyzer (Debian 13); through the built packages, and on Rocky, before the release (TODO.md). (no red test)The agents’ units: a removal of the
.debstops every unit, and a hung agent does not outlive its unit. Thepostrmdisabled the units through a script of the package, which dpkg had already deleted: a removal leftyuneta_agent22running from a deleted binary (and the web server’s unit enabled).prermdoes it inline now. And withKillMode=processthe final SIGKILL of a stop that timed out reached only the watcher: an agent hung in its shutdown stayed in the cgroup;ExecStopPostkills what is left of the unit’s own agent there. Checked on wattyzer with a SIGSTOPped agent22 (no red test).gui_agent’s bottom bar no longer clips its fifth item on a phone (gobj-ui 7.25.24, gui_agent 0.29.8, gui_treedb 0.17.75). Bulma’s
.levelput a 0.75rem gap between items that carry their padding: at 360px the bar scrolled 29px in Spanish, 5px in English. Measured with the real Bulma in Chromium and Firefox.fs_watcher: an end of the queued events said past the stream is reached on a quiet watcher. When
fs_queued_events_end()cannot see the watcher’s read (completions overflowed the ring,FIONREADfailed, the completions kept moving) it counts a read whole, about 8 KB past the real end; a timeranger2 follower that deferred the scan of a key directory to that end waited for as many bytes of unrelated events, for ever on a quiet feed. The watcher now notes such an end and closes it on the next turn of the loop (and after each batch while open): once nothing waits in the ring nor in the kernel, the stream jumps to it with anFS_BATCH_END. And aFIONREADthat fails inside a batch no longer answers the batch’s end alone, which could be short.fs_event_tgains two fields, at its end. Testtimeranger2/test_fs_watcher_overflow(do_test_padded_end, with a__wrap_ioctlthat failsFIONREADonce; red: noFS_BATCH_ENDcomes).dbsimple: an in-place save whose write stops half way writes the old content back. In a directory the yuno cannot write, the persistent attrs are saved in place with ONE
pwrite(): where the room cannot be reserved (NFSv3, FUSE) or the reservation does not cover a rewrite (copy-on-write: btrfs, reflinked XFS), anENOSPChalf way left the new start over the old tail -- a file that does not parse, so the next start loads the defaults and refuses every save. The old content is read first, the write is looped over short answers (the real errno is said), and a write that fails after some bytes writes the old content back, cut to its size and synced; the ERROR says “written back as it was”, or “UNPARSABLE” when that fails too. Testsecret_attrs/test_secret_attrs(8, with a__wrap_pwritethat writes 16 bytes and then fails; red: the old content not back).timeranger2: a follower no longer takes a key directory for gone when it cannot tell which one it is. On the path without a descriptor per directory,
dir_identity()answered FALSE with no log on ANY failure ofstatx(), and its callers skipped the directory’s scan: its links waited for the key’s next record, in silence. Now only a directory gone (ENOENT,ENOTDIR) is silent; astatx()refused (EPERMfrom a container’s seccomp profile,ENOSYS) takes the inode fromlstat(), with the warning of no birth time; any other errno is an ERROR, once per errno in a row. Testtimeranger2/test_dir_identity(red: 5 checks).C_TCP: the
disconnect_causeof a TLS failure says the backend’s reason on mbedTLS too. mbedTLS never wrote the reason whereC_TCPreads it (ytls_get_last_error()), so its causes were a bare “TLS handshake failed”, or "TLS: " with nothing after it. Every TLS cause is now what failed plus the reason when there is one (no dangling": "), the flush of clear data and the decrypt included; “cannot create the secure filter” no longer appends the"???"of a session that does not exist; and a cause longer than its buffer (now 512 bytes) is truncated with a warning, no longer in silence. Testytls/test_handshake_reject_mbedtls(the reason is empty without the fix; it needsCONFIG_HAVE_MBEDTLS).Control center: the rate tick is a real period, checked, and off while paused. It was a
C_TIMER, which checks its deadline on the yuno’stimeout_periodic(1000 ms) and re-arms from when it was served: a tick of 1000 came every 1 or 2 s, and a burst of one second in a tick of two was halved inmaxrxMsgsec/maxtxMsgsec. It is aC_TIMER0now (io_uring, periodic). Atimeoutunder 1 in the CONFIG was not checked (only a write was): the tick was off and the rates read 0 with nothing said; it is refused with an ERROR and 1000 put back. And paused, the rates read 0 instead of the last tick’s value. Testc_controlcenter_scenarios(test 12, red: the rates of the last tick while paused); the config check and the period have no red test.Control center:
__input_side__is not autostarted nor autoplayed in the realm configs of a.com (its operation repo, config 7, not deployed yet). The control center starts and plays it inmt_play, as__top_side__: autostarted, it listened from the yuno’s start, and an agent that connected to a control center run withplay=0logged “Publish event WITHOUT subscribers”. A realm config of your own does the same (the example is incontrolcenter.md).C_QIOGATE: the size of its queue can be read through a C_MQIOGATE.
msgs_in_queueandpending_ackswere said only by the gate’s ownmt_stats, and a C_MQIOGATE asks its children withbuild_stats(), which reads stats attributes and never calls a child’smt_stats:stats-yuno service=__output_side__showed no queue. They areSDF_RD|SDF_STATSgauges now, read live (mt_reading). Andmt_statsno longer reads a closed queue (trq_size()of NULL). Testtest_c_qiogate_stats(red: the counts missing for both children).stats-yunoon aC_IOGATEor aC_QIOGATEservice answered-1with no comment. Theirmt_statsreturned the bare data, written for a parent that reads its children, andC_IEVENT_SRVsent it back where the caller expects the envelope (gobj.h’s contract ofmt_stats). Both build it now (build_stats_response()), and aC_QIOGATEreads its bottom gate’s counters from thedataof the envelope. Testtest_c_qiogate_stats(red: no envelope from either).The msg/s of
C_IOGATEandC_CHANNELno longer depend on who reads them, and when. The rate was the delta since the previous read divided by WHOLE seconds (a read 1.9 s after the previous one divided by 1: +90 %), and every read moved the window, a read under a second too, so the messages of that second went into no rate. The window is now 1 s at least, measured in milliseconds, and a sooner read answers the last rates without moving it. Two readers still share one window (each one sees the rate since the other’s read); the agent’swatch-yuno-statsis the one reader of a node. (no red test)C_YUNO’s
uptimeis the yuno’s, in seconds. Described as “Yuno living time”, it was the machine’s uptime in jiffies, read from/proc/uptimewith every failure silent. It is now the seconds since the yuno’smt_create, on the monotonic clock. The ESP32 yuno does not write it (it reads 0). (no red test)C_YUNO:
info-uptime, the uptime of the machine. Seconds since boot (CLOCK_BOOTTIME, which counts a suspend, as/proc/uptime), the boot time and date it gives, and the yuno’s uptime and start date beside them. Testcommand_shutdown.C_TCP_S: a full server says it once a minute, as a warning, and counts it. With
child_tree_filter, a connection that found no free channel logged an ERROR each time, and peers retry: 600 channels and 1000 simulated controllers made 38,156 of them in minutes. It is now a third cause of refusal, said on the transition like the ip lists (the first one, then one a minute with the count), as a WARNING (a matter of capacity), and counted in a new statnoChannelConnxs--refusedConnxsstays for the ip lists. Testtest_c_tcp_s_ip_lists(phase 4).A yuno raises its soft open-files limit to its hard one, and the descriptors of a follower are watched. With
limit_open_files0(the default) C_YUNO left the limit as it came, so a yuno started from a desktop session ran with the soft 1024 ofsystemd --user, and a timeranger2 follower, which holds a descriptor per key directory of each feed, ran out (every record read failing). Now0raises the soft limit to the hard one (and agetrlimit()that fails no longer goes on with the limit unset). fs_watcher says once when its directory descriptors reach half of the soft limit, and a directory that cannot be opened (EMFILE) is said once, not per key directory: the next ones are counted, and the count is said when a descriptor opens again. Testscommand_shutdown(the limit) andtest_fs_watcher_overflow(dir fds limit).timeranger2: a follower does not hand a record without its body, and a read that finds its file gone is a warning. On the by-path fallback a record of a life already deleted could not be read and was handed with a NULL body to a feed that wanted the body; it is skipped now (the read logged it, and the delete comes next). And
Cannot open file to read/Cannot open md2 filewith ENOENT -- a key deleted under the reader -- are warnings (MSGSET_TRANGER, “gone: its key deleted under the reader”), not CRITICALs. (no red test: the fallback needs a system without /proc)SECURITY: C_AUTHZ’s commands ask a permission, always. They were
SDF_AUTHZ_Xonly, which does nothing whileenable_command_authzis off (the default): any user the entry gate let in created, updated, disabled and deleted users, set passwords, and added trusted signing keys (add-jwk), and linked a role throughcreate-user role=thatlink-nodesrefused. Now a command from a peer (its kw carries the__username__of the entry gate) asks the permission of the users treedb’s own C_NODE, aslink-nodesdoes:read,create(andupdatewhen it links a role),update,delete;-403without it. Internal calls are not asked. Testcommand_delete_user(case 12).SECURITY: dbsimple refuses a persistent attrs file of another user. The data directories are group-writable (02775), so a member of the group could plant the file of any service, and the persistent attrs of the agent and of logcenter held commands run with
system(); it was loaded, with a warning. Now only the yuno’s own file, root’s, or (for a yuno run as root) the data directory owner’s is loaded; any other is refused with an ERROR and the saves with it. A save as root gives the new file only to the data directory’s owner (it gave it, secrets in it, to whoever owned the old one). Testsecret_attrs(case 7, with__wrap_fstat).SECURITY: the commands the agent and logcenter run with
system()are config only.write-attrcould set and persistcert_sync_copy_cmd(run as root through sudo, and at once bycert-sync-now, which had no flag) andrestart_yuneta_command. They areSDF_RDnow, andcert-sync-now/cert-sync-statusareSDF_AUTHZ_X-- which, like everySDF_AUTHZ_X, guards them only where a yuno setsenable_command_authz(off by default, and no yuno sets it yet): the protection that holds today is the attrs being config only. (no red test)emailsender’s
skip-email,remove-emails-failed,disable-alarm-emailsandenable-alarm-emailsareSDF_AUTHZ_X: they move or remove emails, or silence the alarms. The flag does nothing until the yuno setsenable_command_authz(off by default): it marks them for the day the gate is on. (no red test)A frame before the session is not parsed when it is big. C_IEVENT_SRV parsed a peer’s whole frame before checking its identity card, and a frame of
[{},...]takes ~100 times its size to parse (16 MB: 1.7 GB): the max block of a yuno let an unauthenticated peer ask for ~20 GB. A frame bigger thanmax_pre_session_frame(64 KB) is refused with a warning, unparsed and not dumped. Testc_ievent_srv_identity_card(card 8).C_WEBSOCKET no longer reserves the length a peer announces. The payload buffer was created with the frame’s length from its header, before any of it came: a few connections each claiming the max block reserved it each. It starts at 4 KB and grows with what arrives, up to a new
max_payload_size(0: the max block); a bigger frame is refused with a warning and 1009. Testc_websocket/test2.gbuffer: a buffer grown from small reaches its max. The growth doubles, and a doubling past
max_memory_sizewas refused even when what was needed fitted (“MAXIMUM SPACE REACHED”): 4 KB with a max of 10 KB stopped at 8 KB. It grows to the max now, when what is IN it plus what is needed fits (a websocket frame of exactly its max, read into a smaller buffer, is the case that needs it). Testgbuffer(test_grow_to_max).timeranger2: closing a follower’s rt_disk feed under load no longer leaves its directory fed for ever.
tranger2_close_rt_disk()removeddisks/<rt_id>/where it was, while the master kept linking new records into it: a key directory made in between made the removal fail (“rmdir() FAILED ... Directory not empty”, every orderly stop of adb_history_cereading 3000 records/s), the directory stayed, and the master, which closes its side only when it hears the directory go, went on making a directory and a link per key and file into it (an id that never came back grew for ever). The reader now renames the directory away (disks/.closing.<pid>.<seq>) and removes the renamed one, where the master cannot reach it; the master’s watch ofdisks/hears the rename as the close (newFS_FLAG_MOVED_AS_DELETEDof fs_watcher), its link of a feed that is closing is quiet, and it removes at its next read ofdisks/the.closing.*of a reader that died while removing it. A leading dot is no longer a valid rt id (“Invalid rt id (a leading dot is reserved)”, a warning): no id of the SDK or the projects has one. Testtest_delete_key_propagation(close_races_master; red: the directory left, the master still feeding it).Agent: a yuno alive but not connected is no longer launched again.
yuno_running/yuno_pidsay only what the agent saw on the yuno’s channel and are not persistent, so a yuno that lost its channel, or that outlived a restart of the agent (a yuno still loading 13.7 M queued messages, 2026-09-26), was launched a second time by the boot sweeps or byrun-yuno: the second instance opened the persistent queues of the first “as not master” and aborted, and when the first reconnected the agent killed it as the intruder. Before a launch the agent now looks in/procfor a process of that yuno (its role asargv[0], its configuration files in itsbin/on the command line -- notyuno.pid, which a yuno writes where its own config says) and skips it with a WARNING naming its pids (“yuno alive but not connected to the agent: not launched again”); arun-yunothat finds only such yunos says so. A node bounce (restart_nodes(),deactivate-snap) kills them too, SIGKILL with a WARNING, before it runs them all again.kill-yunostill acts only on yunos with an open channel. Also: a yuno whose pid file cannot be written logs it, instead of a SegFault infprintf(). (no red test; checked on the local agent: two yunos stopped by SIGSTOP across a restart of the agent were not launched again, the WARNING named their pids, and they reconnected on their own when resumed)Docs:
kill -9on a yuno’s child ends it for good.ENTRY_POINT.md§7 andYUNO_LIFECYCLE.md§5.3 said the watcher relaunches a child killed by SIGKILL; it exits on it (ydaemon.c, as the table ofENTRY_POINT.mdsays). Any other signal is a crash and is relaunched.mbedTLS: a failed output callback no longer frees its gbuffer twice. The gbuffer is the callback’s whatever it answers, and C_TCP’s releases it with the kw of
EV_SEND_ENCRYPTED_DATAon every path; the backend released it again (a C_TCP that refused the output, inST_WAIT_STOPPEDwith its TLS session alive). OpenSSL ignored the answer: it now logs it. The ownership is written inytls.h. (no red test: found by reading)
The reborn key of an rt_disk follower (the HIGH open since 7.25.20)¶
An rt_disk follower behind a master in another process no longer hands a key’s later life before its delete. It took the link of a new record and read the key’s files BY PATH, so after the master deleted the key and wrote it again, a late event of the old key directory consumed the NEW life’s link and read it against the old life’s cache:
[R2 DEL], the live key out of the cache, the next append handing R1..R4. The identity check of 7.25.21 (inode and birth) could not tell the directories apart on the nodes: ext4 on Debian’s 6.x kernels gives a directory made again the inode and the birth time of the old one, andtest_delete_key_propagation(“race in the batch”) failed there 9 runs in 10. Now each key directory is watched with a descriptor of its own, the very inode of its watch (newFS_FLAG_DIR_FDSandfs_watcher_dir_fd()in fs_watcher, newevent_wd/subdir_wdon every event); a link is taken through the descriptor of the directory its event came from (openat/unlinkat), and its records are read through descriptors checked to be the link’s life. Under a random load of a master in another process (appends, deletes, rewrites 0-20 us apart, a follower slowed 0-100 us a record, 3 feeds and 4 keys), 7 to 11 of the 12 feed and key pairs got a life out of order; none now, on the dev kernel and on wattyzer’s 6.12, and the last life of every key is heard whole. Cost: none on the append path (22.4 us a record and feed, both), +5% on the 69632-key flood oftest_rt_disk_overflow(18.1 -> 19.0 s, an open and a watch through /proc per new key directory). Teststest_delete_key_propagation(new “a stale link”, fails on the code before it locally too:[R2 DEL]; “race in the batch” passes 30 in 30 on wattyzer, 1 in 10 before) andtest_fs_watcher_overflow(new “dir fds”).Out of descriptors, that read no longer uses memory freed or a descriptor closed. The check of a life borrowed its record of what was checked, then opened the content: out of descriptors, that open closes every read descriptor, the md2 just checked and that record too, and the check wrote into the freed record and handed the md2’s number, by then the content’s, to be read as md2 rows. The md2 is now looked for again after the open: closed, the check starts again, once; a second time, the records are not read, with an ERROR. (no red test: found by reading)
A link consumed already is not read again by path. The scan of a key directory and the link’s own event can both meet a link; the second one found it gone and read the key’s files by path, which is the race this section fixes (a new life read against the old one’s cache) and re-read what the first one had read, or had left unread for being another life. The first one unlinks and THEN reads, so nothing is lost: a record made later makes a new link. (no red test: a window of microseconds)
fs_watcher: a watch that cannot be made is said as what it is. A failure of the watch through
/proc/self/fdwas always taken for a missing/proc, said once per process and retried by path; with no watches left (ENOSPC,ENOMEM) that spent the only/procmessage and a second watch. Those two now fail at once with their own ERROR. The test “dir fds” now checks that no descriptor is left open after the stop.
Risks the review named¶
timeranger2: a delete no longer walks every open key. To close the descriptors of the deleted key, the master (at
tranger2_delete_key()) and each follower feed that heard it walked every key with an open file. With 50000 keys open, 2000 deletes and 4 feeds: the master 9.2 s -> 0.17 s, the follower 4.3 ms -> 17 us per delete and feed. What stays is linear in the feeds (fs_queued_events_end()per other feed, to note what each one owes): 17 us per delete and feed with one feed, 30 with sixteen; it was measured and kept, a mark taken once per batch could make a debt be forgotten unpaid. (no red test: a benchmark, figures above)C_TCP: a stop waits
timeout_stop_txat most for a write in flight. A peer that does not read (a window of zero) held the write, the stop of the connection, and so aC_TCP_Sstopped and started again: its clisrv waited inST_WAIT_STOPPEDfor ever and the server never listened again. The socket gets aTCP_USER_TIMEOUTwhen the stop has to wait for a write. Testtest_c_tcp_s_stats.C_TCP_S: a server of the new method leaves to another running one the channels it serves, with an ERROR naming it; it took them all, silently (they accepted on its socket, the other server accepted nobody, and its stop and its counts missed them). Test
test_c_tcp_s_stats.SECURITY: the agent’s audit record redacts a name with a secret’s part whatever else it holds (new
is_secret_name_any()):api_key_max,token_mode,secret_typewere written in clear, being names ABOUT a credential foris_secret_name(). Testtest_audit_record.SECURITY: the subscription traces mask their
__filter__/__config__/__global__. The traces of a subscribe, of an unsubscribe and of the filter at a publish (subscriptions/ev_kw/machine) printed a credential in them as it was. Testsecret_attrs.timeranger2: a filesystem with no birth time is said (a WARNING, once per process): the identity of a key directory is then its inode alone, and a key deleted and written again can take the freed inode. (no red test: needs such a filesystem)
timeranger2: the note of a deferred key directory scan fits its numbers (it held 96 bytes for up to 105), and a truncation would be an ERROR. (no red test: unreachable now)
Traffic dumps: the masking of received bytes is per buffer, documented with an example: a credential cut between two reads shows in the second dump. What we send is marked secret at its source and never dumped.
v7.25.21 (2026-10-01)¶
What changed after 7.25.20. Each behaviour change has a test that fails on the code before it, except the few this list marks “(no red test)”.
Upgrade steps (operators, read first)¶
Nodes that build from source: rebuild the external libraries first. linux-ext-libs is 1.23 (openresty now comes from its release tarball; no library changes version): run
cd kernel/c/linux-ext-libs && ./extrae.sh && ./configure-libs.sh, thenyunetas clean && yunetas build(it refuses to build until then).Rebuild every project against the new headers.
gbuffer_tgained a field (secret) in the middle of a public struct whose accessors are inline: an object built against the 7.25.20 headers reads the wrong fields.Deploy gui_agent 0.29.7 before, or together with, the agent and the control center. The agent now answers a
watch-yuno-statsonly to a requester that says it takesEV_YUNO_STATS(__relays__), and the control center passes that on only when its client says so. An older gui_agent is refused the watch, directly and through a control center: it still works, polling the stats instead, with one warning per node in the browser console. gui_agent 0.29.7 and gobj-js 7.25.9 work with the 7.25.20 agents too.A host of
C_FSmust not subscribe to it by hand any more.C_FSnow subscribes its parent (or itssubscriberattr), as every child gclass does; a host that still callsgobj_subscribe_event()on it gets “subscription(s) REPEATED, will be deleted and override”. watchfs is updated. AC_FSwithrecursive: falseno longer reports subdirectories.emailsender stops (exit 0, not relaunched) when the SMTP server refuses this client, like refused credentials: a 5xx to the greeting (the 554 of a provider that blocked the address) or to EHLO, and a second
334toAUTH PLAIN(the first is answered once, as RFC 4954 allows). Every further attempt would be one more refused login in the provider’s logs. Fix the cause and run the yuno again.Send a test email after upgrading emailsender, and after
set-email-user. It no longer logs in at start, so wrong credentials or a server that refuses this client are found at the first email, which then stops the yuno (exit 0, not relaunched).The agent needs write permission on each yuno’s
bin/: a config file is now written to a temporary file there and renamed over the old one. The files are 0640: a reader outside the group of the agent’s user loses access.emailsender’s failed queue fills much later. A server that fails its handshake no longer sends the head of the queue there every 8 s or so, one message after another: the messages wait at the head of the queue, paced up to 10 minutes between attempts, for as long as the outage lasts (a refused default sender for up to
max_retriespaced attempts), and one that refuses every message is paced from its third refusal in a row. An alarm on the size of the failed queue sees that change. A stuck head of the queue can be moved there by hand with the newskip-email.A lone stop of a
C_TCP_Swithoutchild_tree_filtercloses its live connections (it kept them accepting on a closed socket): to drain a listener, do not stop it alone.emailsender says when its SMTP server fails: a WARNING at the first failure, an ERROR once it has failed for
timeout_failing_alarm(1 h), an INFO when it works again. An alarm on its ERRORs sees the new one.C_TCP_S
connxsandtconnxsread real values (they always read 0); a dashboard that showed them changes.
Performance, against 7.25.20¶
The report of this release measures it against 7.25.20, the release of the last report:
performance/reports/7.25.21.html, and the trend on doc.yuneta .io /performance. Both releases built whole, run alternately, 8 rounds (24 for the timeranger2 tests). Nothing slower beyond its noise: the masking of secrets runs only when a trace prints, a secret gbuffer costs one flag test, and a TCP send one key lookup. Appends 223,031 -> 225,883 per second (+1.3%, noise).
Three figures moved outside their spread, all faster, and none is claimed: the gobj TCP echo +4.8%, the BFF logins +3.1%, 600 000 appends to one key -3.7% of time. No change takes work off those paths (the append code is the same as in 7.25.20): placement of whole-release builds.
The static binaries are 0.6-0.9% larger (stripped).
SECURITY: secrets that still reached a reader or a log¶
SDF_SECRETon the credentials that were not flagged.C_IEVENT_SRV.http_cookieholds the browser’s whole Cookie header, the BFF’s httpOnly access_token included, and view-gobj, view-gobj-tree and view-attrs showed it for every connected browser. It is now masked, as are the MQTT 5auth_dataof C_PROT_MQTT2/C_PROT_MQTT, webstatsvisitor_salt, the esp32 transport’sjwtand C_ESP_WIFI’swifi_list. (Each flag is a declaration; the masking it drives has red tests,http_cookieandauth_datatoo.)A secret is masked whatever its json type.
gobj_mask_secret_attrs()andview-attrsof one attribute masked only non-empty strings, so"pin": 1234was shown. Only an absent value or an empty string stays.gobj_mask_secret_attrs()now logs an ERROR and returns -1 for a NULL gobj or attrs that are not a dict (it returned 0 in silence).view-configmasks the secrets. It dumped the config as it is, so a password or client_secret in a service’s kw, a child’s kw or aglobalkey reached whoever asked. The newgobj_mask_secret_config()masks, on a copy, every value that lands on anSDF_SECRETattr, of any json type; aglobalkey named by a gobj that is not a service, or of an unknown gclass, is masked when its attr isSDF_SECRETin any gclass; and a__json_config_variables__entry is masked when it feeds a secret or has a secret’s name. The[^^children^^]template of a service is masked too (its__content__as a node, its__vars__as config variables), here and in thecreate_delete2traces.Command parameters can be
SDF_SECRET, and a trace masks every key with a secret’s name. Thecommandstrace printed the command line, and withev_kwthe command kw, in clear: passwords given toset-user-pwd,create-user,set-email-user, and through the agent’scommand-yuno ... password=X, whose free kw no table describes. A parameter declaredSDF_SECRET, and any key named like a secret (is_secret_name(), new inhelpers.h: the one list of the SDK, which the agent’s audit uses too), is shown as********in all three traces (newcommand_mask_secret_kw()/command_mask_secret_line()), as is thevalueof awrite-attrof a secret attr. A positional secret that holds a=is masked, and a part of the line that cannot be parsed is shown as<...>instead of being dropped, and a quoted value keeps its closing quote. The parser’s error answers no longer echo a secret (a plain extra word is still named; any text after a secret parameter is masked whole). A parsed key containing a blank (password= hunter2 x=1reads the keyhunter2 x) refuses the command and is not echoed; anSDF_WILD_CMDcommand took it as a free key. Flagged: C_AUTHZpassword, C_IDP_KEYCLOAKkc_admin_client_secret, ycliuser_passw, C_PROT_MQTTcreate-userpassword, emailsenderset-email-userpassword(a red test covers the last; the others are the same one-word declaration).Every kw the kernel dumps on an error or in a trace masks the credentials. The
kw_get_*()errors (“path MUST BE a json str” with the whole kw, so a password read with the wrong reader landed in an ERROR log), themachinetrace withev_kw, the authz traces, “No subscription found”, “Publish event WITHOUT subscribers”, and theievents/ievents2traces of C_IEVENT_CLI/C_IEVENT_SRV, which printed a command’s password in clear both ways. They go through the newgobj_trace_json_masked()(json_mask_secrets(),mask_secrets_inline()): a key with a secret’s name is masked at any depth (down to 64 levels, a dict met twice the same everywhere, a cycle as"<cycle>") and of any type, and so is aname=valuecredential inside a string (an unquoted secret value runs to the nextword=, so a password written with blanks is masked whole; a quote inside a value is part of it); the 64 levels count from the first time a shared object is met; a name that only describes something about a credential (token_endpoint,cookie_domain,jwt_public_keys,*_count) is not. The ievents traces also use the command table of the destination service when it is local. gobj-js 7.25.9 masks its traces the same way. Withgbuffersthe kw’s gbuffer is dumped, masked (a secret gbuffer as hidden, one that parses as json through the same masking). Masking is linear and bounded, because a peer reaches it before its session (the dump of the kw of an event before the identity card). A name is the word before the=, judged by its last 128 bytes. Every json node walked costs one unit of the budget, a key or a string its bytes too: 4 MB perjson_mask_secrets()call, ormax_bytesfor the newjson_mask_secrets_capped(). Once the budget is spent the walk stops and one “” stands for the rest of each container; a text over 4 MB, or one whose masked copy cannot be allocated, is that placeholder too, never shown in clear. The work is bounded by the cap, not by what was sent.C_IEVENT_SRVcaps that dump at four times the 256 bytes it shows. gobj-js masks with the same bounds.The traffic dumps mask the credentials they can recognise, received data included. No sender can mark what it receives, so the
traffictrace of a server gate printed a browser’s Cookie header or a form password in clear.gobj_trace_dump*()now dump through the newmask_secrets_in_text(), which writes*over the value of an HTTPCookie,Set-Cookie,AuthorizationorProxy-Authorizationheader, of a secretname=valueand of a json"name": valuewith a secret’s name (a list or a dict up to its closing bracket), keeps the length, and says"masked": N(gobj_trace_dump()gives{"len", "masked", "data"}). A secret in received bytes that fits none of these is still dumped as it came.The
create_delete2trace no longer prints secrets: its kw and sdata dumps showed the passwords a gobj is built with (the smtp child’s, ycommand’suser_passw).Secret gbuffers: hidden from the traffic trace, wiped on free, never serialized. The C_TCP
traffictrace dumps outgoing bytes before TLS, so a credential sent as data -- the SMTP AUTH line -- reached the log. Newgbuffer_set_secret()/gbuffer_is_secret(): every trace dump prints a secret gbuffer as<N bytes hidden>, and it is zeroed when freed or grown;gbuffer_append_gbuf()flags its destination before it copies. C_TCP takes"__secret__": truein theEV_TX_DATAkw, and the flag survives the tx queue. A secret gbuffer is not serialized:gbuffer_serialize()answers NULL (logged), and a kw carrying it crosses without it. The emailsender marks its AUTH line.The persistent-attrs file is written to a new 0600 file and renamed over the old one. It was truncated and written in place, into whatever inode the name had (another user’s file, a hard link), and a password set once and never saved again stayed in a 0664 file. A save now writes
<file>.tmp-XXXXXX(0600), renames it and syncs the directory; a failed save leaves the old file as it was, and a link in its place is replaced, with nothing written through it. A readable file of another user is taken over by the yuno’s user (logged at INFO); when the yuno runs as root, the new file is given back to the old owner (no red test for these two: a non-root test cannot make a file of another user). A file that cannot be read refuses the save and the removal (its attrs would be lost); an empty file is no data. A file of the yuno’s own found at load that is wider than 0600 or a hard link is replaced the same way (not narrowed in place, which through a hard link changed another name). In a data directory the yuno cannot write, the save goes in place, into a file of its own only: room is reserved without growing the file, so a full disk leaves the old file as it was where the filesystem can reserve it (not on NFSv3/FUSE without fallocate, nor on copy-on-write filesystems), but a crash or a short write there can leave an unparsable file, which then refuses the saves until it is removed or repaired (the message says how). A<file>.tmp-XXXXXXan interrupted save left is removed at the next load.write-attranswers a failed save with -1 (“written, but NOT saved”) instead of “done”, and says “NOT persisted” for an attr of a gobj that is not a service.A yuno’s materialised config files are 0640, never world-readable. The agent wrote
bin/<n>-<role>^<name>.json, which carries the yuno’s secrets, as 0664, and an existing file kept its old mode. Each is now written to a temporary 0640 file inbin/and renamed over the old one, so a file of another owner or a symbolic link is replaced, not truncated and then refused, nor followed, and a failed write leaves the old file whole. Files an earlier launch left behind are narrowed, and the temporary files of an interrupted write are removed (logged).The agent’s audit never writes a peer field raw when it cannot redact it. A failed allocation inside the redaction wrote the JWT or
password=it was meant to hide, or 4 KB never scanned. It now writes<not written: no memory to redact it>and logs it. A peer field that is not UTF-8 is a warning, not an internal error.C_IEVENT_SRV: the dump of a peer’s kw before its session hides every credential, at any depth and of any type, and
name=valuecredentials in a command line (it dumped the whole kw throughtrace_inter_event).MQTT: the CONNECT decode trace no longer prints the password, and prints the username by its length (it read past a field that is not NUL-terminated, so a long password also appeared on the username line). The same over-read is fixed in the SUBSCRIBE/UNSUBSCRIBE traces and the “invalid will topic” warning. The dump of a refused CONNECT no longer carries the credentials: a CONNECT with one extra byte after a valid password put the password in the log; of a CONNECT only the protocol name, level, flags and keep alive are dumped now, of an AUTH only its reason code.
A value with no closing quote refuses the command.
password='abcwith no closing quote dropped the parameter, and the command ran without it and with no log. It is refused now (“value with no closing quote”), naming only the key.
Kernel¶
A renamed subscription is no longer a repeat of a plain one, nor of another rename. Since 7.25.5, subscribing
EV_Xwith__rename_event_name__found an existing plainEV_Xsubscription of the same subscriber as REPEATED and replaced it, and two renames of one event collapsed into one (withdrawing either removed both). The renamed event is part of the match now. A kw with no rename is still a wildcard, so a plain subscription made over a renamed one replaces it, and a plain unsubscribe removes both (seepublish.md). gobj-js 7.25.9 renames as C does.gobj_unsubscribe_list()removes the subscription it is given, not the first one that looks like it. A stale plain subscription (already removed) removed a live one, e.g. one with a__filter__. The object itself is looked up; one no longer there removes nothing, is not passed tomt_subscription_deleted(), and gives one warning per call (“Subscription(s) already removed, nothing to remove”), and one withdrawn bymt_subscription_deleted()meanwhile is not taken for a hard one kept.A subscription the publisher refuses leaks nothing. When
mt_subscription_added()answered -1,gobj_subscribe_event()removed it but never dropped its creation reference: each refusal (a subscription C_IEVENT_CLI could not send to its peer) leaked one.A subscription list that outlives its publisher is safe.
_delete_subscription()read the publisher out of the subscription it was given, freed memory once the publisher was destroyed (a list fromgobj_find_subscriptions()keeps the subscriptions alive). A removed subscription now has itspublisherandsubscriberset to 0, and such a list removes nothing. A subscription removed during a publication is no longer delivered from the copy being iterated.C_IEVENT_SRV: an identity card it refuses is the peer’s, one capped warning. Before the session, a card with no routing, for another role/yuno/service, or with no src role/service was an ERROR with the whole kw (jwt included), plus a second ERROR “event UNKNOWN in not-session state”. Each is now one warning (protocol category, peername, kw capped, credentials hidden) and the channel is closed. A jwt that is not a string is refused instead of authenticated, and an event before the card gets one warning.
C_TCP_S:
connxsandtconnxscount the connections. Both always read 0:connxswas never decremented, and servers withoutchild_tree_filternever saw their accepts. Now they are the connections held and those accepted since the gobj was created (a stop and a start do not reset it), for both accept methods, as attrs and instats, and a connection counts for the server whose clisrv it is (new C_TCP attrtcp_s), so two servers sharing channels and a port on two hosts do not count each other’s. C_TCP’s ownconnxsis readable too. The clisrv names stay unique across restarts.C_TCP_S: a stop and a start in the same turn listen again. The start made a second socket while the first was still being cancelled; its bind failed and, by default, the yuno exited. It now waits in ST_WAIT_STOPPED and listens when the stop ends.
C_TCP_S: a TLS server started again keeps its ytls and reloads its certificates. Each start freed the ytls and made a new one, while the connections that outlived the stop (
child_tree_filter) still used it: the next record of such a connection was decrypted with freed memory. A restart now applies the new certificates to new connections, asreload-certsdoes, and live connections keep theirs until they close. It reloads only when thecryptoconfig (trace_tlsincluded) or one of its files changed (inode, size, mtime, ctime), so a pause and a play no longer log “TLS certificates reloaded” each time; a failed reload is still an ERROR, and an update of the system CA bundle (ssl_use_system_ca) needsreload-certs.C_TCP_S: a lone stop of a server without
child_tree_filterstops its clisrvs. They kept accepting on the closed socket, and the next start logged “GObj ALREADY RUNNING” for each one. In that method the clisrvs are its connections, so a lone stop now closes its live connections: a host or an operator that stops only the listener to drain it (disable-gobjof the listener,gobj_stop()of theC_TCP_Salone) closes its sessions instead of keeping them. Withchild_tree_filterthe connections still survive a stop of the listener.C_TCP: a clisrv of the new accept method started again leaks nothing and waits for its accept in ST_DISCONNECTED. Its start created a new accept event over the one of its last start, never destroyed (a leak, and “Destroying a running event” at shutdown), and it stayed in ST_STOPPED, so its next stop never cancelled the accept, which kept the port bound.
yev_loop: freeing a duplicated accept event no longer closes the listener’s socket. It closed that socket number, and took back the submissions on it, even when the number had been reused by then.
C_IOGATE: “send to all” delivers the whole message on every channel. With
send_type1 (or__send_type__1) a message in a gbuffer was shared by all the channels: the first read it out and the others sent an empty frame, on which aC_PROT_TCP4Hpeer drops the connection. Each channel but the last now gets its own copy (with the secret flag, label and address), and the last takes the original.C_TCP: data sent while a connection closes is said once, with its count, on every end of the close, destroy included. It was dropped with no trace at the default levels; now one warning per connection gives the messages and bytes dropped.
C_TCP:
connect_on_startlets the owner make the first connection on demand;disconnect_causesays why a connection ended (new attributes).connect_on_startfalse keeps a client inST_DISCONNECTEDat its start until its owner sendsEV_CONNECT. Only that first connection waits: after an end the client reconnects aftertimeout_between_connectionsunless that is-1, which makes every connection on demand.disconnect_causekeeps the FIRST cause of an end (the read cancel a drop or an inactivity close makes does not replace it): the socket error, a TLS failure with its reason, “Local dropping”, “Inactivity timeout”, “Local stop”, or a write or connect that could not start. It is emptied when a connection begins, a clisrv’s included, and the Disconnected trace shows it as its cause. (no red test for the TLS and fault-injection causes)C_TCP: an
EV_CONNECTcancels the pending reconnect timer, so a connect on demand no longer gets a strayEV_TIMEOUTinST_WAIT_CONNECTED.C_GSS_UDP_S starts its C_UDP_S again when it stops by itself. The service stayed deaf and every send logged “Event NOT DEFINED in state”. The stop is said once. A send is refused (one warning per stop, the count at the restart or at the stop) whenever the C_UDP_S is not in ST_IDLE, including while it is still stopping. The C_UDP_S restarts after a backoff:
timeout_basefirst, doubling up to 5 minutes while a try does not hold a minute, and a failed start takes the same path. A C_UDP_S stopped from outside is not restarted. Atimeout_base<= 0 is refused with a warning and 5000 is used. A stop and a start in the same turn (logcenter’s pause and play, with a read in flight) found the C_UDP_S still holding its socket: an ERROR, and the service never received again. The start now waits for the C_UDP_S to end its stop.A build with mbedTLS compiles again.
C_ASSETSincluded<mbedtls/md5.h>, which mbedTLS 4 moved undermbedtls/private/, even when OpenSSL was compiled in too: any.configwithCONFIG_HAVE_MBEDTLSfailed to build. On mbedTLS the md5 of a signed url is nowpsa_hash_compute().Static resolver: v4-mapped literals answered as glibc answers them. “127.0.0.1” in
AF_INET6withAI_V4MAPPEDreturns::ffff:127.0.0.1, and “::ffff:1.2.3.4” inAF_INETreturns1.2.3.4; both wereEAI_ADDRFAMILY. An IPv4 literal is numeric in every form glibc takes:"127.1","10.1.2","0x7f.1","0177.0.0.1"and"2130706433"went to DNS, andAI_NUMERICHOSTrefused them; a blank around a literal still makes it a name.C_FSpublishes by the value of the event type, watches the tree once, recursively only when asked, and subscribes its parent. The fs_watcher types are values, not bits: a deleted directory published nothing and leaked its kw, and other types published by accident. A recursiveC_FSalso added a recursive watcher per subdirectory, so a change N levels down was published N+1 times, and"recursive": 0reported the subdirectories anyway. Now one watcher on the root, recursive only withrecursiveset. It subscribed nobody (its host had to do it by hand); it follows the CHILD model now, with asubscriberattr.size_dl_watchreads 1 while watching, 0 otherwise. A watcher that cannot be created, or whose read cannot be armed, fails the start (it answered 0 with nothing watched; the second case has no red test).
timeranger2 and its tools¶
A feed that overflowed hears the deletes another feed of its topic heard, and hears each delete once. The feeds of a follower share the topic’s cache, and the first to hear a key-delete removed the key from it; a feed whose inotify queue overflowed then compared the cache with
keys/, found nothing gone, and never got itskey_deleted. Each feed now keeps the deletes it still owes. A delete signal queued behind the overflow marker, for a key the overflow already reported, is not reported again (it firedkey_deletedtwice, also with a single feed); every event of a watcher has its position in the watcher’s stream (fs_event->offset), andfs_queued_events_end()says where the events queued so far end, a read that the kernel completed and the loop has not delivered included (newyev_get_waiting_completion()). A feed that starts watching after a delete was signalled does not owe it. A debt is forgotten by a record of the key that reaches the feed in its place in the stream: a link heard in the feed’s directory, or a record found by the deferred read of a key directory (see the reborn-key bullet below). In a master, hearing a delete makes no debts;tranger2_delete_key()makes them.A key deleted and written again (once or more) before a follower reads the delete is handed right in far more cases, for any key, one the follower never saw included (with the master in another process and the follower lagging, a race remains: an open defect, see
TODO.md): the feed is tolddeleted(once or more), then the records of the key’s last life from rowid 1, and the key stays in the cache; records of a life deleted before the follower read them are not handed. A follower reads a key directory only if it is still the directory it looked at before it asked where its queue ends, and only once its stream is past that point (each event carries itsoffset_end; newFS_FLAG_BATCH_END/FS_BATCH_END_TYPE); a file linked in a directory that waits to be read is read with it, in order; on a delete the follower closes its descriptors on the key’s files. A record found then forgets the debts made before the directory appeared; a feed opened while a delete is in flight, and signalled after it was watched, no longer takes that delete as new; a delete owed at an overflow whose key is back on disk is paid, not taken as new. Before: a directory read too early ([R1 DEL],[DEL R1 DEL], the live key out of the cache, the same after an overflow), a second file of a new key first ([R1 R1]), a feed made to owe a delete again ([DEL DEL]at its next overflow), and a key read before, deleted and written again in the same day file, read through the old file’s descriptor (short reads, CRITICAL, lost records). An fs_watcher read takes up to 32 events; the 69632-key overflow flood drains ~7% slower (one morestatxper key).fs_watcher: a root that cannot be watched gives no watcher (logged): with
max_user_watchesused up, or no permission, the watcher ran watching nothing.Only the first feed to hear a key-delete forgets the key (the shared cache, its segments, the watermark of every feed). Every feed did it, so a slow feed hearing an old delete after the key came back removed the live key from the cache and dropped its fresh watermark.
In a master, the watcher’s echo of a key-delete no longer removes the key: a key the master wrote again before the echo came was dropped from its cache.
tranger2_delete_key()has already removed it; the echo only calls the feed’s callback. And a master’s own rt_disk feed whose inotify queue overflows is told the deletes it lost:tranger2_delete_key()makes the debts for the master’s watched feeds, whose cache cannot show a lost delete.A watcher whose read fails is logged and tells its owner before it goes (
FS_WATCHER_GONE_TYPE). It went silently, and a timeranger2 feed later used, and stopped, the freed watcher. Its owners now drop it and say they are deaf: the rt_disk feeds, the master’s watch ofdisks/,C_FSandutils/c/fs_watcher. So does a read canceled by someone other than its owner (an ordering bug: no shutdown does it, the owners stop first), and a read that cannot be armed again; a watcher its owner stopped still goes silently. timeranger2 checks the start of its watchers (a failed start is logged and no watcher is kept), andutils/c/fs_watcherexits with an error when its watcher is gone. (no red test for the read that cannot be armed, the start checks andutils/c/fs_watcher: none of them can be forced)A key-delete reaches only the feeds whose
rkeymatches the key. Lists, iterators and the rt_mem and rt_disk feeds use the same filter as their records; onlykeywas checked, so a feed with anrkeywas told that every key of the topic was deleted.A directory deleted and recreated during an inotify overflow is watched again, and so is the watched root. Its
IN_DELETE_SELFandIN_IGNOREDwere lost with the overflow, the table kept the dead wd, and the rescan took the path as watched; a root reborn so went deaf, recursive or not.IN_IGNOREDnow clears its wd, and the rescan re-adds every watch, the root’s included.No rescan pass for a watcher its owner stopped on overflow.
The keys listing names the call that failed: “lstat() FAILED”, not “stat() FAILED”.
tr2checkresolves the topic path, and never gives a verdict on a partial load. A relativedb/topicor a baretopicopened the wrong directory; a key whose sequences could not all be kept in memory ended its load silently and could still PASS (now exit 2). The open-files limit raises the soft limit up to the hard one, and never the hard limit, root included; it says when it cannot.
MQTT¶
A client’s wrong ack is a WARNING, not an ERROR. A PUBREC for a QoS 1 message (“QoS mismatch”) and a PUBREL for a packet id with no QoS 2 message waiting (“Message not found”, with a stack) were ERRORs naming neither the client nor the peer. Both are WARNINGs with client_id and peername; the broker still answers DISCONNECT 0x82 and PUBCOMP. As a client, a PUBREC for an unknown packet id is only the caller’s WARNING (it was also an ERROR with a stack).
C_PROT_MQTT (deprecated): an UNSUBSCRIBE of several topics removes the topics sent. Each topic was read as a C string inside the packet and ran into the next one. The same gclass stopped logging a WARNING on every connection (an undeclared
connectedattr) and an ERROR for every absent MQTT 5 property, and its logs print the command as a string (twelve used"%d"on aconst char *; no red test for those twelve).
Agent¶
watch-yuno-statstrusts only the routing hop it can, and every requester must opt in. The watch was named from the last hop of the ievent stack, which the client writes, so a forged hop could stop or replace another user’s watch; it is now named from the agent’s own hop, or the control center’s (withcc_connection). A directycommandwatch was sent events it cannot handle (“Event NOT DEFINED” until the watch expired): every requester must now send__relays__: ["EV_YUNO_STATS"](gui_agent 0.29.7 does, directly and through the control center). The first readings follow the answer and go to the renewed watch only. Newmax_watch_ids(256) caps one watch; a cap under 1 refuses every watch and is logged once per bad value.
Control center¶
What only an agent sends is taken only from the agents’ side. The answers of
command-agent/stats-agent, the console PTY (EV_TTY_OPEN/DATA/CLOSE) andEV_YUNO_STATSare public events; a web client could send them too, with a route it wrote, and push answers, console frames or readings to another client’s channel. They are accepted only from__input_side__; one sent by a web client is warned about once a minute (withdropped=), not once per event.An agent’s answer goes only to who asked. When the route named no web client, the answer went to a service named by the client’s own hop: any local service that listens, and never the right one, so a
command-agentmade through the control center’s link to its own agent lost its answers. The requester is now the control center’s own hop: a__top_side__channel, or itsC_IEVENT_CLIlink to its own agent, and only for an agent that link sent a request to. Anything else is dropped with a warning (PTY frames: once a minute). Acommand-agentanswers at once “Command sent to 1 nodes”; the node’s answer comes later (ycommand -w <s>waits for it).command-agentsays a client takesEV_YUNO_STATSonly when the client says so. It wrote__relays__: ["EV_YUNO_STATS"]for every client, so aycommand ... command-agent cmd2agent="watch-yuno-stats ..."was sent readings it logs as “Event NOT DEFINED”. It passes the marker on only when the client’s own__relays__names it.Every answer and stream an agent sends back checks the web client’s connection, not only the stats: a slow answer or a console stream of a client that had left reached the next client of that channel.
ac_tty_mirror_openalso stored the client name from a freed frame.What a capped warning counted is said when its window ends, or when the control center stops (
when=,dropped=), never per connection close, which a client could loop; a count with no later event after it was never logged. The window is the newdrops_warning_window(ms, default 60000) of every capped warning of the control center.Stopping the control center outside the shutdown no longer logs “GObj NOT RUNNING”:
mt_stop()stopped timers thatclear_timeout()ormt_pause()had already stopped (an ERROR with a stack each).The message counters and rates count what it relays again.
rxMsgs/txMsgs,rxMsgsec/txMsgsecand their maxima always read 0: nothing incremented the counters, and the 1 s tick that computed the rates was never armed. They count, once each way, a client’scommand-agent,stats-agent,write-ttyandrun-scenario, each request and run step sent to an agent, and each answer or stream relayed to the client, the run’s answer included. The rate is computed eachtimeout(1000 ms) over the exact interval since the previous tick, so a reading is the same whoever reads and however often; the maxima keep the highest tick, bursts nobody read included;stats=__reset__zeroes them all.Several console mirrors through one agent’s connection: the clients are kept per console; when the agent goes, each is dropped once.
A run’s step is answered only by the agent it went to. A client with
command-agentcould putcc_run/cc_stepin its own command and end or advance another user’s run.A step may not carry a framework key, and every run checks its steps again.
__md_iev__=xin a step replaced the routing of its answer (the run waited forrun_step_timeout),__username__=the stamped user, and a scenario saved by 7.25.14 skipped every step check. gui_agent refuses the same steps, naming the same parameter.save-scenarionamed a scenario it had freed in its answer, when the scenario came as a string.A step an agent does not answer in time is a warning, not an error. (no red test)
emailsender and webstats¶
emailsender: a pause or a shutdown with a message in flight loses nothing and frees nothing twice. The yuno kept a pointer to the queued message the pause had freed, and the close that came after it read it (a crash at
kill-yuno); the smtp session’s held message survived the stop and could be sent twice. A play right after a pause no longer fails on a closing C_TCP (“Initial wrong tcp state”), and it sends what the pause left queued.emailsender: a message the session refuses is resolved once. A message with no valid recipient (
to=","throughEV_SEND_EMAIL) was answered and also returned as a failure, so it was resolved twice, the second time on freed memory. It goes to the failed queue with no retry spent.emailsender: no failure is retried faster than the paced schedule. After any failed session or connection the next connection waits
timeout_retry(2 s), doubled per failure in a row up totimeout_retry_max(10 min): a 4xx refusal, a 4xx or 421 to the end of DATA, a refused default sender, a reply that never comes, a connect or TLS handshake that does not end withintimeout_response, a server that closes by itself, a refused or timed-out connect. An email queued during the wait waits too, and so does a pause and a play. A session that ends with no failure resets the doubling. The C_TCP never connects by itself any more, not even at the start (newconnect_on_start): every connection is the session’s, made when it holds a message, and a yuno with an empty queue logs in to nobody. The doubling counts every failure in a row and is not reset when a message goes to the failed queue (a message refused with a 5xx counts only from the third refusal in a row, see below): with the defaults the first failing message gets 2+4+8 s, the next 16 s and then 32+64+128 s, and from the fourth on 600 s between attempts. Every retry came after a fixed 2 s, and a 4xx at the end of DATA uploaded the message again at once, four times.emailsender: a failing SMTP server is said. The first failure of a streak is a WARNING with its cause (a refused connection included, which C_TCP logs only when traced, so a message could wait with nothing said); a streak longer than the new
timeout_failing_alarm(1 h) is an ERROR, said again at most once per that period; the first delivery ends it with an INFO. A streak exists only while an email waits; one of failed connections and handshakes ends at the next handshake that works (“SMTP server answers again”), any streak at a session closed with no failure, and a url changed (set-url-fromand a pause and play) resets the pacing.emailsender: a stalled connect or TLS handshake is dropped after
timeout_response, said and paced: the email waited for ever and nothing was logged.emailsender: a failure before the mail transaction spends no retry. The transaction starts at the first RCPT TO. A server failing its handshake spent a retry every 2 s, sending the head of the queue to the failed queue in about 8 s, then the next; the message now waits at the head of the queue, paced, for as long as the trouble lasts (a refused default sender spends one retry per attempt). A transient refusal of the login (a
454 4.7.0) is retried: any reply but 235 to AUTH PLAIN was taken as rejected credentials, the yuno exited 0 and was not relaunched. Only a 5xx is a refusal now, and a second334(see the upgrade steps);EV_ON_CLOSEcarries the server’s reply text (reply), and the AUTH buffers are wiped before they are freed (no red test for the wipe).emailsender: a message the server refuses with a 5xx no longer costs the next one a reconnection. A 5xx to RCPT TO, DATA or the end of DATA dropped the session, so the next message logged in again. The message now goes to the failed queue after one attempt and the session goes on (RSET; if RSET is refused, QUIT, and the next connection waits
timeout_retry), for the first two refusals in a row. From the third, with no delivery between, the server is refusing the messages: each refusal still goes to the failed queue, but the next message is paced and “SMTP server refused the last messages in a row: each goes to the failed queue, the next ones are paced” is logged (refused_in_row), never the “emails are NOT being sent” ERROR. Every refusal counts, a bad address included, so a queue refused one by one never gets a login per message; a good email behind such a run waits the paced delay until the next delivery. A 5xx to one recipient of several no longer refuses the message: it goes to the others, each refused one a WARNING, and the “email sent” line says who got it (to/ccthe accepted addresses,refusedthe refused ones, bcc only asbcc_count/refused_bcc_count). A 4xx (rate limit, greylist, 421) is still paced and retried.emailsender: a refused sender is judged by the reply. A MAIL FROM refusal is the message’s (to the failed queue once, with a WARNING naming the
from) only when its sender is its own (not the default, compared ignoring case) and the reply is about the form or existence of that address: a 501, a 5.1.7, or a 5.1.8 / 553 saying the domain or address does not exist (553 5.1.8 <a@b>: Sender address rejected: Domain not found). Everything else is the account’s, whatever thefromand whether or not the reply quotes the address (Postfix quotes it in every sender reject): any 5.7.x (a quota, access denied, not owned), a 5.1.8 that does not say the address does not exist, a 4xx, the default sender. These are paced and charged to the message as a retry per attempt, so at most one email goes to the failed queue permax_retriescycle, and a head stuck pasttimeout_failing_alarmraises the ERROR. Every attempt that ends in a refusal of a message sends it to the failed queue or spends a retry. Any 5xx to MAIL FROM sent every queued email to the failed queue in turn.emailsender: new command
skip-emailmoves the email at the head of the queue, the one every other waits behind, to the failed queue at once (with a WARNING), and answers with itstoand subject; it works while paused:ycommand -c 'command-yuno id=<id> service=emailsender command=skip-email'.emailsender:
timeout_retryandtimeout_responseunder 1000 ms, ortimeout_retry_maxundertimeout_retry, are refused with an ERROR and the default taken (0 gave a tight loop, or a timer never armed).emailsender: a run of messages resolved at once is drained by a loop. A message refused inside its own send (no valid recipient, bad content) sent the next one from inside its own resolution, about 1 KB of stack per message: a persisted batch of some 9000 of them overflowed an 8 MB stack.
emailsender: delivery is documented as at least once. A connection lost after the final “.” and before the 250 (or a 250 later than
timeout_response) is a transaction failure: the message is sent again, so the recipient can get it twice. (Same behaviour as 7.25.20; no red test.)emailsender: an idle session closed by the server no longer logs in again with nothing to send (it did 2 s later, for ever with
timeout_inactivity-1).emailsender: bytes read after the session decided to drop are ignored: a 250 read after the watchdog dropped the session resolved the message as sent, and the next MAIL FROM went on the dying connection. (no red test)
emailsender: an email queued during the SMTP handshake waits for it; it was refused (“Event NOT DEFINED”) and retried in the same instant, spending all its retries before the server had greeted.
emailsender: the send logs name the recipients and the url the message went to. They logged an empty
to, after a refused send bytes from freed memory, and the new url afterset-url-fromwhile the running session still sent to the old one. They giveto,ccandbcc_count(never the bcc addresses).emailsender: a session the SMTP server ends is a WARNING of C_SMTP_SESSION, with the reply capped at 512 bytes, and an over-long reply line is one WARNING, with no “gbuf FULL” ERROR before it. An ERROR of the session is our own failure; the emailsender’s ERRORs (moved to the failed queue, exit on refused login) are what an operator must act on.
emailsender:
set-email-user url=andset-url-fromreach the SMTP session (it started on the old url); a running session takes the new url at its next start, and the answer says so. Every command answer names the yuno.webstats: the addresses a sentence dot, a port,
::ffff:, a colon or a dash left bare are bracketed ([a.b.c.d].,[a.b.c.d]:443,[::ffff:a.b.c.d],client:[a.b.c.d]), which could bring back the OVH “phone number” drop; an IPv6 literal that ends in an IPv4 is bracketed whole ([64:ff9b::a.b.c.d]). Versions and times are left as they are.webstats: yesterday is the calendar day before, not
now - 86400, and a scheduled run reports the day before its slot, not before the moment the timer fired (spring and autumn DST days,report_hour0).webstats: days are counted on the calendar, not as N*86400. The oldest day that
keep_dayskeeps (the prune and the report-day refusal) and the age of awhois_cache_daysanswer werenow - N*86400, one hour off per change of hour in between: near midnight the oldest day kept moved by one day twice a year, and a cached answer expired an hour early or late.webstats: a scheduled slot runs, and is mailed, once. A timer that fired before its slot, with a run that ended before it, armed the same slot again, and on the day of the autumn change a
report_hourin the repeated hour ran twice. A slot is now judged by the day it reports, and each one is built from its date andreport_hour(after the spring change a 02:30 stayed 03:30 the day after). The new statnext_rungives the time of the next run.webstats: a run that keeps the stored day publishes it, not the empty one it read.
gobj-js 7.25.9 and gui_agent 0.29.7¶
gobj-js: the two subscription fixes above (renamed subscriptions, the subscription given), renames as in C (
__original_event_name__, the renamed event is published), traces and logs mask credentials as C does (is_secret_name,json_mask_secrets,mask_secrets_inline,trace_json_maskedinhelpers.js; only json is walked: a gobj, a widget, a DOM node or a class instance is passed through, an object met twice is masked everywhere, a cycle is"<cycle>", and the masking never throws), andsubs_flagcarries the C bits.gobj_unsubscribe_list()of a subscription already removed gives one warning per call (“Subscription(s) already removed, nothing to remove”). Publishing costs the same as before (measured: 701 against 703 ns per delivery).gui_agent: a watch of the yuno stats says it takes
EV_YUNO_STATS, directly and through a control center, and a scenario step may not carry a framework key.
Tooling and house rules¶
linux-ext-libs 1.23: openresty from its release tarball.
extrae.shdownloads it from openresty.org (sha256 pinned inrepos2clone.shasSHA256_OPENRESTY) instead of cloning the git repo. Building from the git tag ranutil/mirror-tarballs, which fetches about 45 modules as tarballs from github.com, one download per module. The release tarball is the output of that step, with the same modules at the same versions, configured with the same flags, so the binary is built from the same sources.A SessionStart hook prepares a cloud container for the C suite and the JS packages (
.claude/hooks/session-start.sh). It does nothing outside a cloud session.No two test binaries share a port, and the c_mqtt tests have a work dir of their own run. Ten ports were shared (7778 by the c_tcp/c_tcps families, 18801/18802 by the c_auth_bff binaries, 3333 by the yev_events tests, among others), so
ctest -jfailed with “bind() FAILED”; the newscripts/check_test_ports.pyexits 1 when two binaries share one. Each c_mqtt test wiped a fixed/tmp/test_mqtt_<name>and left it behind; each run now makes its own and removes it.yunetas CLI 0.20.3:
sync-binaries,sync-configsandset-start-prioritiesrefuse anaccess_tokenthat is not a string, a token answer or an OIDC discovery document that is not a json object, and a discovery that fails, with a clear error and exit 2 (Python raisedTypeError,AttributeErroror a traceback).search_process()(behind--stop) allocates throughgbmem_*and no longer leaks the/proc/*/commglob when an allocation fails. The executable-name check ofentry_point.ctakesgbmem_strdup()too. Both pairs were balanced (no heap mixing), but they bypassed the allocator’s limits andCONFIG_DEBUG_TRACK_MEMORY. (no red test)Braces on every body in
c_authz.c,c_uart.c,c_websocket.c,run_command.c,c_yuno.c,entry_point.candydaemon.c, and// Error already loggedafter thegclass_create()ofC_ASSETS,C_AUTH_BFF,C_PROT_RAWandC_PTY. They compile to the same machine code as before. (no red test)About 45 comments and test READMEs said what the code did “up to this fix” or “before this fix” without saying which release; each now names it.
Known limitations¶
An rt_disk follower lagging behind a master in another process can still hand a key deleted and written again out of order (present since 7.25.20, narrowed by this release, not ended): it reads a key directory’s links by path, and on Debian’s 6.x kernels with ext4 a directory removed and made again can keep the same inode and birth time, so the identity check cannot tell them apart.
timeranger2/test_delete_key_propagation(“race in the batch”) fails there for that reason and passes on newer kernels. The fix (reading through a descriptor of the watched directory) is planned; seeTODO.md. Production masters feed their own lists from memory (rt_mem) and are not affected.A frame before the session is parsed whole, up to the maximum block (also in 7.25.20): an identity card is small, but nothing caps a frame sent before it. The masking of that frame is bounded since this release.
The rest of the open items are in
TODO.md, “Found by the last review before the merge” and “Defects open after 7.25.20”.
v7.25.20 (2026-09-30)¶
The last hole of the secret masking of 7.25.19, and the performance study of the fifteen releases since the last one (7.25.5): the report of this release measures 7.25.20 against 7.25.5.
Kernel: list-persistent-attrs masks the secrets too¶
7.25.19 masked
SDF_SECRETattrs inview-attrsand the rest, butlist-persistent-attrs(and the answer ofremove-persistent-attrs) still showed them in clear: the mask was applied one level above the{<gobj>: {attrs}}it answers.gobj_list_persistent_attrs()masks each gobj’s attrs itself now; every caller of it shows what it answers.
Performance, against 7.25.5¶
The report of this release measures it against 7.25.5, the release of the last report:
performance/reports/7.25.20.html, and the trend on doc.yuneta .io /performance. Both releases built whole, run alternately, 8 rounds (24 for the timeranger2 tests). Faster, all from the
SWITCHSfix of 7.25.7 (it compiled a regex on every entry): the open of 40 treedbs with an unchanged schema 0.667 -> 0.318 s (x2.1), a saved treedb update -25.2%, link/unlink -20.3%, the delete of a parent of 200 children -17.3%, an update in memory -16.2%, the first open of 40 treedbs -8.1%, a create -7.0%, appends with a live reader +12.8% (the benchmark’s callback usesSWITCHS).Nothing slower beyond its noise.
tm_build_appends+3.6% (at the edge of the spread) runs the same append code in both releases: placement.test_topic_pkey_integer+0.3%.make_report.pytakes the words of a report from its json (story), so a report is no longer written for one release.
v7.25.19 (2026-09-30)¶
The end of the lost webstats report of wattyzer, and what it turned up:
the report’s IP addresses as [a.b.c.d] (OVH read one as a phone number and
dropped the mail), secrets masked where attrs are shown, the persistent-attrs
file 0600, and the SMTP refusal text in the log. A kernel change, so the
suite ran on both machines (215/215 each); deployed as emailsender and
webstats on every node -- the other yunos take the masking at their next
build.
Kernel: secrets are not shown, and not left world-readable¶
SDF_SECRET, a new attr flag: the attr is read, written and persisted as before, and SHOWN as********byview-attrs,write-attr(its answer echoed the value written),list-persistent-attrs,view-gobjand the start-up trace of the yuno’s attrs -- through the newgobj_mask_secret_attrs(). Marked: the passwords of the emailsender and its SMTP session,client_secret(auth_bff),sign_secret(assets),kc_admin_client_secret(keycloak), the MQTT password, and the user passwords and tokens ofc_task_authenticate,c_ievent_cli,c_tcp, mqtt2 and the CLI tools.view-attrsshowed the relay’s password in clear to whoever asked.The persistent-attrs file is written 0600 (
dbsimple.c), and fchmod’ed so the files written before are closed at their next save: it wasjson_dump_file()with the process umask -- 0666 on every node -- and it holds those secrets in clear.
emailsender: a refused login says why¶
The SMTP server’s reply TEXT is logged with its code when the login is refused.
535alone could not tell a wrong password from what it was on 2026-09-30:535 5.7.1 Authentication failed, OVH blocking the account.
webstats: the report’s IP addresses go out as [a.b.c.d]¶
OVH’s outbound relay read
34.140.132.132as a Spanish phone number (+34 and nine digits) and, in a report full of addresses marked “banned”, accepted the mail (250 queued) and delivered it to NOBODY: no bounce, not in Junk, not at gmail or outlook either. wattyzer’s report of 2026-09-29 was the first with such an address, bisected down to one row of Top clients; Google Cloud addresses start with 34, so it recurs. The mail (andpreview-report) now writes every standalone IPv4 as[a.b.c.d]-- proven on the relay; the usual defanga[.]b[.]c[.]ddoes NOT pass. The stored record keeps the plain address.
v7.25.18 (2026-09-30)¶
emailsender and webstats only: no kernel change. Found chasing a webstats
report of wattyzer that stopped arriving (OVH accepts it -- 250 queued --
and it is lost after the relay; the cause is still being narrowed down, see
TODO). Each node takes it with install-binary of the two roles and the
usual promotion.
emailsender: the smtp trace no longer writes the credentials¶
C_SMTP_SESSION’ssmtptrace wrote every command line,AUTH PLAINincluded: the base64 of the user and password of the relay, that is the password in clear, in the yuno’s log and from there in the logcenter’s. It names the mechanism now (>>> AUTH PLAIN <credentials not traced>). Found tracing a report that did not arrive; the two lines that had been written on wattyzer were overwritten in place.
webstats: a rebuilt day with no access log keeps (and sends) the stored report¶
The guard against storing an empty rebuild over a good day measured the lines KEPT from every source; the fail2ban and error logs of a day outlive its access log, so rebuilding such a day (
report-day ... send=1after the rotation) passed the guard and stored -- and mailed --NO DATAover a day of 4632 requests. It measures what the report saw now (requests + errors, the NO DATA of the subject), and a rebuild asked to send mails the STORED report.load_reportanswers the newest record WITH data when a later one is empty (the store is append-only: the good day was never lost), which also mends the days already written that way.
webstats: send-yesterday¶
Builds the report of yesterday and mails it, whatever
send_emailsays, with nothing to type --report-daywants the date andsend=1.analyze-nowsays in its help that it mails only whensend_emailis on.
v7.25.17 (2026-09-30)¶
Kernel: C_TCP_S and C_UDP_S answer help¶
They were the only gclasses with commands (
reload-certs,view-cert) and nohelp:command-agent service=agent_secure_port command=helpanswered “command not available”, and the commands could only be found in the code. Every gclass with a command table in the SDK and in the projects now has one (the two command-parser test fixtures excepted). Deployed with the agents and the control center; every other yuno gets it at its next build.view-certsays in its description that it answers the listeningurltoo. The transport page documented ahelpthat did not exist and aview-servicesthat does not; it lists the three real commands now.
JS: gui_agent 0.29.5¶
“For TreeDB” gives each agent its own domain. The export built an agent’s url from the node’s name (
wss://artgins:1993), because the scan read the certificate fromview-config, where the agent’s is a global override it does not see. It asks the agent’s secure gateview-certinstead: the loaded certificate’s CN (wss://agent.artgins.com:1993) and the gate’s port. A CN that is not a hostname -- the package’s own self-signedyuneta_agent.yuneta.io-- falls back toview-configand the node’s name, as before.
v7.25.16 (2026-09-30)¶
A lite release (control center and agent only): the two points the review of 7.25.15 left in TODO.
Control center: two editors of one scenario¶
save-scenariotakes therevisionthe scenario was read at and refuses the save when it was saved (or deleted) since, naming who -- instead of the last writer silently winning. The revision is the record’sg_rowid, whichscenariosandsave-scenarionow answer in__md_treedb__(updated_atcounts seconds: two saves in one second look the same). Withoutrevision, or 0, nothing is checked: a new scenario, or an overwrite on purpose.
Agent: a yuno that runs but is not connected is answered for¶
command-yuno,stats-yunoandauthzs-yunosent nothing, and answered nothing, for a matched yuno whose row has no channel (starting, or going): the requester waited for ever, a control center running a scenario for its step’s whole deadline, with every other run held behind it. Now the command is answered at once with an error when it reached none of the yunos, and logged when it reached only some.
JS: gui_agent 0.29.4¶
The live view keeps the revision of the scenario it shows (from the list and from each save) and sends it with a save of the same scenario; a refusal says somebody saved it after it was opened (
scenario changed since read). A save under another name, confirmed first, is an overwrite and sends none.
v7.25.15 (2026-09-30)¶
A lite release (control center and agent only, rule of 2026-09-29): only
those two binaries change, and only they are deployed. It closes what a
review of the scenarios and their live view found, after running the
yunovatios-stress scenario end to end from the console.
Control center: a web client is its connection; runs end when their agent goes¶
A stats reading for a closed tab could reach another user. The relay of
EV_YUNO_STATSfound the web client by the NAME of its channel in__top_side__, and that name is taken by the next client once the first closes -- while the agent’s watch goes on pushing until it expires (watch_ttl, 60 s). Each connection now gets a number when it opens,command-agentstamps it on what it forwards (cc_connection, kept by the agent with the watch), and a reading -- like the answer of arun-scenario-- is delivered only while the channel holds the same connection.Use after free in that relay: the name of the channel was read from the popped stack frame after it was freed, in the warning for a gone client.
A run whose node’s agent disconnects ends at once, not at the step’s deadline (
run_step_timeout, 30 s, during which every other run waited).Pausing the control center in the middle of a run answers its requester: the run was ended after the web side had been stopped, so the answer was always dropped.
save-scenariorefuses whatcommand-yunowould misread. A step parameter named like one ofcommand-yuno’s (id,service,command) or like a column of the agent’syunostopic (date,yuno_name, ...) became the filter that selects the yuno -- another yuno, or “Yuno not found”. A service and a yuno id are letters, digits and_ . ^ -; a scenario id is letters, digits and_ . @ -, not starting with a dot (timeranger2 refuses it as a key), 200 at most (a run id that would not fit is refused, not truncated).Documented: the steps run on the control center’s session, so
write-scenariostogether withrun-scenariosis as much ascommand-agent.
Agent: watch-yuno-stats tells two tabs apart, and takes what it can¶
The requester of a watch is told by its first hop AND the channel it came in by at the control center (
input_channel): two tabs of one browser were one requester, the second replacing the first watch at each renewal, and a hidden tab stopping the other’s.One yuno can be watched through several services (
ids=5120,5120:db); the last one named used to replace the others.An id that is not a yuno of the agent no longer refuses the whole watch: it is watched, its
statesaysmissing, and the answer names it (data.missing).
JS: gui_agent 0.29.3¶
The live view stops the agents’ watches on Disconnect, when it is stopped and when another scenario is shown (behind the control center a watch pushed on until its ttl); readings asked for, or pushed by the watch of, the scenario shown before are dropped (
monitor_gen) instead of landing on a card with the same key.One control at a time, with its number on every request; a restart goes in phases (stop, resets, start), each waiting for all the answers of the one before; a control not answered in time or cut by a drop says so.
A watch the control center could not dispatch is asked again at the next renewal instead of turning that node to polling for good.
Restart only when the scenario has a stop; the editor refuses the same step parameters as the control center; units of the selectors translated; the list opens a row while it re-reads and re-reads after a change that landed during a read; the direct link tears itself down outside the publish that refused it and lets go of the login’s refreshes when stopped.
JS: tabulator-tables ^6.6.0 in every SPA; gui_agent 0.29.2, gui_treedb 0.17.73¶
6.6.0 is additive for these apps (an opt-in range fill handle, an
initialValuefor editors, a Bootstrap 5 border fix). Every SPA of the ecosystem raised its floor, was rebuilt, deployed and checked live; the gobj-ui lockfile and its test-app follow (the library’s^6.5.3range already takes it, so no publish).
JS: gui_agent 0.29.1¶
The scenarios list reads itself again whenever its tab is shown, so the runs it counts are current (they only moved on a save or a delete).
v7.25.14 (2026-09-29)¶
A lite release (controlcenter only, rule of 2026-09-29): only the control center’s binary changes, and only it is deployed.
Control center: the scenarios, in its treedb; its dead topics gone¶
treedb_controlcenterkeeps the scenarios (schema_version3,treedb_schema_controlcenter.c): topicscenarios-- a set of yunos on one or several nodes, how the messages flow between them (links) and the commands of each action (start,pause,resume,stop,report), a document saved whole -- and topicscenario_runs, each run linked to its scenario. New commandsscenarios,save-scenario(validated: the same rules as the console’s editor),delete-scenario(with its runs),run-scenario scenario_id= action=andscenario-runs, with the permissionsread-scenarios,write-scenariosandrun-scenarios, checked by the commands always. The parameter isscenario_id, neverid: through an agent’scommand-yunotheidof the kw is the yuno’s.run-scenarioruns an action on the nodes, one step after another, each ascommand-yuno id=<yuno> [service=] command=<command>to its node’s agent, as the user who asked. A step that fails or does not answer inrun_step_timeout(30000 ms) ends the run; the run is written toscenario_runs(every step with its answer, and for areportwhat each step answered) and the command answers when it is over. One run at a time.Removed, unused since the webix GUI: the topics
systems,nodes,services(declared inventory, viewer launcher) andusers(a copy of each user at login, read by nobody: the users are theauthzstore’s), thelists/viewer_enginesdraft,ac_treedb_node_*(they acted only on atreedb_purezadbof another project), the attrsenabled_new_usersandenabled_new_devices, and the authz entrieslist-groups,list-tracks,realtime-track. On a.com the three inventory topics were empty.A login is no longer refused while the control center boots.
ac_user_login()answeredEV_AUTHZ_USER_LOGIN-- a veto point -- with -1 until its treedb opened, only because it could not copy the user yet; that refusal went with theuserstopic, and so did the “Treedb Controlcenter not ready” warnings a restart logged while the agents reconnected (YUNO_AUTH.md§4.9).C_AUTHZ:EV_AUTHZ_USER_LOGIN/LOGOUT/NEWareEVF_NO_WARN_SUBS. Their subscribers are optional (the agent, the mqtt broker); a yuno that has none logged “Publish event WITHOUT subscribers” on every login -- the control center did, from the moment it stopped subscribing (seen deploying this release on a.com: 29 in the first minute). Reaches each yuno at its next rebuild.
JS: gui_agent 0.29.0, the Scenarios workspace, in place of Monitor and Statistics¶
The console’s side of the control center’s scenarios: Scenarios (the list the control center keeps), Yunos (ticking yunos watches them as cards of every counter -- what the Statistics workspace was) and Live (the graph and charts, or the cards). Actions replace the old test block; a saved scenario is run BY the control center (
run-scenario) and its runs are listed; any other one runs from the console. Against a control center older than 7.25.14 it keeps the scenario in the browser, as before.The Statistics workspace is gone (it is a scenario of cards now), and the rail has five workspaces again.
JS: gui_agent 0.28.0, the Users workspace¶
The users each yuno lets in, and their roles, from the console. Every yuno that authenticates keeps its own store (each node’s agent, each control center: 1996 and 1997 are two), so Users is a per-yuno workspace: the nodes->yunos tree marks the yunos that keep users (a
C_AUTHZwith itstreedb_authzs) and whether each is the master or a read-only replica, and a tab lists the users of one store and creates, enables, disables, deletes and gives or takes roles. Every write is aC_AUTHZcommand over the agent; a role is alink-nodes/unlink-nodesontreedb_authzs, becauseupdate-user role=autolinks and REPLACES every role the user holds by that one.It makes the open per-command authz gate visible:
create-user& co. run for any operator the control center lets runcommand-agent, even with no role in that yuno’s store, whilelink-nodeschecksupdateon its own (TODO.md, “per-command authz gate”).First step of the control-center treedb plan (
TODO.md, “Controlcenter: scenarios in its treedb”).
v7.25.13 (2026-09-29)¶
Agent: watch-yuno-stats, the stats of yunos pushed to whoever watches¶
New agent command
watch-yuno-stats ids=<id>[:<service>],... period=<ms>(stop=1ends it): the agent keeps the requester’s route -- the wayopen-consoledoes -- and SENDS itEV_YUNO_STATSevery period, per yuno astate(running, playing, disabled), acpu(its__yuno__stats) and anapp(its service’s stats). The readings are taken on loopback once per yuno and service however many watch it, at the shortest period asked;watch_min_period(1000 ms) andmax_watches(30) bound it. The watch ends with its requester’s connection. A subscription could not do this: they go from a client to a server only, and the agent is the server of its yunos.New kernel event
EV_YUNO_STATS(g_ev_kernel).One watch per requester = its connection plus the client at the far end of the route (every web client of a control center shares its connection), and it expires after
watch_ttl(60000 ms) unless asked again: behind a control center the agent never sees a browser leave. The answer carries thettl.The control center relays
EV_YUNO_STATSto the web client, like the PTY mirror.command-agentnow adds__relays__: ["EV_YUNO_STATS"]to the kw it forwards (after removing one the client may have sent), and the agent refuses a watch through a control center without it: a control center that gets an event it does not know drops the agent’s connection. A reading for a web client that is gone is counted and said once a minute.gui_agent’s Monitor (0.26.0-0.27.0) uses it directly and through the control center, per node, and polls a node whose agent or control center refuses the watch.
ac_timeout_periodicofC_AGENTnow tells its periodic timers apart bysrc(cert-sync and the watch), and logs a tick of an unknown one.
JS: gui_agent 0.23.0 - 0.27.0, the Monitor workspace¶
A test, live, from its agent. gui_agent gains a fifth workspace, Monitor: a left-to-right graph of the yunos of one scenario (the agent url, the yunos, the links the messages follow) with cpu %, msg/s in and out, queue and state per yuno, and two uPlot charts over a 5-60 minute window. First written for yunovatios’ stress tests (sim_controllers -> gate_central -> db_tracks_ce), and it matches what
ycommandmeasures on them.It is the first transport of the console that does not go through the control center: it talks to the agent’s
wss://<node>:1993with the access_token it gets from the BFF’s/auth/token, so that plane’s BFF needsexpose_access_token(YUNO_AUTH.md§2.2). The agent must serve a certificate a browser trusts; the self-signedyuneta_agent.yuneta.iodoes not.A test gets controls (0.24.0): a
testblock in the scenario declares Start / Pause / Resume / Stop as lists of commands of the generator (sent ascommand-yuno), and Restart is stop +stats-yuno stats=__reset__to every yuno + start. Confirmed in a dialog that lists the commands; a production scenario (notest) shows none. A service that keeps its counters in private fields behindmt_readingdoes not honour__reset__unless itsmt_statsdoes it (the yunovatios yunos do not yet).Through the control center, several nodes (0.25.0): a scenario that says
nodeinstead ofagent_urlsends every command ascommand-agent agent_id=<node>on the console’s own link — no token leaves the BFF and the nodes need no browser-trusted certificate. Propose links reads the yunos’view-configand proposes the links from what they listen on and connect to.The readings are polled (the Statistics exception, extended to this view) until the agent can publish stats to a subscriber (design in
TODO.md, “Stats pushed to a subscriber”);TODO.md“Stats: three things a live monitor cannot trust” lists what a monitor finds in the SDK stats today.
tr2check: check a topic filled by a load test¶
New tool
utils/c/tr2check(installed in/yuneta/bin, shipped in the.deb/.rpm). In one pass per key it gives what an acceptance test asks in SQL: record count, duplicates, gaps and missing records (--expected), out-of-range sequences, checksum mismatches (--checksum-field, sha256 of the record as compact json with sorted keys), the storage rate and the latency percentiles (__t__ - __tm__, exact to the ms on asf_t_ms|sf_tm_mstopic). The sequence counts per key (--seq-scope=key, the way a device numbers its frames) or for the whole topic; a record that arrives out of order is not a duplicate. One json document on stdout; exit code 0 pass, 1 a check failed, 2 the topic could not be checked. Test:timeranger2/test_tr2check.
CLI 0.20.2: sync-binaries uploads the file it compared¶
With
--yunos-dir, the table was built from the staged binary but the upload sent$$(<role>), whichycommandresolves inoutputs/yunos: on a node where two projects build the same role (hidraulia’sgate_caudal), the row saidREBUILD 1.9.0.0and the upload was the other project’s 1.6.1.0. The command now carries the file’s path. The wait for a yuno to stop before a same-versionupdate-binarygoes from 15 to 60 s.
CLI 0.20.1: yunetas init keeps the ctest logs¶
yunetas testleaves onebuild/<timestamp>.txtper ctest run, the history a release compares timings against, andinitrecreated the rootbuild/withrm -rf: the clean rebuild of a version bump wiped it (7.25.12 lost it on the dev machine and on wattyzer).initnow carries those logs across.
v7.25.12 (2026-09-28)¶
timeranger2 (fs_watcher): a pass says where its time went¶
The INFO that closes a pass after an inotify overflow adds
slices,ms_owner,ms_watcher,ms_loopandmax_loop_ms: the owner’s callbacks, the walk, and the loop’s own work between slices. On yunovatios’ central some passes took 251 and 680 s while the yuno digested a queue; measured with it, such passes are 65-78 % owner (handing over the backlog of records), 22-35 % loop and 1.1-1.7 s walk: not the watcher.
emailsender: rejected SMTP credentials stop the yuno, they are not retried¶
When the SMTP server refuses the
AUTH PLAIN, the emailsender logs one ERROR (“SMTP credentials rejected: exiting, NOT relaunched”, with the reply code, the url and the username) and exits with code 0 (LOG_OPT_EXIT_ZERO), so neither the watcher nor the agent relaunches it. It used to treat the refusal as a link drop and retry every few seconds: 768 refused logins in 54 minutes on hidraulia on 2026-09-28, the pattern that got an address banned from the whole OVH mail cluster in July. The in-flight message stays at the head ofemails_queuewithout spending a retry.C_SMTP_SESSIONreports the refusal onEV_ON_CLOSEasauth_rejected, apart from the per-messagecode.
CLI 0.20.0: secret overlays removed (BREAKING)¶
yunetas list-secretsandsync-configs --secrets-dirare gone, andsync/sync-configs --nodeno longer read~/.yuneta/secrets/<node>/. A credential a yuno needs goes in its config file again and is pushed as written. The"__SECRET__"placeholder was filled only bysync-configs; a node reinstall script registered the placeholder itself, the emailsender logged in with it and the mail provider banned the node. A config that still carries"__SECRET__"is now pushed with it: put the value in the file first.
emailsender: a blank password leaves the SMTP side stopped until set-email-user¶
Batch configs leave
username/passwordblank; the operator sets them once withset-email-user(persistent attrs, loaded over the config at every start). With either one blank the yuno now runs, queues emails and logs one ERROR, but does not start the SMTP session. It used to fail twice over: both attrs wereSDF_REQUIRED, so the service refused to start andset-email-userhad nothing to talk to; and a session started with no password skipped the AUTH, so the server refused every message into the dead-letter queue.set-email-usernow applies at once: it hands the credentials to the SMTP session and starts it, and the queue is sent, with no restart.The cached
username/password/url/frompointers are refreshed on every write (mt_writing);set-email-user ... url=...lefturlpointing at freed memory, which the send logs read.
Packaging: colas2.sh works again when called without arguments¶
The
/yuneta/bin/colas2.shthat the.deband the.rpminstall passed its forwarded arguments as"${FWD_ARGS[@]:-}", which, with none, is ONE empty argument, not zero.list_queue_msgs2takes exactly one path, so every queue answered “Usage: list_queue_msgs2 [OPTION...] PATH”: the script failed in its plainest use, on every node, since the.debversion of 2025-09-04. It now expands${FWD_ARGS[@]+"${FWD_ARGS[@]}"}, which is nothing when there is nothing and safe underset -u.
v7.25.11 (2026-09-27)¶
Found by the third gate-outage test of yunovatios’ central: three “Event NOT DEFINED in state” that a connection going down produced, all in the window between a transport deciding to close and the layers above learning it; and a pass after an inotify overflow whose own cost grew with the tree. Rebuild every yuno (TCP and websocket clients, rt_disk followers) and the agent.
root-linux: data that arrives or leaves while a connection closes¶
C_WEBSOCKET, receiving while closing:
ws_close()moves the gobj toST_DISCONNECTEDand sends its Close frame; a client keeps the connection until the server drops it ortimeout_closeruns out, as RFC 6455 asks, and a server’s pending read still completes. What arrives then (the rest of a frame given up on, the peer’s Close) came asEV_RX_DATA, which the state did not declare. It is now discarded, nothing published (debugtrace level). Seen on a sim of controllers too busy to read a frame withintimeout_payload. Test:tests/c/c_websocket(new).C_TCP, sending while closing: a drop or a disconnection waits in
ST_WAIT_STOPPEDfor its last io_uring operation before it publishesEV_DISCONNECTED, and until then the layers above send. ThatEV_TX_DATAwas undeclared there -- ten thousand ERRORs in an hour on a sim of 300 controllers whose central went away. It now goes where the pending queue of a dead connection goes (away;connectionstrace level). Test:c_tcp/test7.Documented in
api/gclass/protocol.md(C_WEBSOCKET: the states corrected,timeout_close, a Closing section) andapi/gclass/transport.md.
timeranger2 (fs_watcher): the pass after an overflow, at a cost that does not grow with the tree¶
7.25.10’s sliced pass rebuilt, in EVERY slice of 20 ms, the index of the watched paths by path (the watch table is indexed by
wd): 50000 entries per slice. The watcher’s own cost per directory grew with the tree -- 254 us at 69632 directories -- and on yunovatios’ central a pass over 50501 key directories took 7 minutes, and the next one more than 20. The index is now built once per pass (and grows with the watches the pass adds): 73 us per directory, a pass of 16 s instead of 29 in the test.test_fs_watcher_overflownow measures that cost (the pass less the owner’s time) and fails above 200 us per directory.
yuno_agent, controlcenter: an answer for a client that left¶
A command or stats answer, or a counted one (
kill-yuno,run-yuno, ...), for a client that left before it (aycommandthat timed out) was sent to its channel, CLOSED by then: “Event NOT DEFINED in state” with a stack. The agent and the controlcenter now check that the requester is still listening (its channelopened, or a link in session) and drop the answer with an INFO “requester gone before its answer, answer dropped”.
v7.25.10 (2026-09-27)¶
The recovery from an inotify overflow (7.25.9) now keeps the yuno answering:
the pass over the watched tree runs in slices of 20 ms per loop turn. Found by
the second gate-outage test of yunovatios’ central, where one pass held a
db_history_ce deaf for four minutes. Every yuno that reads a timeranger2
topic of another yuno must be rebuilt against this SDK.
timeranger2 (fs_watcher): the pass after an overflow gives the loop back¶
7.25.9 recovered from an inotify overflow in one piece: the owner rescanned the whole tree inside the single
FS_OVERFLOW_TYPEcall. A timeranger2 follower of 50000 keys on the busy disk of yunovatios’ central took 56 s, then 243 s, and the yuno answered nothing meanwhile -- its agent, its commands, its timers.Now the watcher runs the pass itself, a slice of 20 ms per loop turn: every directory of the tree goes to the owner as the new
FS_RESCAN_DIR_TYPE, and a directory born in the overflow is watched first (one walk does both: the separate walk that re-set the watches is gone).FS_OVERFLOW_TYPEstays, for what is global and cheap. Another overflow during a pass adds one more whole pass after it. INFO “watched tree rescanned after lost inotify events”.A follower’s rt_disk feed finds the keys deleted while the events were lost by reading
keys/once (it was astat()per key), and reads one key directory perFS_RESCAN_DIR_TYPE. The master andC_FSignore the directories; thefs_watcherCLI prints them.Test:
timeranger2/test_fs_watcher_overflow(a real overflow, an owner slow on purpose and a timer probing the loop: every directory told, a pass of ~30 s, the loop deaf at most 50 ms).test_rt_disk_overflowstill covers the follower. Documented inapi/timeranger2/fs_watcher.md.
v7.25.9 (2026-09-27)¶
An inotify queue overflow no longer aborts the yuno: the watcher recovers in
place and its owner rebuilds its view from the filesystem, where what the lost
events said still is. Found by yunovatios’ stress test of its central, where a
db_history_ce following a busy db_tracks_ce lived in a crash loop under a
burst. Every yuno that reads a timeranger2 topic of another yuno (an rt_disk
feed) must be rebuilt against this SDK.
timeranger2 (fs_watcher): an inotify overflow no longer aborts the yuno¶
An
IN_Q_OVERFLOWof a watcher aborted the yuno to be relaunched and reload clean. Under a sustained burst the reload met the next overflow: in yunovatios’ stress test of its central, adb_history_cefollowing adb_tracks_ceat ~3000 records/s over ~60000 keys aborted five times in two hours, and each relaunch spent 2-3 minutes catching up before falling again.Now the watcher recovers in place: a WARNING, a watch set on every directory of a recursive tree that was not watched, and the owner called once with the new
FS_OVERFLOW_TYPEto rebuild its view from the filesystem, where what the lost events said still is.A follower’s rt_disk feed scans every key directory of its
disks/<rt_id>/(a link still there is a record not handed over yet) and compares its cache with the topic’skeys/(a key gone is heard deleted, itskey_deletedcallback fires). Idempotent: nothing is handed over twice. INFO “rt_disk feed rescanned after lost inotify events”.The master’s watch of
disks/closes the feeds whose directory went and opens one for each directory without it (find_rt_disk()skips anrt_idalready fed).C_FSpublishesEV_FS_CHANGEDfor the root;utils/c/fs_watcherprints it, and no longer reads the event types as bits (a “file created” printed as “directory created” and “directory deleted” too).Test:
timeranger2/test_rt_disk_overflow(a real overflow: 69632 new keys with the loop stopped, oneIN_CREATEeach, and a key deleted while the queue is full; every record once, the delete heard once, a key born in the overflow watched after; with the old code it aborts). Documented inapi/timeranger2/fs_watcher.md.
v7.25.8 (2026-09-26)¶
A TLS fix in C_TCP: under a burst, a connection could have two writes in
flight, and a short one then sent its rest out of order -- the peer answered
“bad record mac” and dropped the link. Found by yunovatios’ stress test of its
central, where the drain after a gate outage barely converged because of it.
Every yuno that speaks TLS must be rebuilt against this SDK (the gclass is
linked into each binary).
TCP (C_TCP): TLS “bad record mac” under a burst -- one write in flight¶
A TLS connection could have two writes in flight, and a short one then sent its rest after the other: the byte stream went out of order and the peer failed the record MAC (“SSL_read() FAILED”,
error:0A000119, “decryption failed or bad record mac”), dropping the link.ytlshands its encrypted output over from several places (the message being written, the handshake, the records after it) andac_send_encrypted_datastarted a write for each at once; and the completion of any write took the next message while another was still in flight. It shows under a burst on a full socket -- clients resending their windows when a gate comes back -- and how often depends on the kernel: yunovatios’ stress test of its central (Rocky 9, kernel 5.14) dropped links 13606 times in 25 minutes, and the drain after a 4-minute outage barely converged.Now a connection has ONE write in flight, TLS included: encrypted output waits in a queue while a write is in flight, it goes before the next clear message, and the queue is discarded with the TLS session it belongs to. The plain path already kept the rule.
New stat
max_tx_in_progressofC_TCP: the most writes ever in flight at once,1by the rule.Test:
c_tcps/test5(10 bursts of 400 x 64 KB over TLS to an echo server, every echo checked in order, andmax_tx_in_progress == 1asserted: with the old code it is2on every run, while the echoes alone came back intact in most runs). Documented inapi/gclass/transport.md(One write in flight).ac_send_encrypted_dataanswers 0 when its write cannot start (logged, the connection dropped): the mbedTLS backend frees the gbuffer again on a negative answer, which the kw had already released.
v7.25.7 (2026-09-26)¶
A leak fix and nothing else in C: every return from inside a SWITCHS case
lost a compiled regex, and every gate with an output queue returned from one per
message. Found by yunovatios’ stress test of its central. The JS submodules move
to maplibre-gl 6.11.2 and vite 8.3.1. Every yuno must be rebuilt against this
SDK to lose the leak.
Helpers: a return inside a SWITCHS case leaked a compiled regex¶
SWITCHScompiled a regex (".*") on entry and onlySWITCHS_ENDfreed it, so everyreturn(orgoto) out of aCASESbody lost glibc’s compiled automaton, a few KB per switch, allocated byregcomp(3)OUTSIDE gbmem: invisible to the gbmem audit and tocur_system_memory. 29 of the 63SWITCHSof the tree return from a case, several once per message:C_MQIOGATE’sac_send_message(every gate with an output queue lost ~1.1 KB per message; agate_centralof yunovatios’ stress test reached 1 GB in 25 minutes),c_prot_modbus_m,tr_treedb,tr_msg2db,yev_loopand the frame decoders of the projects. Found by a gdb breakpoint onmalloc(gbmem allocates withcalloc, somalloccatches only what bypasses it).Now
SWITCHSallocates nothing andCASES_REmatches through the newstr_match_regex(), which compiles, matches and frees on the spot: leaving a switch from anywhere is safe, and no caller changed. The usage is the same.Every yuno that uses
SWITCHSmust be rebuilt to lose the leak: the macros are expanded where they are used, so an old binary keeps it.Test:
helpers/test_switchs(100000 x 3 returns from inside cases: +1.06 GB of libc heap with the old macro, +1 KB with the new one; matching unchanged). Documented inapi/helpers/string_helper.md(str_match_regex()and the switch).
JS: gobj-ui 7.25.23, gui_agent 0.22.99, gui_treedb 0.17.72¶
maplibre-gl 6.11.2 and vite 8.3.1. gobj-ui 7.25.23 raises its peer floor to
maplibre-gl ^6.11.2and builds with vite^8.3.1; no API moved. gobj-js builds with vite^8.3.1(7.25.8 is not republished). gui_agent and gui_treedb take gobj-ui^7.25.23and vite^8.3.1, and are deployed. maplibre 6.11.2 changes the glyph protocol between its worker and the page: a host that emits the worker keeps the version in the file name (maplibre-gl-worker-6.11.2.js), as the gobj-ui test-app does, or a browser that cached an older worker loses the map labels.
v7.25.6 (2026-09-25)¶
What changed after 7.25.5: the tests that failed on the nodes, one of them over a real defect of the static resolver, and the open-files limit of the agents. No change in timeranger2, treedb or the transports.
Event loop (yev_loop)¶
The static resolver (
CONFIG_FULLY_STATIC) sent a numeric address of the other family to DNS:getaddrinfo("::1")asked forAF_INET(asrc_urlof[::1]bound for the IPv4 address of a destination) fell through to an A query, and a slow nameserver stalled the event loop (3.1 s on wattyzer, WARNING “getaddrinfo() BLOCKED the event loop”). It answersEAI_ADDRFAMILYat once now, as glibc does, andAI_NUMERICHOSTwith a name answersEAI_NONAMEwith no lookup. Test:test_static_resolv_numeric.
Packages (deb, rpm)¶
The init script started the agents with 200000 open files, while pam_limits gives a login
fs.nr_open(4000000): it raises them tofs.nr_opennow, and logs withloggerwhen it cannot./etc/profile.d/yuneta.shasked forulimit -n unlimited, which Linux refuses for open files (a no-op); it raises the soft limit to the hard one now, which also lifts the 1024 that systemd gives a desktop terminal. Takes effect at the next boot or login.
Tests¶
New
raise_open_files_limit()(testing.h): raises the soft limit of open files to the hard one.test_iterator_index,test_c_treedb_literal_winsandperf_c_treedbopen more than 1024 files and failed with “TOO MANY OPEN FILES” in a terminal that systemd started with its soft limit of 1024.test_torn_tail_check_fails5a/5c passed only withCONFIG_DEBUG_TRACK_MEMORY: the lexer’s 64 KB buffer fits a 64 KB block when no tracking header is added. The block of case 5 is 48 KB.test_unlisted_relist_once4 failed on kernels before 6.13, where a ctime moves in ticks of the coarse clock (4 ms): thechmodfell in the tick of the flag. It repeats thechmoduntil the ctime moves.test_c_treedb_literal_wins: a child that cannot take its own io_uring ring prints why before it exits.
v7.25.5 (2026-09-25)¶
What changed after 7.25.4. Each behaviour change has a test that fails on the
code before it, except those listed under “No red test” in TODO.md.
Upgrade steps (operators, read first)¶
Nodes that build from source: rebuild the external libraries first. linux-ext-libs is 1.22 (a patched jansson): run
cd kernel/c/linux-ext-libs && ./extrae.sh && ./configure-libs.sh, thenyunetas clean && yunetas build(it refuses a staleoutputs_extuntil then).Before upgrading a node, look for md2 files that are not whole rows:
find /yuneta/store /yuneta/realms -name '*.md2' -printf '%s %p\n' | awk '$1 % 32'. A file that 7.25.4 wrote after a torn row is not cut by this release: its key fails every load (CRITICAL “...not cut, repair it by hand”) until it is repaired. treedb.md, “A topic that did not load whole”, gives the repair: a script that is linear (86 400 rows in 0.13 s) and runs in aset -esubshell. It names its backup~/<topic>.<key>.<file>.origand refuses to overwrite one, writes through$f.tmp+mv(a failure leaves the.md2unchanged), and writes nothing when more than one boundary passes. A scan of ~21 500 md2 files on four stores (local, wattyzer, both yunovatios) found none. Seedeploying-yunos.md.Before upgrading, run
save-schemafor any draft you want to keep. An unsaved draft that added a topic or column is taken as left by 7.25.4; a topic deleted as an unsaved draft is restored by the first open when the file in use is the literal. Seedeploying-yunos.md.The agent removes audit files older than 7 days at its first start. To keep more, set
"agent.audit_keep_days": <days>in theglobalsection of/yuneta/agent/yuneta_agent.jsonBEFORE you start the new agent (0keeps all). See DEBUGGING.md 5.5.Migrate the timeranger2 topics once, after upgrading. Until a topic created by 7.25.4 or earlier is migrated, no file’s tm range is trusted in it, and a
tmquery (from_tm/to_tm) reads every md2 row of the key. On the master, for each tranger service of each yuno (find them withcommand-yuno id=<id> service=__yuno__ command=services):command-yuno id=<id> service=<tranger> command=mark-tm-order all=1. The agent is not a managed yuno (command-yunodoes not reach it); its own trangers are three:command-agent service=tranger_treedb_yuneta_agent command=mark-tm-order all=1,command-agent service=tranger_system_schema command=mark-tm-order all=1andcommand-agent service=tranger_authz command=mark-tm-order all=1(its C_AUTHZ is the master oftreedb_authzs, whoseusers_accesseshas'tkey': 'tm'). The migration reads every md2 file once, is linear in the number of files, and blocks the yuno while it runs (1 key x 30 files x 20 000 rows: ~19 ms). The price until then is under “Performance”. Seedeploying-yunos.md.If the main agent does not come back after the upgrade, look in its log for “treedb ‘treedb_yuneta_agent’ did not open, its schema was refused” (syslog has “Cannot start agent treedb: ...”): when its treedb’s schema is refused, the agent now exits 0 and is not relaunched. Reach the node through
yuneta_agent22.The first open of each treedb after the upgrade reads what 7.25.4 left in
__system__. It writessaved_schemas/<treedb>.upgrade.jsonunder the__system__tranger; do not delete it. Expect once per treedb, where a first projection of 7.23.0-7.25.4 died after its stamp: WARNING “Restored from the schema from C: ...” when the file in use IS the literal, or, when that open installs a newer literal, INFO “What system misses of the schema file in use is no deletion of the operator’s: ...” (the literal is then projected). When that open’s literal is newer, also WARNING “Removed from system what an older release left there: ...” andwithdrawn_at_openkindleft_by_older_release. None of them is operator work.Do not roll a node back to 7.25.4 or earlier for topics created or migrated by this release without running
mark-tm-orderagain after coming forward: an older binary appends out-of-ordertmrows without writing the marker, and a tm scan can then hide rows.denied_ipsrefuses at accept now, on every TCP and UDP listener of the yuno. In 7.25.4 an ip in a yuno’sdenied_ipswas refused only at the login of a gate that authenticates. After the upgrade C_TCP_S closes its connection at accept and C_UDP_S drops its datagrams, also on a gate that does not authenticate (see “Security”). Before upgrading, review each list:command-yuno id=<id> service=__yuno__ command=list-denied-ips(the agent:command-agent service=__yuno__ command=list-denied-ips). To ban an ip:add-denied-ip ip=203.0.113.7 denied=1.deniedis required: without it the command answers -1, “<role^name>: Denied, TRUE or FALSE?”, anddenied=0writesfalse, which denies nothing. A C_UDP_S withonly_allowed_ipsnow drops every peer thatallowed_ipsdoes not name (7.25.4 never read the attribute): checklist-allowed-ipson those yunos too. Loopback is exempt from both lists. C_UDP_S says a refused peer at its first datagram of each cause and then at most once a minute; the new statrxRefusedMsgscounts every datagram it drops (these WARNINGs are new: 7.25.4 dropped no datagram and said nothing).The stored ip lists are rewritten once, at the first start. A list entry is now kept in the form a peer is looked up by (see “Security”). At its first start with this release each yuno normalises its persisted
allowed_ips/denied_ipsand saves each list once:an entry written in another form of a valid ip is renamed (WARNING “ip list entry renamed to the form a peer is looked up by”):
2001:DB8::1->2001:db8::1,::ffff:203.0.113.7->203.0.113.7;an entry that is not an ip is removed (WARNING “ip list entry dropped, it never matched a peer”, with the entry and the cause): a port (
203.0.113.7:443), brackets ([2001:db8::1]), a host name, or a link-local address without its interface inallowed_ips. These entries never matched a peer in 7.25.4 either;two entries of the same ip become one: the deny wins in
denied_ips, the refusal (false) wins inallowed_ips. After the first start, look in each yuno’s log for the dropped entries and add them again in a valid form: a dropped203.0.113.7:443isadd-denied-ip ip=203.0.113.7 denied=1.remove-denied-ip/remove-allowed-ipof an ip that is not in the list now answer -1 (7.25.4 answered success): a script that removes an ip “in case” sees a failure.
A subscription from a peer keeps only
__first_shot__of its__config__. The C and gobj-js clients of this SDK send nothing else. A client of your own that sends other keys (__hard_subscription__,__own_event__,__rename_event_name__) has them ignored, with a WARNING in the server’s log (“SUBSCRIBING config keys a peer may not set, ignored”).A yuno no longer exits when a log file cannot be opened after its start. It prints one line on stdout and syslog (“_rotatory(): Cannot open ‘<path>’ file, <err>”; on a full disk the new day’s file fails at its create, “_rotatory(): Cannot create ‘<path>’ file, <err>” or “Cannot create ‘<dir>’ directory, <err>”; “_rotatory_truncate(): ...” for a truncate), goes on, and prints “_rotatory(): ‘<path>’ is open again” when the file opens again. Look for those lines in syslog: they are the only sign that a yuno is not writing its log (7.25.4 exited, see “Agent, gobj-c and tools”).
A peer holds at most 5000 subscriptions on a channel of a gate. New C_IEVENT_SRV attributes
max_subscriptions(default5000,0no limit) andmax_subscription_size(default16384bytes of compact json, for the__filter__and the__global__of a peer’s subscription). A subscription beyond them is refused, with a WARNING logged once (“SUBSCRIBING refused, the peer holds max_subscriptions”). A gate whose peers subscribe per device must raise the cap before the upgrade: the SPAs of hidraulia subscribe twice per device. Set it in thekwof theC_IEVENT_SRVof the gate’s channel tree, for example{"name": "input-(^^__range__^^)", "gclass": "C_IEVENT_SRV", "kw": {"max_subscriptions": 20000}}(see ievent.md, “What a peer may hold”).A peer’s
__global__keys that start with_are dropped, with a WARNING (“SUBSCRIBING keys a peer may not set, ignored”, withkeys, at most one each 10 s), as are agbufferkey and the peer’s__local__: of a peer’s subscription C_IEVENT_SRV keeps only__filter__, the allowed__config__keys and the__global__keys of the peer’s own (see “Security”). The C and gobj-js clients of this SDK send no such key.The logcenter’s C_GSS_UDP_S has caps now, and the logcenter runs on their defaults:
max_channels1024(peers, a source ip:port each, held at once; a new peer beyond it is dropped),max_pending_bytes8388608(8 MB of unfinished frames of all peers together; beyond it a datagram is dropped with its peer’s unfinished frame) andmax_frame_size1048576(1 MB; a bigger frame is delivered cut). A logcenter that hears more than 1024 yunos at once setsmax_channelsin its config.A client of your own must send its routing on every frame of a session. C_IEVENT_SRV closes the channel when a session frame has no routing: no
__md_iev__ievent stack, a top record whosesrc_*fields are not strings, or adst_*or__msg_type__that is not a string. It logs one WARNING (“Frame without its routing (md_iev ievent stack), channel closed”,MSGSET_PROTOCOL, the kw capped to 256 bytes), at most one each 10 s. The C and gobj-jsC_IEVENT_CLIof this SDK send the routing on every frame. In 7.25.4 such a frame logged an ERROR with a stack and the whole kw, and was processed (see “Security”).A C_UDP_S with a
udps://url stops its yuno at the start. There is no DTLS: the start is refused with an ERROR (“A secure url (udps://) is not supported by C_UDP_S: there is no DTLS”), and withexitOnError(the default) the yuno exits with 0, so it is not relaunched. In 7.25.4 the server listened with no TLS session, and the first datagram crashed the yuno. Before upgrading, look for such urls and writeudp://instead:grep -rln 'udps://' /yuneta/realms --include='*.json'.An MQTT peer that acks a message with the wrong packet is disconnected. A PUBACK for a QoS 2 message, a PUBCOMP for a QoS 1 message or before the PUBREC (and, in the client, a PUBREL of the wrong QoS) is a protocol error: a WARNING (“QoS mismatch”,
MSGSET_MQTT, withclient_idandpeername), the connection is closed (an MQTT 5 peer first gets DISCONNECT 0x82, protocol error), and the message stays in flight for the next session, as mosquitto does. In 7.25.4 it was an ERROR, and the message was removed as delivered. An ack of a packet id that is not in flight is a WARNING now (“Message not found in trq_out_msgs”; it was an ERROR). A device that sends such acks is disconnected at each one: look for “QoS mismatch” in the broker’s log after the upgrade, and fix the device.A C_TCP_S says a refused connection once a minute per cause. The first connection that the ip lists refuse is logged (INFO “TCP_S: Ip denied” / “TCP_S: Ip not allowed”), then at most one line each 60 s per cause, with
refused(the connections of that cause refused since the last line, that one included),refusedConnxsandnext_log_in_ms. The new statrefusedConnxscounts every refusal. An alert that counts those lines counts causes now, not connections: readrefused, or the stat (command-yuno id=<id> service=__yuno__ command=view-attrs gobj=<full name of the C_TCP_S>). In 7.25.4 each refusal wrote its line.A C_TRANGER handle opened through the agent is its user’s. A handle (
open-rt,open-iterator,open-list) is owned by the channel it was opened by and the user it was opened as. Through the agent (command-yuno) every operator reaches the yuno by one link, so the user decides: another user, or a relayed command that carries no user, gets -403 onclose-rt,close-iterator,close-list,get-pageandget-list-data(“<role^name>: <kind> ‘<id>’ is not yours: another session opened it”). Two sessions of the same user through the agent are one owner. A handle opened through the agent is not closed when the operator’s session ends: close it withclose-rt/close-iterator/close-listas the same user.Use the
yunetasCLI 0.19.4 (pipx upgrade yunetas).find-new-yunosnow marks a row already registered at the new release, andyunetas upgrade-yunos0.19.4 counts those rows apart (N created, M already registered); 0.19.3 works with this agent but counts them as created.
Security¶
A remote peer’s subscription changed the event of every later subscriber (HIGH; 7.25.4 too).
gobj_publish_event()gave every subscriber the same kw, and applied each subscription’s__local__(keys removed) and__global__(keys added) to it. A peer that subscribed through a gate could forge or strip keys of the event for every subscriber after it and for the publisher, make a later subscription’s__filter__drop events or pass events it should not, and have the gate’s__md_iev__(another user’s name and channel) reach the local subscribers after the gate. Agbufferkey in a peer’s__global__crashed the yuno: it was merged into the kw, and the gate’s serializer took the integer for a pointer -- any authenticated peer could crash the yuno remotely. Now a subscription with a non-empty__local__or__global__gets its own kw, akw_twin()(a new top level; the values shared, agbufferincrefed); the__filter__andmt_publication_filterof every subscription see the publisher’s kw, and the other subscribers still share it, at no cost. The gate’s and the client’smt_inject_event()work on a twin of a shared kw, and copy the__md_iev__they write into (7.25.4 also wrote the message type into the subscription’s stored__global__). A twinned delivery costs ~0.23 us more, whatever the size of the event (see “Performance”). gobj-js 7.25.8 has the same fix (see the JS section). Testtest_c_ievent_srv_peer_subs.C_IEVENT_SRV keeps of a peer’s subscription only what a peer may set: its
__filter__, the allowed__config__keys (__first_shot__, below) and the keys of its own in__global__. A__global__key that starts with_(the framework’s:__md_iev__,__md_yuno__,__service__, ...) or isgbuffer, the peer’s__local__and any other key of the frame are dropped, with a WARNING (“SUBSCRIBING keys a peer may not set, ignored”, withkeys); up to 7.25.4 each of them reached the publish. A key of the peer’s own comes back to that peer only. Testtest_c_ievent_srv_peer_subs.A peer could hold any number of subscriptions (7.25.4 too). Each one costs a scan of the publisher’s subscriptions when it is made, and one more on every publish of its event: 20 000 subscriptions of one peer blocked the event loop for 80 s, and every later publish took 13 ms. New C_IEVENT_SRV attributes:
max_subscriptions(default5000,0no limit; beyond it a subscription is refused, WARNING “SUBSCRIBING refused, the peer holds max_subscriptions”, logged once until the peer is under the cap; a repeat of a subscription the peer holds takes no room) andmax_subscription_size(default16384bytes of compact json, for the__filter__and for the__global__; WARNING “SUBSCRIBING refused, bigger than max_subscription_size”). See “Upgrade steps”. Testtest_c_ievent_srv_peer_subs.A peer no longer sets the size and the rate of the gate’s log lines. The WARNING of an unsubscribe that matches nothing dumps at most 256 bytes of the kw, with
kw_size; it, the WARNINGs of dropped keys, the size refusal and the INFO of a withdrawal of a refused subscription are logged at most once each 10 s per kind and channel, with asuppressedcount; an authz refusal (“No permission to subscribe event”) is logged once per service and event on a channel. These frames of a peer are one WARNING each 10 s per kind (the peer’s strings capped to 128 bytes, the kw to 256), where 7.25.4 logged a line per frame, an ERROR with the whole kw but for the WARNING “Service not found”: an event to a service the channel may not reach (“event ignored, dst_service not authorized for this channel”) or that does not exist (“event ignored, service not found”), a command or stats request of a service that does not exist (“Service not found”), a stats request of another service (“Not authorized to request stats of a different service”), and a (un)subscription of an event that is not public (“SUBSCRIBING event ignored, not PUBLIC or PUBLIC event”, “UNSUBSCRIBING ...”). Such a subscription also left the channel connected and deaf in 7.25.4 (the gate answered -1 and did not drop it); it is refused and the channel goes on. A role or name that is not the gate’s (“It’s not my role, channel closed”, “It’s not my name, channel closed”) is a WARNING with the routing capped (7.25.4: an ERROR with the whole routing); the channel is closed, as before. A command or stats request that names no command, or whoseserviceis not a string, gets a negative answer (“Request without the command or stats it asks, refused”); 7.25.4 logged ERRORs with stacks and ran the command"". A peer that repeated a bad frame wrote one line (and one stack) per frame, of any size. Testtest_c_ievent_srv_peer_subs(a repeated withdrawal with a 2 KB key: one line, capped).C_IEVENT_SRV stores of a peer’s routing only its own hop (7.25.4 too). The back-metadata of a peer’s subscription (what routes each event back to the peer) is built from the top record of the frame’s ievent stack alone: the peer’s hop, with the gate’s stamps, reversed. It is measured against
max_subscription_size, and a bigger one is refused (WARNING “SUBSCRIBING refused, its routing is bigger than max_subscription_size”, at most one each 10 s). Up to 7.25.4 the frame’s whole__md_iev__was stored, unmeasured, and sent back with every event: a peer made the gate hold, and send it again with each event, keys and hops of its choice, of any size. Testtest_c_ievent_srv_peer_subs.C_IEVENT_SRV: a repeated subscription of a peer is kept (7.25.4 too). A frame that repeats a subscription the channel holds leaves it as it is: no second
mt_subscription_added()and no second first shot. A frame that overrides a held subscription (it matches it, with another kw) replaces it. Each is one WARNING at most each 10 s (“SUBSCRIBING repeated, the one held is kept”, “SUBSCRIBING overrides one held, it is replaced”, withsuppressed). Up to 7.25.4 each repeat was deleted and made again, with a WARNING that carried a stack and the whole kw. Testtest_c_ievent_srv_peer_subs.C_IEVENT_SRV closes the channel of a session frame without its routing (7.25.4 too): no
__md_iev__ievent stack, a top record whosesrc_*fields are not strings, adst_*or a__msg_type__that is not a string. One WARNING at most each 10 s (“Frame without its routing (md_iev ievent stack), channel closed”, the kw capped). Such a frame cannot be answered nor routed back. Up to 7.25.4 it logged an ERROR with a stack and the whole kw, and was processed. The gate reads the routing with plain json calls: thekw_get_*()readers logged a wrong type with a stack and a dump of the frame. See “Upgrade steps”. Testtest_c_ievent_srv_peer_subs.A peer’s own
__username__reached the service (7.25.4 too). The gate stamps__username__on a command, a stats request and an event withkw_set_dict_value(), which kept a key that was already there: a peer that put its own__username__in the kw was that user for the service, andcommand_parser.c’s authz check reads that key.kw_set_dict_value()now overwrites (see “Agent, gobj-c and tools”). The routing stamps (input_channel,input_service) were already overwritten. Testtest_c_ievent_srv_peer_subs(a peer sends"__username__": "admin"in a command, a stats request, an event and a subscription).C_GSS_UDP_S (the logcenter) caps what its peers make it hold:
max_channels(1024 peers),max_pending_bytes(8 MB of unfinished frames of all peers together) andmax_frame_size(1 MB a frame, whose buffer now starts at 4 KB and grows): a sender of one-byte datagrams from many source ports cannot take the logcenter to its memory ceiling. See “Transports”, the per-peer label. Testc_udp_s_rx(case 4).C_TRANGER: a handle opened by one session is refused to another (7.25.4 too).
close-rt,close-iterator,close-list,get-pageandget-list-dataof an id that another session opened answer -403, “<role^name>: <kind> ‘<id>’ is not yours: another session opened it”, with a WARNING (“Handle of another session, refused”, withkind,id,srcanduser); a peer could read the pages of another user’s iterator, or close its feed. The check runs before the liveness check, so such a request cannot reap the handle either. A handle’s owner is the channel it came by and the user it came as: through the agent (command-yuno, which reaches the yuno over its one C_IEVENT_CLI link) another user is refused (the__username__that the agent stamps), and so is a relayed command of no user; two sessions of the same user through the agent are one owner. A local gobj is trusted. See “Upgrade steps”. Testtest_c_tranger.jansson overflowed its buffers on a long string (v2.15.1; not fixed upstream, and the distribution’s libjansson does the same). Its lexer ignored a failed save and wrote past its buffers: a json string longer than about half of
MEM_MAX_BLOCK(~8 MB with the default 16 MB) crashed the process that parsed it -- anyjson_load*(), json received from a peer included. linux-ext-libs 1.22 carriespatches/jansson/0001-load-stop-the-lexer-when-a-save-fails.patch: the lexer stops withjson_error_out_of_memory, and every allocation failure of the parser now sets that error.configure-libs.shappliespatches/<lib>/*.patchafter the checkout. The ESP32 copy of jansson has the same fix.timeranger2: a record this process has not the memory to parse is reported (“Cannot read the record, this process has not the memory to parse its content (MEM_MAX_BLOCK)”, with its
key) instead of crashing the reader; a torn md2 whose last row names such a record is flagged, not cut.The agent’s audit wrote secrets and keystrokes in clear text. Up to 7.25.4 each record held the whole kw of the command:
check-user-pwdandset-user-pwdwrote the password, awrite-attrof a secret attribute its value, and every console keystroke (write-tty) its bytes in base64, in files kept for ever under/yuneta/realms/agent/agent/audit/. A secret is now<redacted>(also inside a JSON text given as a string, under a name written with json escapes, and the credentials afterBasicorBearer) and a keystroke is only counted (see “Agent, gobj-c and tools”, the audit record). What a peer says about itself is redacted too:user, each field of the hops ofsource,console_purposeand the console name of a console write (a JWT there is<redacted>), and such a field longer than 1024 bytes is written as<N bytes, not scanned, sha256:HEX>. Audit files written before the upgrade keep what they hold: remove or protect them. Testtest_audit_record.denied_ipsis asked at accept. C_TCP_S refuses a peer that the yuno’sdenied_ipsnames before it spends a channel (INFO “TCP_S: Ip denied”, the socket closed), with or withoutonly_allowed_ips, and the deny wins overallowed_ips; withonly_allowed_ipsa peer thatallowed_ipsdoes not name is refused as before (“TCP_S: Ip not allowed”). Every refusal is counted in the new statrefusedConnxs, and said on the transition: the first refusal of each cause, then at most one line a minute per cause, withrefused(the refusals since the last line) andrefusedConnxs(7.25.4 wrote a line per refused connection, so a denied host that reconnects in a loop would fill the log). Loopback (127.0.0.x,[::1],[::ffff:127.0.0.x]) is exempt from both lists (7.25.4 exempted only127.0.0.x, and only from the allow-list). In 7.25.4 only the allow-list was asked at accept: a denied ip was refused only at the login of a gate that authenticates, so a gate that does not (an MQTT or IoT field port) built its channel and ran its protocol. C_UDP_S asks both lists for every datagram in the same way: every drop is counted in the new statrxRefusedMsgs, and said on the transition -- the first drop of each cause is a WARNING (“UDP_S: Ip denied, datagram dropped” / “UDP_S: Ip not allowed, datagram dropped”), then at most one a minute per cause, withdropped(since the last WARNING),rxRefusedMsgsandnext_warning_in_ms(the source of a datagram can be forged, so a WARNING per datagram would let a peer fill the log). 7.25.4 never read its documentedonly_allowed_ips, and heard every peer. See “Upgrade steps”. Testsc_tcp_s_ip_lists,c_udp_s_rx.An ip-list entry is kept in the form a peer is looked up by. The
add-/remove-allowed/deniedip commands read the text withinet_pton()and store its canonical form (2001:DB8::1->2001:db8::1,::ffff:203.0.113.7->203.0.113.7,fe80::1%eth0->fe80::1%2); the answer says it (“<role^name>: ‘2001:DB8::1’ stored as ‘2001:db8::1’”). They refuse a text that is not a numeric ip (a host name, a port, brackets, an interface on an address that is not link-local), with the cause. A link-local address inallowed_ipsneeds its interface, and matches that interface only; indenied_ipsone without its interface is denied on every interface.remove-accepts any form of the ip, and answers -1 for an ip that is not in the list (“<role^name>: ip ‘<ip>’ is not in denied_ips”). Up to 7.25.4 the text was stored as typed and the command answered success, so2001:DB8::1,[2001:db8::1]or203.0.113.7:443were listed and never matched a peer: the operator believed an ip was banned, and it was not. The stored lists are normalised at the first start (see “Upgrade steps”). Testc_tcp_s_ip_lists.C_IEVENT_SRV: a remote peer no longer chooses how its subscription is delivered. Of the
__config__of a peer’s subscription only__first_shot__is kept;__hard_subscription__,__own_event__,__rename_event_name__and any other key are removed, with a WARNING (“SUBSCRIBING config keys a peer may not set, ignored”, withkeys). Up to 7.25.4 a peer could make its subscription hard: it outlived the session and reached the next user of the static channel, past any check of who that user is; with__own_event__, a delivery that failed stopped the publish before later subscribers got it. The close of a channel now removes every subscription the channel made, hard or not. Testc_subscription_authz(a client asks hard, own, rename and first_shot).C_IOGATE: a
channel_namethat does not match every channel no longer hangs the yuno.view-channels,enable-channel,disable-channel,trace-on-channel,trace-off-channelandreset-stats-channellooped for ever on the first channel that did not match (7.25.4 too): the event loop was blocked, and the process ignored SIGTERM. Withenable_command_authzoff (the default), any authenticated peer that reached an iogate could do it, the agent’s own__input_side__included. Achannel_namethat is not a valid regular expression answers “<role^name>: channel_name is not a valid regular expression: ‘<re>’” (it answered “regcomp() failed”). Testc_subscription_authz.gobj-c: the
authzstrace no longer reads a kw the checker has freed (7.25.4 too): with the trace on, every authz check printed the kw after the checker had released it, a use-after-free. The trace now holds its own reference. Testc_subscription_authz.An IPv6 peer can be listed.
is_ip_allowed()/is_ip_denied()key a peer by its ip without the port, for IPv6 too ([2001:db8::1]:443->2001:db8::1) and for an IPv4 peer seen by a dual-stack socket ([::ffff:1.2.3.4]:80->1.2.3.4). 7.25.4 cut the port at the first:, so every IPv6 peer was looked up as"["and no list could name it.The subscription authz is enforced, behind a gate that is off by default. New yuno attribute
enable_subscription_authz(SDF_RD, defaultfalse, apart fromenable_command_authz). With it on, an external subscription (throughC_IEVENT_SRV, whose user C_AUTHZ set) to an event flaggedEVF_AUTHZ_SUBSCRIBEneeds the permission of the publisher’s gclass whosealiasis__subscribe_event__, or else the global__subscribe_event__. C_NODE flags its fiveEV_TREEDB_NODE_*, and C_TRANGER its realtime feedEV_TRANGER_RECORD_ADDED, and each puts that alias onread: with the gate on, a peer needsreadon the treedb or tranger service to get its feed (for C_TRANGER, the permissionopen-rtalready asks). Every other public output event in the tree was checked, and none carries data a peer could not reach otherwise (the table is in YUNO_AUTH.md §4.6). A refusal is an ERROR (MSGSET_AUTH, “No permission to subscribe event”, with the service, the event, the permission and the user); the subscription is not made and the channel stays open. An internalgobj_subscribe_event()is never checked. In 7.25.4 the check was commented out: theEV_TREEDB_NODE_*feed of a treedb went to any authenticated user, and a user withoutreadcould subscribe to a C_TRANGER feed with no filter and get every record of the realtime lists other users opened. Nothing changes until a yuno sets"yuno": {"enable_subscription_authz": true}in its config; that needs a C_AUTHZ role model (the checker is fail-closed), and every user of a treedb GUI then needsreadon the services it watches (a refused GUI keeps its session and gets no live updates). YUNO_AUTH.md §4.6. Testc_subscription_authz.libjwt (vendored): the
jwks_*keyring getters (jwks_item_get(),jwks_error_any(),jwks_error(),jwks_error_msg(),jwks_error_clear()) accept a NULL set, as upstream’s do:jwks_item_get(NULL, 0)crashed. So dojwks_item_count()(0),jwks_find_bykid()(NULL, also for a NULL kid),jwks_item_add()(1) andjwks_item_free_bad()(0), which upstream does not guard either: they carry local guards (// ArtGins: NULL-safety, beyond upstream, listed inkernel/c/libjwt/README.md). C_AUTHZ never passed a NULL set. Testlibjwt/test_jwt_alg_confusion.
Data loss and integrity¶
C_AUTHZ
disable-user,enable-userandset-max-sessionsERASED the user’s local password (since 55266cbb5, 2025-11): they read the user without the hiddencredentialscolumn and wrote the whole view back. They write only{id, <field>}.gc-assetswith a snap ACTIVE took the assets of live nodes written after the snap (since 7.18.1): memory holds the snap’s photo. The gc refuses while any snap is active in the tranger. A refused gc takes no asset row;gc-assetsstill sweeps (and reports) orphan blobs, the bytes no asset row names. The library’streedb_gc_files()takes nothing on a refusal; the newtreedb_gc_files2()returns the report.Damaged keys. A key damaged on disk (an md2 that cannot be opened or read, a record whose content cannot be read) makes every load of the key fail (
load_failed): a row or content that cannot be read fails the load that meets it, and an md2 the cache build cannot count flags the key at the open, also after a restart. The flag of a file is cleared when the file reads again (an append into a file still unreadable is refused). A keylesstranger2_open_list()still loads every readable key and opens its realtime feed, and reports the failure asload_failed/load_failed_keys. treedb remembers those keys and refuses what memory would answer wrong: a create of such an id (it would shadow the stored record); shoot/activate and the snapshot-guarded deletes when__snaps__is partial;gc-assetswhen a topic with afilecolumn or__assets__is partial; a delete of a node, forced or not, whose hooks hold a topic that did not load whole (“Cannot delete node: a topic its hooks hold did not load whole, a child that did not load may hang from it”; 7.25.4 deleted it, and a child that did not load kept naming it).treedb_delete_instance()refuses when it cannot read every row of the key, and the asset snapshot guard fails closed when it cannot read. Recovery procedure in treedb.md (“A topic that did not load whole”), from least to most destructive.msg2db: an id whose history did not load whole used to serve an older message as current. It is now reloaded newest first, up to the damage: a pkey2 whose newest message comes after the damage is served as before; one whose newest message is in the damage or before it is ABSENT (its state is unknown) until its next message -- unless the damaged file is the one new messages go to (the current period’s file; the db_history trangers use one file per YEAR): then every new message of that id is refused (
msg2db_append_message()answers NULL) until the file is repaired (treedb.md) or the period changes, and a second ERROR says so. A torn last md2 row is not damage (below, “An md2 that ends in a part of a row”). An ERROR names the id withserved=N, and the newmsg2db_id_incomplete()tells “unknown” from “none” while the msg2db stays open. For the db_history alarms of the projects (wattyzer, yunovatios, estadodelaire, hidraulia), until an absent alarm’s next message: a device that reports it active gets it announced as new (a notification may repeat), a clear that was in the unread message is not announced, andmsg2db_list_messages()leaves it out.tr_msg2db.mdshows how a consumer can usemsg2db_id_incomplete()to avoid the repeat; no project repo was changed.An append that was never acknowledged is not damage. When the md2 of an append cannot be opened or created, or its seek or write fails, the content written for it is truncated back, BEFORE the critical log (“Cannot append record, its md2 file cannot be opened: its content was cut back” for the open) (with the default
on_critical_errorthe process exits there, and 7.25.4 left the bytes behind). An md2 with 0 rows beside a non-empty.json(a kill or a power cut between the two writes) is ignored with a warning, and its key loads (7.25.4 ignored its rows too, without a warning).An md2 that ends in a part of a row (a power cut during the 32-byte row write) is cut back to whole rows by the master when the file is first examined, with one WARNING (“md2 file of the key ends in a part of a row: an append that was never acknowledged was cut back”, with
old_size/new_size); the key loads whole and appends go on. An append that meets such an md2 cuts it back the same way before it writes its row, so a row always starts on a row boundary; if the cut fails, the append is refused with a CRITICAL (“Cannot append record, its md2 file ends in a part of a row that cannot be cut back: the append is refused”). The cut removes only the part of a row after the last whole row, never an acknowledged row, and only when the last whole row and the row before it are good (each names one whole record of the content, in order: one json value and its NUL, or only zeros for a deleted instance), and the end-aligned 32 bytes, read as a row, neither end the content nor name a whole record. A replica reads only the whole rows and writes nothing. 7.25.4 logged a CRITICAL (“Cannot read last record, md2 file corrupted”) and left the whole file out of the key, on a master and on a replica: the acknowledged rows of that file were missing from every load, and nothing failed. A replica also did this when it caught a live master mid-write.An md2 that must not be cut is flagged. An md2 that 7.25.4 kept appending into after a torn row (its later rows off the row boundary, with or without content after its last row) is NOT cut, on a master or a replica: it is flagged (
load_failed, appends refused) with a CRITICAL (“md2 file of the key ends in a whole row that is not on a row boundary: written by 7.25.4 after a torn row; not cut, repair it by hand”, with acause). The check of a node and the repair are in “Upgrade steps”. An append that meets such a file refuses (“...that must not be cut back: the append is refused”) and flags it in memory at once (ERROR “md2 file of the key flagged unreadable at an append: every load of the key says load_failed”); with the defaulton_critical_errorthe yuno then exits at the CRITICAL of the refused append. A check that cannot RUN (the.jsoncannot be opened or read, no memory) refuses the append (“Cannot append record, the torn tail of its md2 file cannot be checked now: the append is refused, the file is not flagged”) and does not flag the file; the next append checks again. At an open the file is still flagged. A candidate range larger than the largest memory block of the reader is read in parts of 64 KiB, never allocated whole: a NUL before its end still says it is not a record, and one whose only NUL is its last byte is never cut on -- the file is flagged (it can be a record written by a yuno with a largerMEM_MAX_BLOCK, which is set per yuno).A record with a NUL in a string (
json_dumps()writes it as\u0000) is read back. 7.25.4 took it at the append, and every read of it failed with a CRITICAL (“Bad data, anystring2json() FAILED.”) and handed the record to the callback as NULL. A read that fails now logs “Bad data, the content of the record is not json”, with the janssonerrorandposition.A short write of
topic_var.json/topic_cols.json, or of the zero wipe oftranger2_delete_instance(), is logged as a short write (“...short write: the file size limit or the disk is full”), not with a stale errno (7.25.4: “Cannot write in json file” for the json files, “write() zero payload FAILED” for the wipe).A read that returns fewer bytes than asked logs “... short read” with
read/expected(it logged “read FAILED” with a stale errno); a short write likewise logs “... short write”. A read error or short read of an md2’s first or last row at the cache build flags the file (it went on with a zeroed row).A treedb write whose save fails is taken back in memory. In 7.25.4
treedb_update_node(),treedb_link_nodes(),treedb_unlink_nodes(),treedb_autolink(),treedb_clean_node(),treedb_replace_links()and the child unlinks of a forced delete changed the node (fields, fkeys, parent hooks) before the save and kept the change when the save failed: a read then answered a value the disk never had, and a retry of the same write found nothing to write. Now the node goes back to what the disk has, and the call answersNULLor -1; a write taken back tells no event.treedb_autolink(),treedb_clean_node()andtreedb_replace_links()answer -1 when the save fails (7.25.4 ignored it); a link or unlink that fails part way is taken back whole. Link and unlink events, and the parent’sEV_TREEDB_NODE_UPDATEDin the compatible mode, are told only after the child is saved, in the same order among themselves (7.25.4 told the link events before the save). A write that is taken back puts every child back in its place in the hooks of its parents (list and dict hooks), in every instance of the parent that held it, each in its place (a topic with apkey2has several instances of one id, and an unlink takes the child out of all of them), and a stale fkey ref it removed (its hook gone, or re-pointed to another column) goes back into its field alone, never linked. When memory cannot be taken back whole, an ERROR says so: “A write that did not reach the disk could not be taken back whole in memory: the links in memory differ from the disk until the treedb is opened again”.treedb_autolink()refuses a ref whose hook fills another column than the one the ref arrives in (“fkey reference: its hook does not link into this column”), astreedb_replace_links()does; 7.25.4 linked it through the other column.A forced
treedb_delete_node()that is refused changes nothing. A child that cannot be saved unlinked stays linked (“Cannot delete node: still has down links”), a key that cannot be deleted refuses the delete too, and the children already unlinked are put back and saved again while the node keeps its parents. In 7.25.4 the parent was deleted while a child on disk still named it, and a delete refused by its key left the node unlinked in memory and every child unlinked on disk. A child that cannot be put back stays unlinked, in memory as on disk, with an ERROR (“A refused delete cannot put back a child it had unlinked: the child stays unlinked, in memory as on disk”), and the events of its unlink are told. A forced delete tells the events of its unlinks after the node has left the indexes and beforeEV_TREEDB_NODE_DELETED, and until it returnstreedb_save_node()refuses the node (“Cannot save a node that is being deleted”): a subscriber that saved it from one of those events could bring it back on disk.treedb_delete_instance()answers -1 when a tombstone fails, and keeps the instance. 7.25.4 tombstoned the rows newest first, went on after a failure, answered 0 and dropped the instance, which came back at the next open from an OLDER row. The rows go oldest first and the first failure stops the delete (“Cannot delete instance, a row of it cannot be tombstoned: the instance stays, its newest rows alive”); the older rows tombstoned before it are gone from the history.A delete of a node with several instances (a
pkey2topic) looks at all of them.treedb_delete_node()deletes the key, every instance of it, but counted and unlinked the children and the parents of the node it was given alone (7.25.4 too): a child of another instance was left naming a key that is gone (“Node not found” at the next open), and another instance stayed in the hook of its parent, so a forced delete of that parent saved it back into the deleted key and the node was back after a reopen. The children of every instance are counted (withoutforcethe delete is refused) or unlinked and saved (withforce); every other instance that a parent’s hook holds refuses a delete withoutforce, and withforceit leaves that hook. A refused delete puts all of it back. The instances of a child that name the key and that no hook holds (a non-primary instance that nothing links) are counted too: withoutforcethe delete is refused (“Cannot delete node: has down links”, now withchildrenandunheld_instances), and withforcetheir ref is cleared and saved first, and put back if the delete is refused (7.25.4 left them naming a key that is gone). The newest record of each key that those saves and the unlinked children touch stays on the instance that wrote it, so the next reload takes the same primary as without the delete (in 7.25.4 the child that a forced delete unlinked last became the primary of its key at the next reload, and a new instance created since lost to it).treedb_delete_instance()takes the instance out of the hooks of its parents, and hands its children to the primary (7.25.4 did neither). A non-primary instance sits in a hook when it is linked into a slot that does not hold the primary (an array hook, or the empty hooks of a new parent instance): a delete of that parent withoutforcewas refused for an instance that was gone, and a forced one saved the instance back -- its newest row, the primary of its key after a reopen. The children that a deleted parent instance held hung from no visible parent until a reload. Now the instance leaves every parent hook, the primary takes its place where it names that parent too, and the children go to the primary’s hooks, where a reload puts them. Nothing is written.After a restart, a pkey2 lookup of a node’s primary returned a copy of it (7.25.4 too). The load built the slot of the primary’s own pkey2 value from its record as a second object, with no links, until the first save of the primary. An update through the instance lookup (
treedb_get_instance(), or C_NODE’supdate-nodewhen the primary does not match its filter) changed that copy: the primary kept the old values, and its next save wrote them back over the new ones and freed the copy while the caller still held it. C_NODE’sdelete-nodewith a pkey2 value took the copy for another instance and tombstoned the rows of the primary’s version before deleting the key; when that delete was refused (a link, withoutforce), the node came back at the next open at an older version, without its links. The load now puts the primary node in that slot: one instance, one node, with its links. A new instance still does not become the primary in memory; the reload makes the newest record the primary, as before.treedb_save_node()refuses a node that no index holds (ERROR “Cannot save a node that no index holds: its record would bring back what was deleted”): a deleted key or instance still reachable through a pointer that a hook or a caller kept. In 7.25.4 such a save wrote a record into the deleted key, and the node came back at the next open. Testtest_tr_treedb_delete_instance(every case of these five bullets, each red before its fix).treedb_save_node()refuses a node whose pkey2 value was changed in place (ERROR “Cannot save a node whose pkey2 value changed in place: its record would be another instance”, withpkey2_name,old_valueandnew_value): a caller that wrote a new value of a pkey2 column into the node and saved it. In 7.25.4 the save wrote a new instance on disk while memory kept the node in the slot of its old value; when the new value was another instance’s, the node also took that instance’s slot, the primary’s included. Change a pkey2 value withtreedb_update_node(), which refuses it as in 7.25.4, or create the new instance. A node that no slot of a pkey2 holds is saved as before.treedb_create_node()indexes the pkey2 value its record holds (a columndefault, a converted value); 7.25.4 indexed the raw kw value, so the instance of a defaulted value was missing until a reload.test_c_node_link_eventstest 21 asserts the invariant C_NODE’sdelete-noderelies on: a pkey2 lookup of the primary’s own value returns the primary, before and after a reopen. It holds with more than one pkey2 too (a save no longer re-points a slot that the primary holds, where a second pkey2 took the primary’s slot of the first), and for a key with instances and no primary (below, a snap).A reference is never cut in silence. 7.25.4 wrote
topic^id^hookintochar[NAME_MAX]and decoded it into parts ofNAME_MAXwith no check. A reference with a part that cannot be decoded back is refused and nothing moves: built, “Cannot build the reference of a node: a part of it is too long, or holds a ‘^’”; decoded, “Wrong reference: a part of it is too long”.treedb_create_node()refuses, for a topic with hooks, an id holding^(“Invalid ‘id’: it holds a ‘^’, the separator of a reference”) or ofNAME_MAXbytes or more (“Invalid ‘id’: too long to be part of a reference”); a topic without hooks keeps any id.A delete of a parent no longer leaves a child pointing at a node that is gone. A topic without hooks keeps any id, also one with a
^or one ofNAME_MAXbytes, and such a child can hang from a hook like any other. The delete guard now counts every child a hook holds (count_node_children()), so a delete withoutforceis refused (“Cannot delete node: has down links”) and a forced delete unlinks the child. In 7.25.4 a dict hook did not count a child whose id holds two^: the parent was deleted, forced or not, and the child’s fkey named a node that is gone, also after a reopen (“Node not found”). The guard no longer builds a reference string for each child only to count them (that change alone takes 2-4% offdelete_parent, on top of the ~6% against 7.25.4; see “Performance”). Testtest_tr_treedb_hook_hygiene(test_children_without_ref()). Therefs/hook_refsoption of a hook lists every child, also one whose id holds a^or isNAME_MAXbytes long ("users^a^b"); in 7.25.4 the reference of a long id was cut in silence, and named another node or none.decode_child_ref()refuses such a reference, and logs why.A hook tells two topics apart. 7.25.4 tested hook membership by the bare id: the group
xafter the userxwas a skipped duplicate in a list hook (its fkey written all the same) and took the user’s place in a dict hook. A list hook holds both; a dict hook refuses the second (“Cannot link, the dict hook holds a node of another topic with this id”), a load keeps the first (“A dict hook holds a node of another topic with this id: this link is not loaded”), and an unlink does not delete the other topic’s entry.C_NODE
update-nodewithautolinkis ONE write (the newtreedb_update_node_and_links()): fields, links and save; a failed save takes all of it back and tells nothing. In 7.25.4 it was three calls, and a failed save left memory with the fields and links, withEV_TREEDB_NODE_LINKEDtold. It is also what C_AUTHZ does when it creates a user with a role.The repair of several active snaps at open keeps a snap active in memory when its deactivation cannot be saved (7.25.4 made it inactive in memory only), and a replica does not try it.
A link to a parent whose id cannot make a reference (an id of
NAME_MAXor more, or holding^, loaded from an older store) is refused and nothing moves (the “Cannot build the reference of a node” above); in 7.25.4 the reference was cut, the link answered 0, and it was lost at the next open. A save that meets a wrong reference in an fkey column leaves out that reference alone, and logs it: it saved the record without the whole column, so the valid links of the column were lost at the next open (7.25.4 too).An unlink clears the parent’s ref in every instance of the child (7.25.4 too). A topic with a
pkey2has several instances of one key, and each carries the fkey it was saved with; the unlink cleared the ref only in the instance it was given, so the other instances still named the parent, and a reload linked them back. Now every other instance that names the parent is cleared and saved first, and the child the unlink was given last, so the child’s record stays the newest; if any of those saves fails, everything is put back and saved again (once each, in the order saved, and the newest record of the key goes back to the instance that wrote it), and the unlink answers -1. Afilecolumn is left alone: each instance holds its own asset. Testtest_tr_treedb_delete_instance.A stale ref in a string fkey column is removed (7.25.4 too). A ref whose hook was renamed or re-pointed is removed from the child by the first relink, clean or forced delete; in a column whose fkey is a single string, the removal wrote the value through
kw_set_dict_value(), which kept the old value (see “Agent, gobj-c and tools”), so the clean and the forced delete of such a child failed for ever (“Cannot clean the links”). Testtr_treedb_hook_rename(a string fkey with a renamed hook, clean and forced delete).With a snap active, the create of a key that the snap does not hold takes the slot of its value (7.25.4 too). A key created after the snap has instances in the secondary indexes and no primary in memory; its create made a second object for the same instance, and C_NODE’s
delete-nodeof that key then went half way (“delete_primary_node() FAILED”, then “Node not found”). A create that makes the primary of such a key now takes the slot of its value, so the slot of the primary holds the primary. A delete of a key that only the secondary indexes hold no longer logs “delete_primary_node() FAILED”. Testtest_tr_treedb_delete_instance.gc-assetskeeps the asset that a non-primary instance names (7.25.4 too): only the primaries were read, so the gc took the asset of an older instance that still names it. Testtest_tr_treedb_delete_instance.A relink of one instance of a child no longer takes a sibling instance out of a dict hook (7.25.4 too): a move of the instance to another parent (
treedb_link_nodes()) removed the slot of the old parent by the child’s id, so a sibling instance that still named that parent left the hook with it. The slot is removed only when it holds that very node. A directtreedb_unlink_nodes()still takes every instance out, and clears the parent’s ref in each (above). The two normal cases are no longer errors: the hook holds another instance of the child, or the child is a non-primary instance that nothing holds (7.25.4 logged the ERROR “Child data not found in ... parent hook”; one of the list-hook cases said “dict”, and now says “list”). A new instance of a child held through a list hook no longer warns “Duplicate fkey on load, deduping parent hook”. Teststest_tr_treedb_delete_instance(test_relink_through_a_dict_hook_keeps_the_sibling, a dict hook over a string fkey).A snap shot no longer chooses the primary of the next reload (7.25.4 too). A shot that clones a primary already tagged by an earlier snap made the clone the newest record of its key, so a new instance created since (e.g. by an
install-binary) lost to the photo at the next reload. The newer instance’s record is written again after the clone. Testtest_tr_treedb_delete_instance.A link that fills the hook of a new instance no longer warns “Parent ref already in child fkey, skipping duplicate” (7.25.4 too). The ref names the parent’s KEY, which all its instances share, so it is already in the child’s fkey when a new instance of the parent takes a child that an older instance holds, or when a new instance of the child inherited it: every
create-yunoof a new release warned for its configuration (and for its binary, when the binary was installed after the old release linked it). The link fills the hook of the new instance and nothing else. The warning stays for a real duplicate: the same pair linked twice. When no other instance of the parent key explains the ref (none holds the child, and this one is not a new instance with an empty hook), the hook LOST a child its fkey names: the link repairs it and warns “Parent hook had lost a child its fkey names: repaired” (with the topic, parent id, hook and child id); 7.25.4 repaired it under the duplicate warning. Teststest_c_agent_find_new_yunos(its fixture now links binaries and configurations ascreate-yunodoes) andtest_tr_treedb_hook_hygiene(a hook emptied by hand on a parent with one instance warns once).Lost lock. A master that lost its lock while stopped (another process took the store) writes nothing: every write path, including the three md2 flag rewriters (
tranger2_write_user_flag,tranger2_set_user_flag,tranger2_set_system_flag, which on a replica logged a CRITICAL and exited), retakes the lock first or refuses. If another process holds the lock the revive is a CRITICAL aton_critical_error, as at startup; any other lock failure demotes the tranger to replica (masterfalse,master_losttrue) with an ERROR and the yuno keeps running -- watch for it. C_TRANGER’smasterreads the effective state, and so do the newmasterandstoppedfields of C_TREEDB’streedbsrows.topic_var.json(the rowid counter) andtopic_cols.jsonare replaced through a fresh.newfile (O_EXCL|O_NOFOLLOW) and a rename, never rewritten in place; a topic_version change fsyncs, writes cols before var, and on a failed write leaves the version on disk unchanged.tr_queue and tr2q_mqtt: a queue whose key did not load no longer saves a
first_rowidpast the pending messages (they were skipped for ever, even after the md2 was repaired), and its periodic backup is refused until a load reads every pending message.C_NODE link, unlink and delete are refused on a replica before memory moves (direct C callers; the commands already checked).
A key directory that cannot be LISTED (EMFILE, EACCES, ENOMEM, or no memory for an entry of the listing) loaded in 7.25.4 as an EMPTY key: the load said nothing (only the helper’s ERROR “Cannot open directory” when the directory could not be opened; nothing at all for a lost entry). After a restart a treedb node was absent, a create of its id was accepted over the history nobody read, and an append counted its row as the first of its file. The key is flagged (
"unlisted", ERROR “key directory cannot be listed when its cache was built: every load of the key says load_failed”): every load of it saysload_failed(a keyless list names it inload_failed_keys, so treedb’s create guard fires), and an append lists the key again first -- refused while it still cannot be listed (“Cannot append record, its key cannot be listed: ...”). A topic whosekeys/cannot be listed is not opened (tranger2_open_topic()answers NULL, “Cannot open topic: its keys cannot be listed”; the next open tries again); it opened with no keys.find_files_with_suffix_array(),walk_dir_array()andget_ordered_filename_array()answer -1 with the listing empty when an entry cannot be kept (7.25.4 dropped the entry, or kept it as NULL, and answered 0), and the last two also when the ROOT is a directory that cannot be opened (7.25.4: 0 and an empty listing; a root that is not a directory already answered -1). Areaddir()that fails is a failed listing too, in those three and inwalk_dir_tree()(-1, “Cannot list directory, readdir() FAILED” / “Cannot read directory, readdir() FAILED”): 7.25.4 took it for the end of the directory, and timeranger2 could load a key without some of its md2 files. A subdirectory that cannot be opened is skipped, with a WARNING (“Cannot open subdirectory, it is skipped”; 7.25.4 was silent forEACCESandENOENT), only forEACCES,ENOENT,ENOTDIRandELOOP; for a transient cause (EMFILE,ENFILE,ENOMEM,EIO), or when it cannot be read, it fails the walk (7.25.4 skipped it and answered 0 with a short listing), and so does an entry whoselstat()fails for a cause other thanEACCES/ENOENT(7.25.4 logged it and dropped the entry).walk_dir_array()with a NULL pattern lists every entry (7.25.4 crashed), and a root ofwalk_dir_tree()that cannot be opened is always logged (anEACCESroot answered -1 with no log). On a filesystem withoutd_typethe keys are asked withlstat(): a symbolic link inkeys/is not a key (7.25.4 followed it withstat(): a key with another key’s records); anlstat()that fails withEACCEStakes the entry as a key, asDT_DIRdoes, and one that fails for another cause thanENOENTfails the listing of the topic’s keys (“Cannot list the keys of the topic, stat() FAILED”; 7.25.4 dropped the key in silence), so the topic does not open.find_files_with_suffix_array()withoutd_typelists an entry whoselstat()fails withEACCES, as the entry type does (7.25.4 skipped it with no log: a key directory of moder--loaded EMPTY and unflagged).rmrcontentdir()andrmrdir()fail on areaddir()error, with a log (7.25.4:rmrcontentdir()answered 0 in silence, andrmrdir()blamed thermdir()withENOTEMPTY). Teststimeranger2/test_unlistable_dirs,tr_dt_unknown(cases 4 and 5: a link inkeys/, a key directory of moder--),helpers/test_dir_array_nomem,helpers/test_dir_read_error(case 18: anlstat()that fails withEACCESwithoutd_type).A flag goes when its cause goes, on a replica too. A load tries the flags of its key first: it lists an
unlistedkey again when its directory changed since it was flagged (ctime) or opens now after an open failure (INFO “key directory listed again: its files are counted, and the key is not flagged”), and counts a flagged md2 again when it changed since it was flagged (ctime or size) or opens now after an open failure. A flag that looks unchanged is not tried, so a damaged key or file is not read and logged again at every load; a directory whosereaddir()recovers without changing stays flagged until it changes, an append lists it, or the topic is opened again. A replica never appends: its loads are what clear its flags. The notification of an append (the rt_disk follower) lists the key first: its file’s rows are published with their rowids in the whole key (for a key the replica could not list, 7.25.4 counted the notified file alone as the whole key, and the feed got the key’s 4th row with rowid 1), and nothing is read while the directory cannot be listed (“New records of the key not read: its directory cannot be listed”). Teststimeranger2/test_unlistable_dirs(cases 4, 5),timeranger2/test_unlisted_relist_once.A queue backup that FAILED killed the queue (7.25.4 too):
trq_check_backup()/tr2q_check_backup()took the NULL oftranger2_backup_topic()as the queue’s topic and answered 0 -- every read logged “What topic?”, every ack answered -1, no backup happened again.tranger2_backup_topic()opens the topic again after any failure that follows its close (WARNING “Backup of topic failed: the topic is opened again as it was, not backed up”; it was left closed), no longer leakstopic_varon a failedrename(), and logs a missing topic directory (“Cannot back up topic, its directory not found”; it returned in silence). The queues take the topic again and answer -1 (“Queue backup failed: the queue goes on in its topic, not backed up”). A backup whose new topic cannot be created after the move removes what the create left, moves the backup back and opens the topic again (7.25.4 left the queue with no topic, and never tried the backup again). This works in every tranger: the create of a backup runs with the exit bits ofon_critical_erroroff, so the MQTT broker’s queues (LOG_OPT_EXIT_ZERO) no longer exit at the CRITICAL of the failed create, before the backup is moved back, with the messages left in<queue>.bak. When the backup cannot be moved back, a CRITICAL says where the data is (“Backup of topic failed, and the backup cannot be moved back: the data of the topic is in the backup”). A queue whose topic cannot be opened again either (“Queue backup failed, and the queue has no topic”) takes it again by name as soon as it can be opened: at the nexttrq_check_backup()/tr2q_check_backup(), read, ack or load (INFO “Queue topic taken again”); until then those calls fail: only the first one says it (ERROR “Queue without topic, it cannot be opened”), the next ones ask the disk quietly and log nothing (the broker callstr2q_check_backup()every second per session). A topic the tranger has open again (an append opens it by name) is taken as it is. Otherwise the open is tried again only when what stopped it may be gone: atopic_desc.jsonthat loads as no json once it changes (each change tried once, so a broken file is said once as well); a file that cannot be opened or akeys/that cannot be listed (EACCES,EMFILE,EIO) once both can; anything else oncetopic_desc.jsonorkeys/changes. Testtr_queue/test_tr_queue_backup_failed(casestrq_broken,tr2q_broken: 3 open errors each;trq_keys,tr2q_keys: akeys/of mode 0).A new topic is created whole or not at all. When any part of a NEW topic cannot be made (its directory,
topic_desc.json,topic_cols.json,topic_var.json,keys/ordisks/),tranger2_create_topic()removes the directory it made, keeps nothing in memory and answers NULL (CRITICAL “Cannot create topic: it is not whole, what was made is removed”, aton_critical_error, logged after the removal); the next create starts again from nothing. The failures before it are logged without the exit bits, so underLOG_OPT_EXIT_ZERO(C_TRANGER, C_TREEDB, the broker’s queues) the process exits only once what was made is removed. In 7.25.4 each failure was logged and the create went on: a topic without itskeys/ordisks/was opened and returned as created (a replica had nothing to watch), and a queue backup took that half topic as the queue’s new one, answered 0 and left the messages in the.bak. Testtr_queue/test_tr_queue_backup_failed(thekeyscases).
Scans and lists (timeranger2)¶
tmorder is marked: in topics created from now on (or migrated with the newtranger2_mark_tm_order()), a file that receives a record out oftmorder gets a<file>.tm_unorderedmarker, written BEFORE the md2 row (a marker that cannot be written is retried at the next append to that file). A marked file’s tm range is read from all its rows; in a file known to be in tm order a row past the range ends the scan of that FILE (not of the key). A marker name that does not fit makes the file be read whole (logged). In 7.25.4 a reload and a replica took a file’s tm range from its first and last rows: a tm query left out a file whose rows are not in tm order, and its matching rows were missing from the answer. The<file>.unorderedmarker of a late__t__(every topic, as in 7.25.4) is also written BEFORE the md2 row and retried at the next append to the file (mark_file_before_append()); 7.25.4 wrote it after the row, so a process that died between the two left a file whosetgoes back with no marker.tranger2_mark_tm_order()fails, and does not mark the topic, when a key directory cannot be listed.tranger2_list_topic_names()answered[]with no log when the store could not be listed (7.25.4 too): the MQTT broker’slist-queuesandclean-queuesanswered an empty list with 0. It logs (“Cannot list the topics of the store”, with theerrnoof the failure) and answers NULL; the newmark-tm-order all=1answers -1, and so dolist-queues(“<role^name>: cannot list the topics of the store, see the log”) andclean-queues(“<role^name>: cannot list the topics of the store, nothing cleaned, see the log”).list-queues queue=<name>answers -1 for a queue that cannot be opened (“<role^name>: cannot open the queue ‘<name>’, see the log”) and for a queue whose messages cannot all be read (the partial list, with “<role^name>: the messages of the queue ‘<name>’ cannot all be read, the list is PARTIAL, see the log”); 7.25.4 answered 0 with an empty or short list. Testc_mqtt/acl.A scan steps over the holes a tm filter leaves between files (7.25.4 logged a false “next rowids not consecutive” and lost the rows after the hole).
After
delete_keythe iterators of that key drop their segments (7.25.4’s segment stamp could miss a key deleted and written again with the same counts). The delete is announced only once its directory is gone: the key directory is removed first, and a removal that fails announces nothing, reads the key’s cache again from what is left, and indexes the key again in its filtered iterators (ERROR “Cannot index the key again after its files changed: the filtered iterator is empty” when it cannot). A key directory whosestat()fails with anything other thanENOENTanswers -1 (“Cannot delete key, stat() of its directory FAILED”), with nothing deleted, dropped from the cache or announced; 7.25.4 answered 0, dropped the key from the cache and announced it, with its files left on disk. OnlyENOENTis “not found” (logged, 0, as before). The mirror of the delete to the rt_disk feeds (disks/) logs anopendir()orreaddir()that fails (“Cannot tell the rt_disk feeds that a key was deleted, opendir() of disks/ FAILED”, “Cannot tell every rt_disk feed ..., readdir() of disks/ FAILED”); 7.25.4 was silent. The delete itself still answers 0 then. Testtimeranger2/test_delete_key_propagation.rt disk ids: an empty id and a same-creator duplicate are warnings (peer input) like another creator’s;
tranger2_open_iterator()no longer leaks on a bad topic.
Performance, against 7.25.4¶
Measured with the benchmarks under performance/c/ (perf_timeranger2,
perf_tr_treedb, perf_c_treedb, perf_rotatory, new in this release, and
the yev_loop ones) and the test timeranger2/test_topic_pkey_integer, each
linked against the module of 7.25.4 and of this release, run alternated
(means of 8, 10, 14 or 20 rounds, some of them on 4 link layouts;
RelWithDebInfo with memory tracking, ext4, laptop NVMe). The last fixes (keys
that cannot be listed, the delete guard, the schema leftovers, the rotatory
newfile callback; then the instances of a delete, the pkey2 load, the
readdir failures, the schema order, the audit of a JSON text inside a JSON
text; then the pkey2 save guard, the audit’s quote-proof scan, the rotatory
exit_on_fail, the directory walks and the queue backups, and the C_UDP_S
read buffer; then the instances that follow an unlink or a delete of their
parent, the rotatory size limit in bytes, the audit of escaped names, the
yev_loop fd close, and the kw of a subscription that rewrites it; then the
ring entry a stop reads, the newest record of a key, and the match of a
repeated subscription) were measured by an A/B of each change alone, on a
machine shared with other builds; the pkey2 load, the C_UDP_S round trip,
the publish and the subscribe each with a program written for it (none in
the tree). The raw figures, with their
spread, are in performance/c/README.md.
timeranger2.
An append costs what it cost in 7.25.4, with the checks this release adds to it (a flagged file, the tm order, a torn md2 tail, a stopped master). The append looks up the cache of its key once, and it searches the cache cell of its file once;
cmp_file_ids()compares file ids in place, withoutsnprintf()(new testtimeranger2/test_cmp_file_ids).perf_timeranger2, 20 rounds: 400 000 appends into 20 000 files 1733.5 -> 1712.6 ms (-1.2%), 600 000 appends into 30 files of a topic that marks tm order 1635.0 -> 1627.3 ms (-0.5%).test_topic_pkey_integer, appends/s without / with an rt list, averaged over 4 code layouts (n = 80): 7.25.4 220 120 / 166 498, now 221 588 / 167 381 (+0.7% / +0.5%). A single link of each moves up to 2.5% with the address of the libraries after the module (jansson’sjson_dumps()is ~40% of an append): seeperformance/c/README.md.A master opens a store 13% faster (20 000 md2 files: 92 ms -> 81 ms): it finds the order markers in the listing of the key directory it already reads, instead of a
stat()per md2 file.Price kept, by decision: a replica’s open is ~19% slower (90 ms -> 106 ms, one more
stat()per md2 file): it looks for the markers on disk after reading each file, because a listing taken before could miss a marker the master writes during the open. Creating a topic makes 2 fsyncs and atopic_versionchange 4 (7.25.4: none; 10 topics created in 119 ms against 3.6 ms, 10 version changes in 232 ms against 1.7 ms); opening an existing store makes none. That is the price of durable topic files: a power cut never leaves a newtopic_versionover atopic_cols.jsonthat is not on disk.Price until migrated: a
tmquery on a topic created by 7.25.4 or earlier reads every md2 row of the key untilmark-tm-orderruns (one minute of 1 key x 30 files x 20 000 rows: 13 ms on 7.25.4, ~392 ms unmigrated, ~7 ms migrated). See “Upgrade steps”.The last fixes (a key or a topic that cannot be listed, a failed backup, the retry of a flag at each load) move nothing out of the noise. 10 alternated rounds, before -> after: 400 000 appends 1927 -> 1874 ms (-2.8%), the master open 89.1 -> 87.1 ms (-2.2%), the replica open 117.1 -> 116.0 ms (-0.9%); the retry of the flags, 6 rounds: -2.8% .. +1.4% on the open, tm query, append and
mark_tm_ordercases. With nothing flagged it costs two dict lookups per iterator open and one per notification. The fixes after them (a faileddelete_keythat indexes its iterators again, a key listed again only when its directory changed, a backup whose create fails moved back), 8 alternated rounds: 400 000 appends 1517.6 -> 1526.8 ms (+0.6%), the master open 83.1 -> 83.3 ms (+0.2%), the replica open 136.5 -> 136.2 ms (-0.2%), noise. The whole-or-nothing create came after that run; it changes only the failure paths of a create. The directory walks that fail on a long path or a transient error, and the queue that takes its topic again (8 alternated rounds): the master open 81.17 -> 81.90 ms (+0.9%), the replica open 106.47 -> 106.36 ms (-0.1%), noise.
treedb writes (
perf_tr_treedb, 100 000 operations, CPU time, 4 link layouts): an update in memory 3.92 -> 2.92 us, a saved update 11.75 -> 10.81 us, link+unlink 11.27 -> 11.09 us. An update deep-copied the node, with every child its hooks hold, for the system schema’s check on every treedb; now only on the system schema. A forced delete of a parent with 200 children is ~6% faster (2952 -> 2781 us): it tells its children’s events itself instead of holding two per child. The last fix of the delete guard takes 2-4% more off it: it counts the children instead of building a reference string for each (A/B of that change alone, 12 rounds x 4 layouts: median 3139 -> 3076 us, -2.0%; paired mean -3.9% +- 1.4%). Price: a forced delete of a leaf costs ~1.7% more (53.1 -> 54.1 us, ~1 us): the hold of its events and the kept column that let a refused delete change nothing (the guard fix: -0.5% +- 0.8%, noise). The delete that looks at every instance of its key, and the save guard (A/B of that change alone, 12 rounds x 4 layouts, mean of the 48 paired ratios): -0.1% .. +3.5%, the medians under 1% on the save path and +2.5% ondelete_force(one array of the key’s instances per delete), inside a spread of 8-13% per pair: noise. The save guard of a pkey2 value changed in place (A/B of that change alone, 20 alternated rounds, medians): a saved update 10.805 -> 10.836 us (+0.3%), link+unlink -0.4%,create_link_half-0.7%. The benchmark’s topics have no pkey2, so the guard itself does not run there; the pkey2 list it needs is fetched once per save and reused. The unlink and the delete that clear the parent’s ref in every instance of the child key (A/B of that change alone, 12 alternated rounds, paired change): -3.8% .. +1.7%, no loss (a saved update 10.61 -> 10.21 us,delete_parent3040 -> 2944 us); the benchmark’s topics have no pkey2, and the new scans run only for a child topic that has one. The newest record of each key kept on the instance that wrote it (A/B of that change alone, 12 alternated rounds pinned to one CPU, us of CPU an operation, mean +- standard deviation): a saved update 10.955 +- 0.11 -> 10.947 +- 0.53 (-0.1%),delete_force73.07 +- 4.1 -> 74.53 +- 5.4 (+2.0%),delete_parent3308 +- 409 -> 3210 +- 134 (-3.0%): noise. The benchmark’s topics have no pkey2, so there the new code is one lookup per child.A pkey2 topic opens faster: the load puts the primary in its own pkey2 slot instead of building a copy of it. 4000 keys x 3 instances, 12 alternated rounds: a reopen 353.5 -> 327.2 ms (-7.4%), an update through
treedb_get_instance()29.4 -> 28.6 us (-2.7%).perf_tr_treedb, whose topics have no pkey2, moves -1.5% .. +0.5% (8 rounds x 4 layouts, noise).A JSON file is read whole, then parsed (
load_json_from_file(),load_persistent_json(), and tr2migrate):json_loadfd()made oneread()per byte, 60 000 system calls for a 60 KB schema.C_TREEDB opens (
perf_c_treedb, 40 treedbs x 10 topics x 20 columns, seconds for 40 opens, 10 alternated rounds, mean +- standard deviation). Against the C_TREEDB of 7.25.4 linked with the libraries of this release, so both on the new JSON load: the same literal 0.672 +- 0.071 -> 0.613 +- 0.020 (-9%), the first projection (seed) 9.87 +- 0.42 -> 10.84 +- 0.32 (+10%), a newer literal 12.79 +- 0.44 -> 13.89 +- 0.53 (+9%). 7.25.4 with its own JSON load took 1.84 s for the same literal. An open reads only the nodes of__system__of its own treedb, and parses the literal once: 7.25.4 parsed the same object a second time withparse_schema()(~2.9 ms an open), a check left from the releases that read the schema back from__system__(a literal that fails now logs its bad column once, not twice). The id-collision check collects the ids only and describes the two elements when there is a collision.Price kept, by decision: the last two schema fixes (a literal with other content under the file’s
schema_versionsaid at every open, and the move of a rowid-keyed projection before the upgrade record) cost the same literal ~0.9 ms an open: oneschema_diff()of the literal against the file and one more lookup oftreedbs(A/B of that change alone, 8 rounds, ms an open: 15.33 +- 1.46 -> 16.20 +- 0.66, +5.7%; seed +0.3%, newer literal +1.0%, noise). The same literal is then ~0.648 s for 40 opens, ~4% below 7.25.4 (not measured in one run).Price kept, by decision: what an open costs over 7.25.4, measured with probes in the code (ms an open; the probes cover only the steps listed):
The same literal, ~0.9 ms, paid by the removed second parse (net -2 ms, ~-1.1 ms with the ~0.9 ms of the last fixes above): C_TREEDB reads and parses the schema file in use (~0.7 ms for a 60 KB file), because what runs is decided against the file, as
treedb_open_db()decides it, which reads the file again as in 7.25.4; the id-collision check (~0.1 ms); the records of an apply and of an unfinished projection (~0.04 ms).A newer literal, ~25 ms (net ~22 ms): the record of the projection in progress, written whole with its fsyncs before the first write (~14 ms); the index of the treedb’s nodes of
__system__and the search for orphans (~5 ms); the drafts,__system__compared with the file in use (~3 ms); the ownership checks of the projection (~2 ms); the schema file in use read (~1.3 ms).The seed, ~19 ms (net ~16.5 ms): the same record (~12.5 ms); the index of
__system__, the orphans and the ownership checks (~6.6 ms). The benchmark says more than the probes: +24.3 ms an open for the seed and +27.5 ms for a newer literal ((10.84 - 9.87) / 40, (13.89 - 12.79) / 40), so ~6-8 ms an open is spent outside the steps the probes cover. The record is the crash safety of the projection; the index, the orphans and the drafts keep an operator’s draft from being taken for what a projection left, and the reverse.
The schema fixes after them (the order of the file in a save, the positions a save writes, a topic raised past what runs, the imposed tie said), 8 alternated rounds on a machine busy with other builds, seconds for 40 opens: the same literal 0.835 +- 0.135 -> 0.857 +- 0.140 (+2.6%), the seed 2.786 +- 0.270 -> 2.913 +- 0.286 (+4.6%), a newer literal 3.012 +- 0.552 -> 2.836 +- 0.252 (-5.8%). Each difference is inside a spread of 9-18%: noise. The only new work on the same-literal path is one
topic_var.jsonread per topic (and itstopic_cols.jsonfor a topic that runs ahead of its file). The absolute figures differ from the ones above: another day, another state of the machine.rotatory (
perf_rotatory, ns for one agent audit record of 300 bytes, 8 rounds): 6 194 -> 554 (see “Agent, gobj-c and tools”); flushed after each record, as the agent audit does now, 7 623 -> 1 475. The last rotatory fixes (the newfile callback for a new file only, the keep_all retry, the fixed names) move nothing: 12 alternated rounds, +0.4% .. +2.3%, each within one standard deviation. Nor do the ones after them (the callback kept when a new file fails to open, no size rotation at a new name, the retry on the monotonic clock): 10 alternated rounds, the medians ofaudit_record670 -> 640 ns,audit_record_flush1 476 -> 1 474 ns,log_record432 -> 438 ns. The audit of a JSON text inside a JSON text costs +0.22 us a record (6.32 -> 6.54 us to build and write it, +3.5%), only when the kw holds an escaped JSON text; the other records do not move (list-yunos1.63 -> 1.64 us,run-yuno5.34 -> 5.15 us, a kw with a plain JSON text 6.31 -> 6.33 us). The scan that finds such a text whatever quotes come before it (6 alternated runs, us a record): the escaped JSON text 8.28 -> 8.41 (+1.6%), a plain JSON text 8.11 -> 8.18 (+0.9%),run-yuno6.83 -> 6.83,list-yunos1.60 -> 1.54 (-3.8%). A handle whose later opens are printed instead of ending the process (exit_on_fail, 8 alternated rounds, ns a record):audit_record553 -> 547 (-1.1%),audit_record_retention547 -> 545 (-0.4%),audit_record_flush1 280 -> 1 265 (-1.2%),log_record386 -> 387 (+0.3%): only the failure paths moved. The size limit compared in bytes (8 alternated rounds, ns a record):audit_record555 -> 574 (+3.4%, its median +0.9%: three slow rounds),audit_record_retention564 -> 565 (+0.1%),audit_record_flush1 310 -> 1 300 (-0.8%),log_record391 -> 388 (-0.8%), noise. The audit that reads an escaped key as its whole string and a secret value as its whole shell word (10 alternated rounds, us a record): +1.1% .. +1.4%, andlist-yunos+6.9% from one outlier round (its median -1.3%). Price: the audit that decodes the JSON escapes of the names and looks for a write-attr in every string of the kw costs every record +2-3% (8 alternated rounds, us a record:list-yunos1.51 -> 1.55,run-yuno6.44 -> 6.57, a kw with a JSON text 7.74 -> 7.98, an escaped JSON text 8.05 -> 8.29).yev_loop:
perf_yev_ping_pong-1.1%,perf_tcp_test4+1.6%,perf_tcp_test5-0.8% (8 rounds, noise); the last fixes (a take-back that wakes the loop, a stop without memory) -0.4% onperf_yev_ping_pong(6 rounds, noise), and thesrc_urlof another family -0.2% (148.7 -> 148.4 K msg/s, 6 rounds, noise). C_UDP_S has no benchmark in the tree; the read that takes a new gbuffer when the host keeps the received one was measured with an echo server and a client written for it (64-byte datagrams, 20 000 round trips a round, 8 alternated rounds): 11.88 +- 0.77 -> 11.59 +- 0.35 us a round trip (-2.4%), noise. The close of an fd that takes back the submissions of other events on it (8 alternated rounds):perf_yev_ping_pong137.8 +- 5.0 -> 140.5 +- 4.6 K msg/s (+2.0%), noise. The entry of the ring cleared when it is handed out, and the cancel of a running stop prepared before the fd closes (6 alternated rounds):perf_yev_ping_pongmedian 147.7 -> 147.9 K msg/s (+0.1%; one round of 123.0 among 147-153), noise.Price kept, by decision: a subscription that rewrites the kw gets its own.
gobj_publish_event()gives a subscription with__local__or__global__akw_twin()of the event (its own top level, the values and gbuffers shared); every other subscriber shares the publisher’s kw, as in 7.25.4 (see “Security”). No benchmark in the tree; one written for it, onlygobj.oswapped, us a publish (8 alternated rounds at ~250 bytes, 16 at ~20 KB), 7.25.4 -> now:no
__local__/__global__: 1.78 -> 1.78, 3.20 -> 3.21 and 19.86 -> 18.86 (1, 10, 100 subscribers, 250 B); 63.03 -> 62.76, 64.98 -> 64.22 and 79.59 -> 81.28 (20 KB): -5.0% .. +2.1%, noise.__global__on every subscription: 1.95 -> 2.08 (+6.6%), 3.71 -> 5.93 (+59.9%), 21.76 -> 44.94 (+106.5%) at 250 B; 64.37 -> 63.88 (-0.7%), 65.65 -> 67.54 (+2.9%), 82.44 -> 104.93 (+27.3%) at 20 KB. A twinned delivery costs ~0.23 us whatever the size of the event (a deep copy of the kw would cost ~1.8 us at 250 B and ~55 us at 20 KB). Every C_IEVENT_SRV subscription carries__global__(the gate’s back-metadata), so a remote subscriber pays it, beside the serialization of the kw it already paid.
A subscribe matches the kw as it is stored (the fix of a repeated
__own_event__/__rename_event_name__subscription). No benchmark in the tree; one written for it, onlygobj.oswapped, 4 alternated rounds, medians of a subscribe and an unsubscribe, us an operation, before -> after that change: 2000 subscriptions 167.7 -> 158.4 (-5.5%), with__config__226.0 -> 218.6 (-3.3%); 100 subscriptions 8.63 -> 8.47 (-1.9%), with__config__11.24 -> 10.69 (-4.8%). The publish of the same program moves -3.4% .. +2.7% (medians), noise. The gate’s own subscribe adds one scan of the channel’s subscriptions (at mostmax_subscriptions), and stores a smaller back-metadata, which every event of the subscription copies: not measured.ctest times (
build/*.txt, 52 runs ofyunetas testfrom 2026-09-23 10:20 to 2026-09-25 12:26): of the tests whose code did not change after 7.25.4, onlytest_treedb_schema_fidelitymoved more than 10%, 1.14-1.26 s (the 4 runs before the change) -> 1.55-1.81 s (the 48 after it; +36% .. +44%, median 1.19 -> 1.66 s, +39%). It opens four treedbs in four new stores, and its fsyncs (88, ~0.56 s) are the price of the durable topic files and of the projection record. Thec_tcp,c_tcp2andc_tcpstests wait on whole seconds of their retry timers and moved by whole seconds in every period (c_tcp/test29-15 s,c_tcp/test32-4 s,c_tcps/test112-14 s): noise.timeranger2/test_topic_pkey_integerstays at 1.84-2.12 s.
Schemas (C_TREEDB)¶
A newer C literal wins whole (dynamic-schema treedb, impose off). As in 7.25.4, a literal with a
schema_versionhigher than the schema file in use replaces the whole file; an equal or older literal is not installed and the file runs. New:__system__is projected from an installed literal whole -- topics, columns and attributes the literal no longer declares are removed from__system__too (7.25.4 only added and updated). What the literal discards is said in one WARNING (“Schema from C withdrew work on the schema at open”) and exposed aswithdrawn_at_openper topic intreedbsrows andsaved-schemauntil the next open:applied(an applied schema never opened),in_use(an applied schema that ran),saved,unsaved(drafts), andleft_by_older_release(below). A topic whose columns the literal changes without raising itstopic_versionkeeps running its old columns in the store (tranger2 swaps columns only on a version raise): the open warns.A projection says when it is unfinished, and records it. A removal that is refused (a snapshot of
__system__holds the node), or a create, update or link that fails, leaves the projection unfinished: one WARNING withnot_removed/not_writtenand how to finish it (for a removal, delete the__snaps__row), and a record,saved_schemas/<treedb>.unfinished.json(schema_version,not_removed,not_written,leftovers,draft_kinds,leftover_nodes). While the record exists every open retries the projection (INFO “Completing the projection into system, left unfinished by an earlier open” when the file runs, otherwise “Updating TreeDB schema in system”: a newer literal installed, or an imposed one), the leftovers it names, while they stay as the projection left them, are never taken for an operator’s draft nor reported as withdrawn work, a newunfinished_projectionfield says it intreedbsandsaved-schema, andsave-schemarefuses (it would publish the removed topic again). A projection that completes removes the record; so doesdelete-treedb. The record is written whole (one buffer, a.newfile, an fsync and a rename, as the record of an apply); one that cannot be read still means “unfinished” (one WARNING at each read, with its cause inerror;save-schemarefused; retried), and then what__system__holds over the file is reported as anunsaveddraft when the projection completes. A record that cannot be written is kept in memory and written again at every open, and the treedb node saysc_schema_version: -1; after a restart the record is “lost” (a WARNING,save-schemarefuses, and what__system__holds over the file is reported, never deleted in silence). The record keepsdraft_kinds, so the open that finally replaces a draft reports its kind (a saved draft issaved). A treedbs node never stamped (schema_version0, no record, a file in use atschema_version1 or more) is a seed that died: the next open completes it and reports nothing.A draft is reported once, by the open that replaces it in
__system__, whether it was made before the projection failed or while it was unfinished. The part of a draft a projection could not replace stays a draft (indraft_changed), never a leftover; a column delete that fails after its topic was removed keeps its draft. A topic the operator deleted from__system__is a draft too: when a newer literal re-creates it, it is reported asunsaved, orsavedwhen a pending save published the deletion. A topic the operator added and did not save isunsaved, also when a save of another topic is pending. An EDIT of a leftover is operator work too: the record keepsleftover_nodes(what the projection left at each leftover), a leftover that differs from it (edited, unlinked, deleted, moved to another topic, linked to one more, or left in no topic) is a draft (draft_changed), and the open that replaces it reports it (a move reports both topics); an unedited leftover is not reported. The record keeps only the attributes a projection writes, andsystem_schema_version: an id whose write failed (not_written) is taken as left, and a record kept under another meta-schema takes every leftover as left, with a WARNING (“Leftovers of an unfinished projection were kept under another meta-schema: every leftover is taken as left, an edit of one made meanwhile is not told apart”).A projection that dies half way invents no operator work. Before its first write a projection plans every write and records the plan (the record of an unfinished projection, with
"in_progress": true:planned,target_nodes,replaced_kinds); the projection that completes removes it after its WARNING. A process killed between two writes always leaves that record: the next open takes a node that is as it was, or as the projection writes it, for the projection’s, completes the projection, and reports the drafts it was replacing exactly once, also when the retry is killed too (tests:c_treedb_literal_wins, scenarios CR and DC, a kill at every write). Until thenunfinished_projectionlistsplannedandsave-schemarefuses. The move of a rowid-keyed projection (before 7.13.2) to qualified ids runs first, before the upgrade record is built, so the record names what an older release left at its qualified ids (scenario LQ). It can die at any write and completes at the next open; a move with a failed write says so (WARNING “TreeDB schema ids moved to qualified names only in part: nothing is projected at this open, every open retries the move, save-schema refuses until then (see the errors before this)”), projects nothing at that open, records the projection as unfinished, and the next open moves the rest (scenario LX).A projection that 7.25.4 stamped first is completed, and judged node by node. 7.25.4 and earlier wrote the stamp of the treedb node first; a process that died after it left a projection that said it was complete, and it was taken as done at every open. Now it is completed, with
impose_c_schemaand without it: a column the dead projection did not write gets the literal’s content, and only a NODE (a topic or a column) that differs from both the old file and the literal is the operator’s draft -- a topic with one column written and one not is not reported. A node in no tree that is what the literal writes, and that the old file does not declare, belongs to the dead projection (a create whose link never happened). The record of the completing projection keeps the literal (stamped_base), so a retry after a second crash compares with both again (tests:c_treedb_literal_wins, scenarios SE, SH, SI, SJ, SK, O4C, O4A, O4X).The first open by this release reads what an older release left. New record
saved_schemas/<treedb>.upgrade.json. A first projection that 7.23.0-7.25.4 stamped over a file at the literal’s version and did not finish is completed. When the file in use IS the literal, what__system__misses is restored from it, with one WARNING “Restored from the schema from C: ...” naming the ids; when that open installs a newer literal, what__system__misses of the old file is no deletion of the operator’s (INFO “What system misses of the schema file in use is no deletion of the operator’s: ...”, with the ids) and the literal is projected. Nothing is reported as the operator’s. After that first open a stamped projection is complete, and a topic the operator deletes stays a draft. A topic or column that 7.25.4 kept after a literal removed it is no draft: the open that removes it reports it inwithdrawn_at_openas kindleft_by_older_release, with one WARNING “Removed from system what an older release left there: ...”.save-schemaleaves such a node out of what it compares and writes, and names its id indata.left_by_older_release(with nothing else changed it answers “nothing to save”); an edit after the upgrade makes it the operator’s, and then it is saved (tests:c_treedb_literal_wins, scenarios FP, OL, OS).A
delete-treedbcut half way is finished by the next one. Columns first, then topics, then the nodes in no tree, the treedb node last; with the treedb node already gone the nodes left in no tree are still deleted (7.25.4 deleted the treedb node first, and a delete cut after that answered -1 “not projected in system” at every run). It answers 0 with the ids indata.deleted; nothing left also answers 0. The upgrade record is removed too (test:c_treedb_literal_wins, scenario DTK).A topic or column that the treedb’s tree no longer reaches (an operator’s unlink, or a projection link that failed) does not block a projection with “Node already exists”: a schema that declares it takes it over, otherwise it is removed (INFO “Node of the treedb that its tree does not reach: removed from system”); when it is operator work its topic is reported
unsaved.Which treedb a node of
__system__belongs to is read from the node, never from a prefix of its id (a treedbm2would take the nodes of a treedbm2.b). A column whose topic node is gone and that more than one treedb could own (m2andm2.b) belongs to the treedb whose record or schema names it, or to the treedb projected now when none does; only that treedb warns, once per node; no node is left for ever.delete-treedbdeletes every node of its treedb (a column moved to another of its topics and its nodes in no topic included) and only those: a node of another treedb, or of another topic, linked into it is only unlinked.A column the operator moved to another topic, or linked to a second one, is put back where the literal declares it in ONE open. A column linked to a topic that does not declare it is only unlinked when the literal declares it elsewhere or it belongs to another treedb; a topic of another treedb linked here is unlinked (INFO “Topic of another treedb, not declared by the schema from C: unlinked from the treedb in system”).
A topic whose write or link fails leaves what it holds untouched: the columns of a new topic whose link failed are not written under it, and a take-over or create that fails deletes no orphan node of the declared topic and reports it once.
A schema whose elements get one qualified id in
__system__(a name with a dot: the columnx.yofuand the columnyofu.x; the topicb.cofmand the topiccofm.b) is refused with ONE ERROR naming both: the treedb does not open,save-schemadoes not save it,apply-schemadoes not apply it. No id changes.save-schemarefuses a draft with no topics andapply-schemaa saved schema with no topics (WARNING, -1, the file in use unchanged): a treedb without topics does not open.saved-schemaanswerscan_apply: false, with the reason, for such a saved schema (7.25.4 answered true and applied it).The numbers of a treedb’s node in
__system__(schema_version,c_schema_version) are written last, only when the whole projection succeeded; a new node is created with them at 0 (7.25.4 wrote them first, so a process that died half way left a projection that said it was complete).A second
open-treedbof an open treedb is refused first (“already open here: close-treedb first, nothing was changed”); 7.25.4 reconciled__system__first and then failed with “Internal error, tranger client NULL”, and with this release’s whole projection that reconcile would delete topics and withdraw the saved schema. An open that fails before its treedb starts (no valid schema, a collision, no C_NODE) destroys the tranger it created; when the schema is refused, its services stay untilclose-treedb. When the treedb’s schema is refused (treedb_open_db()fails) the answer is -1 (“did not open, its schema was refused (see the log): close-treedb it before opening it again”); 7.25.4 answered 0 “Treedb opened!”. Untilclose-treedb, a second open of it answers -1 (“did not open at its last open-treedb...”) and itstreedbsrow saysopened: false.delete-treedbof it answers that it did not open and to close-treedb it first (7.25.4: “while it is OPEN”). C_NODE gives a treedb thattreedb_open_db()refused no callback and does not close it at stop, so the failed open logs only its cause andclose-treedblogs nothing (7.25.4 logged “TreeDB not found” at the open and twice at the close, with stacks); a write ofwith_link_eventson it sets no callback either. Every answer ofopen-treedb,close-treedbanddelete-treedbstarts with the yuno.The agent stops when its treedb does not open. The agent opens its treedb (
treedb_yuneta_agent) with its schema imposed and exits 0 on a -1 answer (its log has the answer, syslog “Cannot start agent treedb: ...”); ydaemon does not relaunch an exit 0, so the main agent stays down until started again, andyuneta_agent22(which opens no treedb) stays up as the way in. 7.25.4 answered 0 when the schema was refused, and the agent ran without its treedb.The client store decides. If another process holds the client store’s lock, the treedb opens as a replica and nothing is reconciled (INFO).
A literal with another content under the file’s
schema_versionis said at every open (WARNING “Schema from C has the schema_version of the dynamic schema in use but another content: NOT applied, raise its schema_version to publish it”, with thediff), withimpose_c_schematoo (the default: an imposed tie runs the file, and the WARNING carries"imposed": 1). 7.25.4 compared only when__system__'sc_schema_versionwas another number, and never when imposing: a column added to the literal without raisingschema_version, the classic mistake, reached nothing and nothing was said (scenarios SV, SVI). A seed of__system__at an imposed tie is made from the file that runs, so the literal’s difference is not read as the operator’s draft.The saved schema is written whole, through
<treedb>.treedb_schema.json.new, an fsync and a rename: a save that died or found the disk full left a torn file where the pending save had been (7.25.4 too; scenario SW).save-schema/saved-schemawith notreedb_namereleased agbuffercarried in the command kw once per treedb (“BAD gbuf_decref()”, and the caller’s buffer freed; 7.25.4 too). Each treedb gets a copy of the kw with its own reference (scenario GB).EV_OPEN_TREEDB/EV_CLOSE_TREEDBare public events that nothing implements: they did nothing in silence (7.25.4 too). Each logs an ERROR (“EV_OPEN_TREEDB is not implemented: nothing opened, use the command open-treedb”, and the same for close) and answers -1; use the commands (scenario EV).A command that writes through a STOPPED tranger is refused. While the
__system__tranger is stopped (between a stop of C_TREEDB and its next start),save-schemaanddelete-treedb; while a treedb’s own tranger is stopped,create-topicanddelete-topicof it -- all with “<role^name>: treedb ‘<db>’ is STOPPED: its tranger holds no lock until its service starts again” -- andapply-schemaof it (“<role^name>: treedb ‘<db>’ is READ-ONLY, this yuno is not the master of its tranger, or the tranger is stopped”). 7.25.4 read themasterthat the stopped tranger still said:save-schemaread nothing and answered “nothing to save”, anddelete-treedbwent on to delete nodes of__system__.apply-schemarecords what it put in use in a new file,saved_schemas/<treedb>.applied.json:{"schema_version": N, "topics": {"<topic>": "applied"|"in_use"}}, with the record it replaced inprevious. It lives as long as its file is in use. It is written before the rename;apply-schemarefuses when it cannot write it, and restores the previous record when the rename fails.A seed from the schema file stamps
c_schema_versionwith the literal’s version only when the file IS the literal, 0 otherwise. A seed, or the completion of an unfinished projection, reads a file whosetopicsare a list or a dict keyed by name (what a node opened withimpose_c_schemaoff wrote before 7.25.0; scenario DF).A saved draft that is reverted is withdrawn by the next
save-schema(withdrawnin the answer). A saved schema lives only as long as the file it was saved against: applied, withdrawn or superseded files are removed;savednow means “waiting to be applied”,stalemarks an old file,brokenan unreadable one (left out of an apply of every treedb).schema_versionin__system__never goes down;apply-schemawrites no derivedfkeymarks into the file in use; a reorder of topics or of columns is a difference (only when the names both sides declare change order; thesaved-schemadiff carries a__topics_order__row and a__cols_order__row per topic). A storedorderthat says nothing about the node’s place -- absent, or the default 9999 that a node of a projection written before 7.14.0 (every one keyed by rowid, and the qualified ones of 7.13.2) loads with -- is no difference (diff-schemaincluded): such a projection would otherwise read as a draft of every topic. A projection that rewrites the node writes its position, and a reorder the operator makes is still a draft (scenario LQ); a save places such a node as the next bullet says. The INFO “TreeDB schema from C is behind the schema in use, not applied” is logged only when the literal is behind the FILE in use (7.25.4 also logged it for a literal equal to the file while a save was pending, and on an imposed open of a literal behind__system__); an imposed literal behind__system__logs “TreeDB schema from C is imposed, but it is behind system: the projection is kept”.A save keeps the order of the schema file in use. A node whose
ordersays nothing about its place (the default 9999 of every projection written before 7.14.0, or of a column made by hand) was sorted as if 9999 were a position (7.25.4 too): such nodes tied and kept the order the store loaded them in, alphabetical by id, so a save published a reorder of every such topic and column that nobody made, andapply-schemaput it in the file. Such a node now goes where the schema file in use declares it, then where the literal declares it, then last (scenario LQ).A save writes the places too. A reorder is now a difference (above), so
save-schemawrites into__system__the position each topic and column has in the saved schema, where it differs. A column added with anorderthat is not its position (e.g. 99) therefore does not read as unsaved after its save, nor as a draft after the apply. Each topic and column is found by name, as the diff finds it, and written under its own id: a column the operator moved to another topic keeps its old id, and a topic of another treedb linked here keeps its treedb’s. Thetopic_versiona save writes is found the same way (7.25.4 wrote it under the id composed from the names, which missed such a topic). A name of the saved schema that__system__does not hold is an ERROR (“Topic of a saved schema not found in system, its place is not written”, and the same for a column and for thetopic_version): the schema is built from that same tree (scenarios OP, OM). The same save publishes every topic whose saved place is not the file’s: a node placed in front of its siblings shifts them (withorder5 written onusers, the draft isdepartments, users), and the save publishesusersanddepartments, each with anorderrow inchanges(scenario OX).saved-schemasays it before the save: itsdraft_changedis{"users": true, "departments": true}, so the schema editor’s draft marks and Save/Apply agree with what the save publishes. A node that hangs from more than one parent (the meta-schema’s fkeys are lists: a column linked to a second topic, a topic of another treedb linked in and still in its own) gets no place of its own, since oneordercannot say a place in each parent. The save writes itsorderas 9999 (says nothing) and names it indata.places_not_written. The draft places it by the file in use, then the schema from C, then last, and the diff does not compare itsorder. So a node that STOPS being shared (unlinked from one parent: the second step of moving a topic from one treedb to another) goes where the file of the parent that remains put it, and that parent, edited by nobody, reads no draft (scenarios SC, ST).Two topics with one name in one treedb are refused (7.25.4 too). A topic of another treedb linked in under a name the treedb has made two topics of one name; the schema is keyed by name, so the rebuild kept the last and the diff found the first, and a save published the other treedb’s topic, at a
topic_versionit did not raise, with no word.treedb_link_nodes()refuses that link (“Treedb already has a topic with this name”, withidandsibling_id), as it already refused two columns of one name in a topic; it still takes the child that is already linked there, and one keyed<parent id>.<name>(the move of a rowid-keyed projection). What the link does not see (an autolink update, an older store) is refused bysave-schema, a dry run included, and by the rebuild of the draft: ONE ERROR naming both ids (“Schema refused: two topics of the treedb in system have the same name, unlink or rename one of them”, and the same for two columns of a topic), and -1 withdata.twins{what, treedb_name, topic_name, name, first, second}(scenario DT).A failed
open-treedbkeeps the saved schema. An open that installs a newer literal withdraws the pending saved schema (withdrawn_at_open, above) only when the treedb OPENS, so a literal that is refused does not take the operator’s save with it. The open that does not open keeps the file, with an INFO (“Saved schema kept: the open that installs the schema from C did not open; the next one that installs a schema from C and opens withdraws it”), andwithdrawn_at_open.saved_schema_versionis 0. While the refused literal is the file in use the kept save readsstale(treedb_open_db()writes the literal over the file before it checks the rest); with the good file back and the literal rolled back, it is pending again (scenario KS).A save publishes a topic past what runs. A store can run a topic at a
topic_versionabove the one of its schema file (a file written whole over a topic that an apply had raised, by an older release or by a newer literal installed whole).save-schemagave the topic the file’s version plus one, which was not above what ran, so the apply reached nothing and nothing said so (7.25.4 too). It now raises the topic past the higher of the two, and every open that runs such a file says it, per topic (WARNING “Schema file in use declares other columns than the store runs, at a topic_version behind the store’s ...”, withtopic_version,running_versionandpath, the directory whosetopic_cols.jsonsays what runs).__system__holds the file’s columns, so a save with no edit has nothing to save: its answer names such a topic in the comment and indata.store_ahead, which both “nothing to save” answers ofsave-schemacarry ({}when there is none), for example"store_ahead": {"users": {"topic_version": 1, "running_version": 2, "path": "..."}}. To keep what runs, edit the topic in__system__to those columns, thensave-schemaandapply-schema. To run the file’s, raise itstopic_versionaboverunning_version, and theschema_version, in the schema from C (scenario BH).
Event loop (yev_loop)¶
A full submission queue ended the process in 7.25.4: of the 12 places that ask
io_uring_get_sqe()for an entry, 10 used the NULL entry (a crash) and 2 logged “io_uring_get_sqe() FAILED” and aborted. The queue is flushed and the entry asked again; when the kernel still takes nothing (on an older kernel a CQ overflow answersEBUSYuntil the completions are reaped), the submission is KEPT and made at the next cycle of the loop, in order (a timer is armed, a stop reaches its event). A stop of a submission still kept reaches the callback as a cancel (STOPPED,-ECANCELED). One WARNING when the loop starts keeping: “Submission queue full and the kernel takes nothing: kept for the next cycle”. Only a submission that there is no memory to keep fails: CRITICAL “No memory to keep a submission”, then the caller answers -1 with its ERROR “No memory to keep a submission: <what did not happen>” (every event start -- connect, accept, write, read, sendmsg, recvmsg, poll -- “...: event NOT started”; a timer “...: timer NOT started”; a cancel “...: event NOT canceled”; a re-arm inside the loop “...: accept event NOT re-armed” / “...: periodic timer NOT re-armed”, with no caller to answer; a loop stop “...: loop NOT stopped”). Testyev_events/test_yevent_sq_full; the hot path does not move (perf_yev_ping_pong -1.1%, perf_tcp_test4 +1.6%, perf_tcp_test5 -0.8%: within noise; figures inperformance/c/README.md).A submission the kernel did not take is no longer left waiting: a failed
io_uring_submit()can leave its entry in the queue (the submit that hands over the kept entries included); the loop submits those entries again at each cycle, with a 10 ms retry wait. One WARNING when it starts (“Submissions the kernel did not take: submitted again at each cycle”), an ERROR after 100 cycles (“Submissions not taken by the kernel for many cycles of the loop: their operations wait”), and an INFO (“Submissions taken by the kernel again”) when the kernel takes them again after that ERROR.A stop of an event whose submission is still in the queue takes it back, as for a kept one: the entry is found by its event (never by its fd) and becomes a NOP whose completion is not delivered; the callback gets
STOPPED,-ECANCELED. Otherwise the stale operation ran later on the fd the stop had closed, and a new timer that got the same fd number never fired. The loop does not block on the ring while it holds the completion of such a take-back, also when a posted action (gobj_post_event()) made the stop: theSTOPPEDcomes at the next cycle, not with an unrelated completion (testyev_events/test_yevent_kept_after_post). A stop that has no entry for its cancel and no memory to keep it answers -1 and gives up nothing: the event keeps its gbuffer and the fd of a timer or connect, and a read then completes with its data (testyev_events/test_yevent_stop_nomem).A closed fd takes back the submissions of other events on it. A write still in the submission queue after a failed
io_uring_submit()went to the kernel at the next submit that worked (7.25.4 too), and so did a kept submission: when the event that owns the fd (a connect, an accept, a timer) was stopped or destroyed and the fd closed, the operation ran on whatever file reused the number -- the bytes of the old connection went to the new one. Before theclose(), the loop now takes back every submission of another event on that fd that the kernel has not taken, in the kept list and in the queue, and that event completes asSTOPPED,-ECANCELED, at the next cycle, with a WARNING (“An fd is closed with submissions of other events on it that the kernel did not take: taken back, completed as canceled”); with no memory for that completion they are dropped all the same, with an ERROR (“...: dropped, the event will not complete”). An entry of the queue is told by the event that owns it: every entry the loop hands out has its fd and its owner cleared first, so what an earlier operation left in the same ring slot is never taken for a submission on the fd, and the stop of a running timer or connect prepares its cancel before its fd is closed. Testsyev_events/test_yevent_close_fd_kept,yev_events/test_yevent_stop_stale_sqe.A zero-copy UDP send (
io_uring_prep_sendmsg_zc()) gives two completions, the result (IORING_CQE_F_MORE) and a notification (IORING_CQE_F_NOTIF); the loop counted one, so an event destroyed at the first completion -- asC_UDP_Sdoes when all the data is sent -- was read after it was freed by the notification (7.25.4 too). The callback is called once, at the first completion; the notification only ends the operation. A kernel without zero-copy sendmsg (probed atyev_loop_create()) gets a plain sendmsg.An IPv6 peer is kept with its real length, so
C_UDP_Sanswers it and an accept or connect takes it: in 7.25.4 the address was a 16-bytestruct sockaddr, an IPv6 accept or connect was refused, a received IPv6 peer was cut and the reply refused (-EINVAL). A url can hold an IPv6 literal in brackets (udp://[::1]:5000).A stop keeps the gbuffer of an operation the kernel still has, and releases it at the LAST completion of the operation, which is not always before the callback (a zero-copy send’s notification comes after it); the callback of a stop sees the event STOPPED, without a gbuffer. 7.25.4 released it at once while the kernel could still write into it or read it. A stop does not reach the callback of an IDLE event that is not a timer, of an IDLE zero-copy send whose notification is pending, of a stop without memory, or of a RUNNING event destroyed while the loop stops; a stop never closes an accept socket.
yev_set_gbuffer()refuses to replace the gbuffer of such an operation.yev_loop_destroy()frees the destroyed events whose completions have not come: it cancels what the kernel still has, collects the completions for up to 1 s, and frees the rest with an ERROR (“Loop destroyed with events whose completions did not come: freed”, withcancel_submitted); in 7.25.4 they leaked, or were freed while the kernel still had their operation. A zero-copy send whose notification has not come is waited for up to 5 s in all, and then NOT freed, with a WARNING (“Loop destroyed with zero-copy sends whose notification did not come: NOT freed, the kernel may still read their gbuffer”): the kernel may still read its gbuffer. Testyev_events/test_yevent_loop_end_drain(case E). When there is no entry for that cancel, an ERROR says so first (“Submission queue full: the cancel of the events left is NOT submitted, their completions may not come”).A connect binds its
src_url, and a bad one is an error: in 7.25.4 a non-emptysrc_urlwas never parsed -- a dynamic build failed the connect with “getaddrinfo() src_url FAILED”, and a static one bound the socket to the loopback address with a port the kernel picked, whatever thesrc_urlsaid. It can be"host:port","[ipv6]:port"or"schema://host:port", resolved in the family of each address of the destination: an address in whose family thesrc_urlhas none is skipped and the next one is tried; a badsrc_url, a failed resolution or a failed bind is logged, naming thesrc_url, and the connect gets no socket. That error path no longer leaks the destination addresses. Testyev_events/test_yevent_connect_src_url.
Transports and inter-yuno events (C_TCP_S, C_UDP_S, C_IEVENT_*)¶
C_UDP_S sent only the FIRST datagram of its queue (7.25.4 too): the send completion asked for
ST_CONNECTED, a state C_UDP_S does not have, and every later datagram waited for ever. And a send that did not start (a gbuffer with no peer address or a bad address length, or no memory to keep the submission) leaked its event and gbuffer (7.25.4 too), lefttx_in_progressabove 0 (a stop waited for ever) and held the queue. Such a datagram is dropped with an ERROR (“Cannot send datagram: dropped”) and the next one is sent. Testc_udp_s_tx.A datagram the kernel refuses no longer stops C_UDP_S (port 0, another family, a broadcast without
set_broadcast, too big). It is dropped with a WARNING that names the peer and the cause (“UDP: datagram refused by the kernel, dropped”, witherrnoandstrerror) and the next one is sent; a send of 0 bytes does not stop the server either. In 7.25.4 it stopped the whole server, reading too, and only a trace said so. Only a read that fails while the server runs stops it now, with an ERROR (“UDP: read FAILED, the server stops listening”), and so does a read that cannot be started again after a datagram (“UDP: the read cannot be started again, the server stops listening”; 7.25.4 ignored that failure, and the server stayed up and deaf with nothing logged). Such a stop ends inST_STOPPEDand publishesEV_STOPPED, also with a datagram in flight. Testsc_udp_s_tx,c_udp_s_self_stop.C_UDP_S refuses a
udps://url (7.25.4 too): there is no DTLS. In 7.25.4 the server listened with no TLS session, and the first datagram crashed the yuno. The start fails with an ERROR (“A secure url (udps://) is not supported by C_UDP_S: there is no DTLS”), and withexitOnError(the default) the yuno exits with 0.use_sslis always false, soreload-certsandview-certanswer “Listener is not TLS-enabled or not running”. See “Upgrade steps”. Testc_udp_s_self_stop.A stop and a start of C_UDP_S (the logcenter’s
pause-yuno/play-yuno) start a new read on the new socket. In 7.25.4 the old read restarted on the number of the closed socket, which could be any file of the process, and reception ended with no log. A stop drops the datagram in flight and the queue (WARNING “UDP server stopped: the datagrams waiting to be sent are dropped”, with the count), and a send that completes after a restart does not move the new queue: in 7.25.4 the datagram in flight at the stop held every later one. Testc_udp_s_restart.The
EV_RX_DATAgbuffer of C_UDP_S always carries the peer as its label. In 7.25.4 the label was set only when tracing, so C_GSS_UDP_S (the logcenter) put every peer on one channel and joined the pieces of long log lines from different yunos into corrupt records. C_GSS_UDP_S keeps a channel per peer (a source ip:port), and caps what the peers make it hold: a frame buffer starts at 4 KB and doubles up tomax_frame_size(default 1 MB, the old fixed size); a frame with no end within it is delivered cut, with a WARNING at most each 10 s (“Frame without end within max_frame_size, delivered cut”);max_channels(default 1024) drops the datagrams of new peers beyond it (WARNING “Too many peers, datagrams of new peers dropped”, on the transition);max_pending_bytes(default 8 MB, all peers together) drops a datagram beyond it with its peer’s unfinished frame (WARNING “Too many bytes in unfinished frames, datagram dropped with the unfinished frame of its peer”, at most each 10 s, with a count); a channel that cannot be allocated is logged, never dereferenced. A frame that C_GSS_UDP_S publishes carries its peer’s label and address, so a host can answer it:EV_SEND_MESSAGEsends a gbuffer that has an address to that address, one without an address to the known peer its label names, and refuses a gbuffer with neither (ERROR “EV_SEND_MESSAGE without a peer: no address, and its label names no known peer; dropped”, -1). In 7.25.4 a send by address logged “UDP channel NOT FOUND”, a send by label had no address to go to, and a frame could not be answered. Testc_udp_s_rx(case 4: one-byte datagrams from many ports; case 5: answers by label and by address, and a send with neither).C_UDP_S publishes
EV_STOPPEDwhen a stop completes. Its event table declared the STATE nameST_STOPPEDas an output event, and nothing was published (7.25.4 too).EV_TX_READYis declared and never published (transport.md says so).An answer written in the
EV_RX_DATAgbuffer of C_UDP_S goes back to its sender, whole (7.25.4 too). transport.md says a host can answer a datagram by sending the received gbuffer back asEV_TX_DATA, but C_UDP_S cleared that gbuffer after the publish and read the next datagram into it: an answer that waited in the queue was sent empty (“Cannot start event: gbuffer WITHOUT data to write”, “Cannot send datagram: dropped”) or with the bytes and the peer of the next datagram, the same gbuffer could sit twice in the queue (“Wrong dl_item_t, WITH links”, memory not freed), and a zero-copy send could read memory the next read was writing. When the host keeps the gbuffer, the next read now takes a new one ofrx_buffer_size; when nobody kept it, it is cleared and reused as before (one refcount test on the common path). No memory for that new gbuffer stops the server, with an ERROR (“UDP: no memory for the next read, the server stops listening”). Testc_udp_s_echo.C_TCP_S asks
denied_ipsat accept, and C_UDP_S asks both ip lists for every datagram; each counts every refusal in a new stat (refusedConnxs,rxRefusedMsgs) and says it at the first refusal of each cause and then at most once a minute (see “Security”). Testsc_tcp_s_ip_lists,c_udp_s_rx.C_TCP: a write that does not start drops the connection (7.25.4 too). A write whose event could not be created or started (no memory to keep its submission, an empty gbuffer) was ignored: the event and its gbuffer leaked,
tx_in_progressstayed above 0, the later writes waited for ever, and a stop waited inST_WAIT_STOPPEDfor ever. The connection is now dropped, with an ERROR (“Cannot start a write: the connection is dropped” / “Cannot create a write: the connection is dropped”): a TCP stream cannot go on with a piece missing (C_UDP_S drops one datagram and sends the next). Testc_tcp/test5.C_TCP: the stop of a client already disconnected publishes
EV_STOPPED(7.25.4 too). A client stopped while it waited for its reconnect timer (after a failed connect) went toST_STOPPED, published nothing and kept its connect event, so the nextgobj_start()failed with “yev_connect ALREADY exists” and the client never connected again. That stop now frees the connect event and publishesEV_STOPPEDonce, insidegobj_stop(), like every other stop of C_TCP: a host that sets its flags aftergobj_stop()gets the event before they are set. transport.md, “The stop”. Testc_tcp/test6.A C client can withdraw a subscription that carries a
__global__(7.25.4 too). C_IEVENT_SRV compared the withdrawal with the stored__global__, which also holds the gate’s back-metadata that the peer never repeats: nothing matched, and the subscription stayed until the channel closed. The withdrawal is filtered as the subscription is, and its__global__is compared with the stored one without the gate’s_keys. Testtest_c_ievent_srv_peer_subs.A repeated hard subscription is one (7.25.4 too).
gobj_subscribe_event()stores a subscription’s__config__without__hard_subscription__, so a second hard subscription with the same kw matched nothing: it was made again, with no log, and each event arrived twice; a hard subscription over a plain one did not replace it either.__hard_subscription__is no longer compared: a subscription that matches a hard one returns the hard one, with a WARNING (“Hard subscription REPEATED, the one there is kept and returned”), and a hard one over a plain one replaces it. A repeated plain subscription is still replaced (“subscription(s) REPEATED, will be deleted and override”). Both WARNINGs are about the caller’s subscription, not a broken invariant: they carry no stack now, and their kw is capped to 256 bytes (7.25.4 dumped the whole kw, with a stack). A repeated__own_event__or__rename_event_name__subscription is compared as it is stored, so it replaces the one there, andgobj_unsubscribe_event()with the same kw removes it: 7.25.4 made a second one (each event delivered twice) and could not withdraw it. Testc_subscriptions/test3.C_IEVENT_CLI sends nothing to a stopping transport. A subscription added or withdrawn while the client is stopping (after
gobj_stop(), before the close arrives) sent its frame down to the transport that was closing, and C_WEBSOCKET or C_TCP logged “Event NOT DEFINED in state” (7.25.4 too). Nothing is lost: the peer drops the channel’s subscriptions when it closes, and an added one is sent at the next open.C_IEVENT_SRV stops its protocol gobj (C_WEBSOCKET / C_PROT_TCP4H) when it is stopped, as every other layer stops the one below it. In 7.25.4 a gate stopped with a plain
gobj_stop()-- the way the yuno stops an autostart service -- left them running, and the yuno exited with “Destroying a RUNNING gobj” (the agent escaped it because it callsgobj_stop_tree()). The protocol gobj still lives as long as the gate, not as long as one connection, and C_IEVENT_SRV starts it again when it starts:disable-channel/enable-channelbring a channel back whole, becauseC_CHANNEL’smt_enablenow starts its tree (gobj_start_tree(), what its comment andgobj_enable()'s default always said).C_WEBSOCKETandC_PROT_TCP4Hstop their transport only when it runs (no “GObj NOT RUNNING”). Testc_subscription_authz(every channel disabled and enabled again, then a session opens on each).An unsubscribe that matches nothing is said for what it is. C_IEVENT_SRV counts the subscriptions of its channel that the authz refused: the INFO “UNSUBSCRIBING event never subscribed, its subscription was refused” is logged only for one of them, and any other unsubscribe that matches nothing is a WARNING without a stack (“UNSUBSCRIBING event matches no subscription of this channel”), where 7.25.4 logged the ERROR “No subscription found” with a stack.
gobj_unsubscribe_event()warns when a hard subscription it matched is left in place (“Hard subscription not removed, only gobj_unsubscribe_list() with force removes it”); 7.25.4 counted it as removed and said nothing.C_PROT_MQTT2: an incoming QoS 2 message queued at the reload of a persistent session is released (7.25.4 too). The move of queued incoming messages into flight tested the quota the wrong way round (it stopped when nothing was in flight and moved a message when the quota was full; mosquitto stops when the quota is spent), gave the moved message a NEW packet id and sent its PUBREC with it: the client’s PUBREL for its own id found nothing (“Message not found”), and the message was never released. A queued message now goes in flight when a slot frees, with its own id, as in mosquitto. Only a session reloaded with more pending QoS 2 messages than
max_inflight_messages(the limit lowered between two connections) has queued ones. Testc_mqtt/queued_in(a raw client).C_PROT_MQTT2 client: a QoS 2 message received again with DUP=1 is found by its packet id (7.25.4 too). It was looked up by queue rowid: the old copy stayed in flight for ever, the next message with the same packet id delivered the stale copy again, and an unrelated message could be removed instead. Now the copy waiting for its PUBREL is found by packet id, in flight or queued, as in mosquitto, and replaced (WARNING “QoS 2 message received again (dup): it replaces the copy waiting for its PUBREL”, was “removing an inflight qos2 dup message”). And a client’s QoS 2 PUBREL no longer logs “QoS mismatch”: the qos flag bits (32) were compared with the qos level (2). Test
c_mqtt/client_queues.MQTT queues: a message that expires before it is sent gives its slot to the queued ones (7.25.4 too). The send loop dropped the expired message and left the queued messages waiting, although a slot was free: with 20 in flight and the 21st expired, the 22nd was never sent. The release now moves queued messages in flight while there is room, sends, and repeats while an expiry frees a slot. Test
c_mqtt/client_queues.A queued MQTT message moves in flight with its content (7.25.4 too).
tr2q_move_from_queued_to_inflight()moved the message first and read its content after: a read that failed left an incoming QoS 2 message in flight with no content, and no PUBREC was sent. The content is read first; on a failure the message stays queued (ERROR “Cannot load the content of a queued message: it stays queued”) and the call answers -1. Andtr2q_check_backup()does not back up while messages are queued (it answers 0): the backup could lose their content. Testtr_queue/test_tr2q_queued.MQTT: the broker keeps the client’s Receive Maximum (7.25.4 too). With
max_inflight_messages0 (no limit of the broker’s own), the broker sent every queued QoS 1 and QoS 2 message at once, over the Receive Maximum of an MQTT 5 client. The out window is now the queue’s limit, the lower of the two maximums, as in mosquitto. Testc_mqtt/out_flight(a raw MQTT 5 client with Receive Maximum 2: two messages, then two more after the acks).MQTT broker:
list-queues qos=keeps the pending filter (7.25.4 too). Aqos=replaced the defaultpending=1and listed the messages already delivered. The conditions combine now:list-queues qos=1lists the pending QoS 1 messages, andlist-queues pending=0 qos=1the delivered ones. Testc_mqtt/out_flight.MQTT: an ack of the wrong packet is a protocol error (7.25.4 too). A PUBACK for a QoS 2 message, a PUBCOMP for a QoS 1 message or before the PUBREC, and in the client a PUBREL of the wrong QoS, were an ERROR, and the message was removed as delivered. Now it is a WARNING (“QoS mismatch”,
MSGSET_MQTT, withclient_idandpeername), the connection is closed (an MQTT 5 peer gets DISCONNECT 0x82 first) and the message stays in flight, as in mosquitto. An ack of a packet id that is not in flight is a WARNING (“Message not found in trq_out_msgs”, was an ERROR): the id comes from the peer. See “Upgrade steps”. Testc_mqtt/out_flight.
Agent, gobj-c and tools¶
The agent removes old audit files. New attribute
audit_keep_days(default7;0keeps all, as 7.25.4 did) is the retention of/yuneta/realms/agent/agent/audit/. At start and at each new audit file the agent removes the audit files older than that (only names with the shape of the audit mask, never a link) and logs one INFO “Old audit files removed” with the names. Up to 7.25.4 nothing was removed (19 GB on wattyzer). An existing large directory is swept at the first start: to keep more, setagent.audit_keep_daysinyuneta_agent.jsonbefore that start. Newrotatory_remove_old_files(); the sweep at a new audit file runs inside the write of its first record (once a day, or at a size rotation).The agent’s audit record is small, and holds no secret and no keystroke. A read-only command (
list-*,view-*,get-*,info-*,dir-*, help, stats, services, nodes, treedb-info, topics, ...; the list is in DEBUGGING.md 5.5) is written as{command, date, user}only. A command with a__reset__value (stats-yuno stats=__reset__) is a write. The command word is read as the parser reads it (any case, quotes, aliases, looked up in the agent’s command table):WRITE-TTY,'write-tty',EV_WRITE_TTYare console writes too, andCLOSE-CONSOLEends the bursts.command-yuno/command-agentare judged by the command they carry, the one its handler reads: the key exactlycommand(the parser stores a key as typed), socommand-yuno id=gate COMMAND=list-yunoswith a kwcommand=delete-yunoruns delete-yuno and gets the full record; acommandkey in any other case is never read-only, andwrite-tty NAME=xdoes not move the record to another console. Any other command is written as{command, date, user, source, kw}:__md_iev__is no longer written, andsourcekeeps the console purpose and each inter-yuno hop (role, yuno, service, user, host).user, each field of a hop,console_purposeand the console name of a console write come from the peer: each is redacted like any other string, and one longer than 1024 bytes is written as<N bytes, not scanned, sha256:HEX>. Acontent64is never written (blanks around the=included): it becomes<N bytes sha256:HEX>of the decoded content (thesha256sumof the binary). A secret is never written: the value of a parameter named likepassword,pwd,secret,token,jwt,private_key,api_key(set withwrite-attr attribute=api_key value=...on wattyzer’sgate_pvpc),apikey,x-api-key,cookie,session_key,__session_id__,auth_data,passphrase,credential,authorization,bearerorsalt(and thevalueof awrite-attrof such an attribute, also as{attribute, value}in a json object of the kw) is<redacted>, in the kw at any depth and in any string, as are the token afterBearer, the credentials afterBasic(the base64 ofuser:password) and anything shaped like a JWT (also one followed by a.). The value of a secretname=valueis its whole shell word, quoted pieces and escapes included:password=\"two words\",password="a\" S",password='it'\''s S'; the closing quote of a"..."value that ends with the secret stays in the record (command="write-attr attribute=api_key value=S" n=1is writtenvalue=<redacted>" n=1). A name is judged with its json escapes decoded, at any level: a key (\"secret key\":,\"password/db\":, a key or anattributewrittenapi\u005fkey), and awrite-attrgiven as JSON text inside a string of the kw is recognized:{"x":"{\"attribute\":\"api_key\",\"value\":\"S\"}"}writes noS. A json key is read with its escapes, and a JSON text given inside a JSON text (quotes escaped as\"or\u0022, up to 8 levels) is scanned with its escapes decoded, the record keeping the same JSON text (deeper than 8 levels the run is written as<N bytes, not scanned, sha256:HEX>), whatever quotes come before it (a double-quoted parameter, a stray quote). When no quoted run holds it whole (x="{\"password\":...}", where the parser’s value ends at the first\", or'{\"password\":...}'), the key is read by its shape and its value is taken with the quotes of its level. The redaction is stricter than the parser:attribute/valuematch in any case, and a write-tty or a secret carried underCOMMAND=is hidden too. Up to 7.25.4check-user-pwdandset-user-pwdwrote the password in clear text (see “Security”). The record is built before the command parser and the authz, from the text of any peer that can send a command: its scan is one pass in linear time at each level of escaped JSON text (at most 8), and it scans at most 128 MB of strings per record, all levels together (a string past that budget is written as<N bytes, not scanned, sha256:HEX>; of the command text, its first word comes first). A console keystroke (write-tty) keeps only who, when, which console, and how many writes and bytes: one record at the first write of a burst (one user, one console, up to 60 s) and one for the rest when the burst ends; up to 7.25.4 each keystroke wrote the whole kw with the keystroke in base64 (1000 keystrokes: 922 KB, now 945 bytes). A field of a__md_iev__hop from a peer that is not a string is written as"", with one WARNING (“Audit: bad md_iev from a peer, written as empty”). Aninstall-binaryof a 32 MB yuno went from 134 MB to 541 bytes (wattyzer wrote 0.6-1.2 GB of audit on a deploy day). A day of audit that crossesmax_megas_audit_filecontinues in.OLD.1,.OLD.2, ...: up to 7.25.4 each size rotation removed the previous.OLD, so a day that crossed the limit twice lost its first part (wattyzer lost the mornings of 22 and 23 September 2026). Newrotatory_keep_all_old_files()(off by default: the yuno logs keep their one.OLD).The agent flushes every audit record before the command runs (one
write()per command, about 0.75 us more:perf_rotatoryaudit_record_flush1.34 us againstaudit_record0.60 us, the same 12 alternated rounds): up to 7.25.4 a record could wait in the buffer of the file and be lost in a crash.A command whose kw carries a
gbufferreleases it once. The command parser gives the handler a new kw with the keys of the caller’s kw, and it copied them withjson_object_update_missing()(every command, with or without a parameter schema); the default stats parser (build_stats()) handed the kw on withjson_incref(). Neither took a reference of the gbuffer, and eachKW_DECREFreleased it: “BAD gbuf_decref()”, or a gbuffer freed while its owner still held it. The parser uses the newkw_update_missing(), which takes that reference, and so does theEV_ON_CLOSEof C_IEVENT_SRV;build_stats()and C_MQIOGATE’sview-channelsusekw_incref(). C_NODEcreate-node/update-nodehad worked around it (they took a second reference of the bytes of thefilecolumns); they take the command kw’s own now. Testcommand_binary_kw.kw_set_dict_value()overwrites, as its header always said, and answers -1, logged, when it writes nothing (a scalar or a null in the middle of the path, “long path”; an index outside its list; an empty path, “path empty”). Up to 7.25.4 a key that already existed kept its old value, with no sign: a peer’s own__username__survived the gate (see “Security”), a stale ref in a string fkey column could never be removed (see “Data loss and integrity”), andgobj_set_stat()/gobj_incr_stat()/gobj_decr_stat()never changed a stat that existed. All 38 callers in the tree were checked: none relied on “set only if absent”, and none reads the return. gobj-js’skw_set_dict_value()already overwrote. Testkw/test_kw_set_dict_value.New
kw_twin(gobj, kw): a new kw with its own top level, the values shared, and each binary field (gbuffer) increfed askw_incref()does, so the twin and the original each release their own reference. It is whatgobj_publish_event()gives a subscription that rewrites the kw (see “Security”), for a kw written at the top level only:json_t *mine = kw_twin(gobj, kw);, thenjson_object_set_new(mine, "x", json_integer(1));leaveskwas it was, andKW_DECREF(mine)releases the twin (kwid.md).A log call leaves
errnoas it found it (gobj_log_*(), the traces), and a failedmkrdir()leaves the cause inerrno(ENOTDIRwhen a part of the path is not a directory,ENAMETOOLONG,EINVALfor an empty path, or the syscall’s);rmrdir()/rmrcontentdir()setENAMETOOLONG/EINVALon their refusals. A caller that loggedstrerror(errno)after a helper that logged its own failure named another cause, or “Success” (7.25.4 too):tranger2_create_topic()'s CRITICAL “Cannot create TimeRanger subdir. mkrdir() FAILED” always said errno 0, and callers in timeranger2, C_TREEDB and the agent had the same defect. Teststr_queue/test_tr_queue_backup_failed,helpers/test_dir_read_error.gbuffer_base64_to_binary()decodesbase64_lenchars instead of reading up to a'\0': a slice of a longer text (acontent64='...'inside a command line) failed to decode.gbmem_realloc()refused because the new size is larger than the largest block leaves the old block valid and tracked: withCONFIG_DEBUG_TRACK_MEMORY7.25.4 took the block out of the tracking first, and the later free logged “Wrong dl_item_t, WITHOUT links” and wrapped the memory counter. The kept lists of yev_loop reach this path.rmrdir()andrmrcontentdir()no longer follow symbolic links: a link to a directory inside the tree was walked into and the files of its TARGET were deleted (outside the tree); a dangling link made the removal fail. A link is removed as a link; their failures name the path, andrmrcontentdir()no longer fails silently. An entry that another process removes during the walk is not an error (it returned -1 with no log).mkrdir()over a path that exists and is not a directory logs “Not a directory: the path exists and is not a directory” (ENOTDIR) and returns -1; over a dangling link it logs “newdir() FAILED” (EEXIST) and returns -1. 7.25.4 returned 0, with no log and no directory (its check could never be true).rmrdir()/rmrcontentdir()refuse, with a log, a tree whose paths do not fit in PATH_MAX or that is deeper than 1024 levels (they could recurse without end);mkrdir()refuses, with a log, a path of PATH_MAX or more (7.25.4 cut it silently and returned 0).find_files_with_suffix_array()never lists a symbolic link (on filesystems withoutd_typeit usedstat()and listed a link to a file).A directory walk whose paths do not fit in PATH_MAX fails, logged (7.25.4 too).
walk_dir_tree(),walk_dir_array()andget_ordered_filename_array()ignored a failedbuild_path(): the entry’s path was the directory itself, and the walk went into it again until the process crashed (SIGSEGV), or answered 0 with the directory given to the callback as its own entry.find_files_with_suffix_array()on a filesystem withoutd_typestat’ed the directory instead of the file and dropped the file with no log. The walk now answers -1 with the ERROR “Path too long, the directory cannot be walked”, and a tree deeper than 1024 levels (16 on ESP32) with “Tree too deep, the directory cannot be walked”, asrmrdir()does.A walk callback that returns FALSE stops the whole walk (7.25.4 too): only the directory of the callback stopped, and the walk went on in the next one, where helpers.h and directory_walk.md said it would stop.
walk_dir_tree()still answers 0. The transient causes that now fail a walk are under “Data loss and integrity” (a key directory that cannot be listed). Of the callers, timeranger2, the tr_treedb import and blob sweep,dir_listingand c_resource2 already answered the -1; fs_watcher, C_FS and timeranger2’s rt disk search log it and stop there instead of skipping one subtree. Testhelpers/test_dir_read_error.The agent’s
dir-yuneta,dir-realms,dir-repos,dir-store,dir-logsanddir-local-dataanswered an EMPTY list with result 0 for a directory that could not be listed (does not exist, mode 0, EMFILE, a badmatch): they ignored the return of the listing. They answer -1, “<role^name>: cannot list ‘<dir>’, see the log”, with no data (the cause is in the log):dir-store subdirectory=nopeanswers “-1: yuneta_agent^agent: cannot list ‘/yuneta/store/nope’, see the log”. Testshelpers/test_dir_listing,helpers/test_dir_array_nomem(the listing helpers themselves are under “Data loss and integrity”).save_json_to_file()checksclose()-- a failed close is a CRITICAL aton_critical_error-- and logs a missing directory when it may not create it.rotatory (the log files): a record is checked once and written whole. Up to 7.25.4 each piece of a record rebuilt the file name (
localtime()) and ranaccess()+fstat(): an agent audit record cost 6.2 us, now 0.55 us. A record is no longer split between two files at a size rotation. A log file RENAMED by another program is no longer noticed (a removed one still is); the free-disk check runs every 100 records instead of every 100 pieces.rotatory: a full disk stops only the log file on it, and it writes again when the space is back (checked every 100 records; one line on stdout/syslog when it stops, one when it resumes with the number of dropped records). Up to 7.25.4 one full disk stopped EVERY file log and the agent audit of the process until it was restarted. A handle closed by
rotatory_close()/rotatory_end()is never touched again (the logger keeps its handle): a write, flush, truncate or second close through it does nothing; up to 7.25.4 it read freed memory and could close it twice.rotatory: a clock set back across midnight empties no log file. Up to 7.25.4 each new file name was opened with
"w": a step back at 00:00:05 emptied the file of the day before, and the step forward emptied today’s file with its first records (the agent audit and the yuno logs). Now an existing file is emptied only when it was last written before the day its name is used for (last week’s file of aWmask, as intended); a handle withrotatory_keep_all_old_files()never empties a file.rotatory: a new day on a full disk opens the new file and calls the newfile callback (where the audit runs its retention and frees space); a piece of 0 bytes no longer closes the file until the next day; a failed write closes the file and the next record opens it again. An old file of a mask with a day letter or the month and without the year is emptied at the first record after
rotatory_open(): up to 7.25.4 a yuno that started on a Monday appended to last Monday’sWfile, so one file held 8 days. A mask with the year, and a mask with no date letter (a fixed name,logcenter.log), never empties a file; anMM-only mask is judged by its month (a restart on the 15th does not empty the month’s file). The decision waits for the first record, sorotatory_keep_all_old_files()called right after the open applies to the file the open found.rotatory: the newfile callback (
rotatory_subscribe2newfile()) runs for a NEW file only (a new name, or a size rotation), never for the same file opened again after a failed write, a removed file or a failed truncate. The logcenter sends a summary e-mail from that callback, and the agent audit runs its retention scan there: a log file that refuses writes (quota, no inodes, EFBIG, EIO), and so is opened again at the next record, sends no e-mail per log line and scans nothing per command. The callback of a new file whose open fails (EMFILE, an inode quota, a directory that refuses writes for a moment) runs once, at the next open that works, with the old name of before the failure; on a full disk that open is tried every 100 records, since the retention the callback runs is what frees the space. In 7.25.4 a failed open of a new day ended the process of a handle opened withexit_on_fail(next bullet) and stopped any other for the rest of the day, and that day’s callback (the audit retention, the logcenter summary) never ran.rotatory:
exit_on_failis forrotatory_open()only. A file that cannot be opened later -- a new day with EMFILE, an inode quota, a directory that refuses writes, a removed file, the reopen after a failed write, a truncate -- is printed once on stdout and syslog (“_rotatory(): Cannot open ‘<path>’ file, <err>”), the record is dropped, the next record tries again, and “_rotatory(): ‘<path>’ is open again” is printed when it works; the pending newfile callback then runs. In 7.25.4 the first such failure ended a process whose handle was opened withTRUE: the agent (its audit) and every yuno (its file log).rotatory_open(..., TRUE)still exits when the open itself fails. Testhelpers/test_rotatory(each case in a forked child withexit_on_fail).rotatory: the size limit is compared in bytes (7.25.4 too). It was compared in whole megabytes, so a limit of 8 MB rotated a file at 9 MB, and a yuno’s logs reached 126 MB instead of 112 MB (7 days x 2 pieces x 8 MB; DEBUGGING.md 5.3). A file of exactly the limit is not rotated; the next record over it is. Test
helpers/test_rotatory.rotatory: at a new day the file of the day before is not size-rotated. In 7.25.4 a day whose last piece crossed the size limit had that file renamed to
.OLDat the first record of the next day, and the day’s earlier.OLDwas removed.rotatory: with
rotatory_keep_all_old_files(), a size rotation whose rename fails (chattr +a, a read-only bind, a MAC denial) keeps the file, appends to it, and tries the rename again after 60 s of the monotonic clock (a wall clock set back or forward does not move it) or at the next name, with one line when it fails (“_rotatory(): Cannot rename ‘<path>’ to ‘<path>.OLD.1’, <err>, the file is kept and grows, the rename is tried again every minute”) and one when it works again (“_rotatory(): the size rotation of ‘<path>’ works again”). Without it, a failed rename still empties the file, as in 7.25.4.The entry point closes the log files last: up to 7.25.4 they were closed before the final cleaning and the memory leak report, so “system memory not free” never reached the yuno’s log file.
yuno-skeleton: the four C templates logged with
MSGSET_INTERNAL_ERROR, which the unified msgsets renamedMSGSET_INTERNAL(2026-04): a gclass generated from them did not compile. They useMSGSET_INTERNAL.yuno-skeleton: the timer of
gclass_service,gclass_childandyuno_citizenreached nobody. They created it withgobj_create(), andC_TIMERsubscribes its parent only when it is a pure child: itsEV_TIMEOUTwas published without subscribers and the template’sac_timeoutnever ran. They create it withgobj_create_pure_child(), asyuno_standaloneand the kernel gclasses do. The JS template already did.New benchmarks
performance/c/perf_timeranger2,perf_tr_treedb,perf_c_treedb(see “Performance”) andperf_rotatory(the rotatory figures, with a case that flushes each record): each prints one line of JSON per result, and ctest runs them with small sizes.The test harness (
capture_log_write(), gobj-ctesting.c) no longer counts “io_uring_queue_init_params() pinned-memory pressure, retrying” as an unexpected log: it speaks of the machine (other processes held locked pages when the test created its loop), and the loop retries and goes on. A full suite run beside other work failedtest_yevent_listen1on it; the message is still printed.create-yunorefuses a release name too long for it (7.25.4 too). It builtyuno_releaseinto 120 bytes and ignored the failure: a binary version plus a configuration version longer than that stored a cut pkey2 and answered success;find-new-yunosnever found that name, offered the yuno again at every call, andcreate=1failed on it with “Yuno already exists” for ever. The buffer isNAME_MAX, and a name that does not fit answers -1: “<role^name>: Yuno ‘<role>.<name>’: release name too long (‘<binary version>’ + ‘<config version>’)”.find-new-yunossays which rows are already registered. On a resumed upgrade (afind-new-yunos create=1that ran, and nodeactivate-snapafter it) its preview listed those rows again as new, andcreate=1failed on each with “Yuno already exists” (7.25.4 too). A row whose yuno already exists at the new release stays in the preview, marked “already registered, pending promotion (deactivate-snap): create-yuno ...”;create=1skips it, and the comment gives the count (“<role^name>: N yuno(s) already registered at the new release, pending promotion: run deactivate-snap”). The marked rows stay in the preview so thatyunetas upgrade-yunosstill runsdeactivate-snapfor them; the CLI 0.19.4 counts them apart (see “Upgrade steps”). Testc_agent_find_new_yunos(the preview).Build: an archive installed with a changed content is now newer than the tests and yunos linked against its previous copy, so
cmake --build build --target <test>relinks them.install()kept the archive’s build time, cut to whole seconds, and a test linked between a library’s build and its install looked newer than the new archive: a per-module test run could run the old library (7.25.4 too). An unchanged archive relinks nothing.
C_NODE, C_AUTHZ, C_TRANGER¶
Every C_NODE and C_AUTHZ command comment starts with the yuno (except the “No permission ... in service” refusals, which name the service), and none reads
gobj_log_last_message().import-dbcounts its errors by cause indata.errores("node exists","cannot create the node (see the log)"); 7.25.4 keyed them ongobj_log_last_message(), one key per record. A create refused for another cause than an existing node is afailurein every mode (skipcounted it asignored) and stops anabortimport.treedbs,links,hooksandnodehand their kw on withkw_incref()(7.25.4:json_incref(), released withKW_DECREF).activate-snapkeeps the old snap active when the new one cannot be saved.C_TRANGER
mark-tm-order [topic_name=<t> | all=1](master,write), over the newtranger2_mark_tm_order(); withall=1it skips, with an INFO, each directory of the store that is not a topic (notopic_desc.json, e.g.saved_schemas/). A topic that cannot be opened answers -1 (“<role^name>: cannot open topic ‘<t>’ (see the log)”).gc-assetsanswers a report (dry_run,assets,blobs,refused,blobs_refused) instead of the list of ids taken.C_TRANGER
add-recordworks (7.25.4: a stub that answered -1 “Pending to review” and logged an ERROR with a stack). It appends a record that carries the topic’s pkey, on the master only, with the permissionwrite:command-yuno id=<id> service=<tranger> command=add-record topic_name=pp record='{"id":"1","tm":1700000000}'answers “<role^name>: record added to topic ‘pp’, rowid <n>” withdata: {topic_name, rowid, t, tm}.recordis a dict or its json text;__t__(0: now) anduser_flagare optional. A replica answers -1 before the library is called. Testtest_c_tranger.The pushes of a C_TRANGER live list carry its id as
rt_id(7.25.4: no id), so a client with two lists open on one topic tells their records apart. Testtest_c_tranger.instancesno longer leaks itsfilterwhen it answers “What topic_name?”. Anupdate-nodewithcreatewhose create is refused logs “Cannot update node: it does not exist and it cannot be created (see the previous log)” (7.25.4 answered withgobj_log_last_message()).import-db counts a link it cannot make. A record naming a parent that does not exist was saved without the link and counted as a success (
"link failure": 0); it is alink failurenow. And import-db answers -1 when it aborts or anything fails, with a comment naming the counts, for example “<role^name>: import-db incomplete: 3 added, 0 overwritten, 0 ignored, 0 failed, 1 link failure(s) (see the log)” (ABORTEDfor an abort); a clean import answers 0 with the counts. 7.25.4 answered 0 with no comment (tests 15, 16 ofc_node_link_events).export-db with no
filenamereadsschema_versionas a number: an integer version logged an ERROR with a stack (“path MUST BE a json str”) and was left out of the file name, which is now<treedb>-<schema_version>-<date>.trdb.json(for exampletreedb_link_test-1-<date>.trdb.json). A failed write of the export is logged (“Cannot write the export of the treedb”) and answers -1 (it answered 0). Test 17 ofc_node_link_events.link-nodes / unlink-nodes split a child ref at its first
^: the topic is what comes before it (a topic name holds no^), the id is all the rest, up to 255 bytes. A child whose id holds^or is 255 bytes long (legal in a topic without hooks, such as an MQTT client id) was listed by the treedb and could not be named: “Wrong child ref”. For examplelink-nodes parent_ref=groups^g1^users child_ref=users^client^7links the userclient^7. The parent ref keeps its three strict parts (test 18 ofc_node_link_events).import-dbnames the peer of a content that is not json. It parsed the content as json of our own: an ERROR with a stack and a dump of the whole buffer (megabytes), naming no peer (7.25.4 too). It is a WARNING now (Protocol, “frame is not json”,peername, at most 256 bytes dumped), and the answer is still -1 (test 19 ofc_node_link_events).snap-contentreads only the topics of its treedb. 7.25.4 read any topic of the tranger by name, and its overview counted every topic, another treedb’s included, whose records carry the same snap numbers. Atopic_namethat is not a topic of the treedb answers -1, as in the other commands (treedb_is_treedbs_topic()), and the overview walks the treedb’s topics (test 20 ofc_node_link_events).C_NODE’s
EV_TREEDB_NODE_*feed and C_TRANGER’sEV_TRANGER_RECORD_ADDEDaskreadwhen the yuno setsenable_subscription_authz(see “Security”).
JS: gobj-js, gobj-ui, gui_agent, gui_treedb¶
The versions: gobj-js 7.25.1 - 7.25.8, gobj-ui 7.25.6 - 7.25.22, gui_agent 0.22.79 - 0.22.98, gui_treedb 0.17.58 - 0.17.71 (SDK 7.25.4 shipped gobj-js 7.25.0, gobj-ui 7.25.5, gui_agent 0.22.78, gui_treedb 0.17.57). In this section a bare version is the package’s own; the SDK is written “SDK 7.25.4”. Each gobj-js and gobj-ui version named was published on npm, and each gui_agent / gui_treedb version was deployed, so a parenthesis about one of them (for example “gobj-ui 7.25.15 - 7.25.19 applied ...”) is about a version a consumer may have installed or run.
Deployed to artgins
.yunetacontrol .com and .ovh (gui_agent) and artgins.ytreedb.com (gui_treedb); every deploy console-checked (login, session, Schemas editor, forced reconnect, navigation and clicks during loads): 0 errors. Includes the fix of a regression of gobj-ui 7.25.6 that was live (navigating in the schema editor during a reload emptied the model). Schema editor: the loading screen is honest (body cleared, toolbar disabled, navigation waits for the load); a dialog opened on a model that a reload replaced is closed with a message and nothing stamped with the old model is written; a refused reload keeps the previous model; a transport drop ends the cut load or write and the reconnect reloads; stale answers are ignored by round; no load is asked out of session (the reconnect asks it); a late successful write reloads and keeps its marks. A reload that was refused is owed: the operator’s next change runs it first (the change is refused with a toast and the operator does it again once the schemas are in; a move, the editor’s own or the host’s, is not refused and reads there); any load that lands clears the debt (a Refresh included), and an import plan is dropped only by a load that lands. A confirmation that arrives with no model is refused as stale; the import plan lost to a reload is a warning and a toast.
gobj-ui 7.25.7 / 7.25.8, the shell toasts (
yui_shell_show_info/warning/ error): ONE toast per (kind, message) on screen. The same string of the same kind while one still shows restarts its time instead of stacking a copy (one close of a transport showed a column of identical “the connection dropped”); each caller holds it with its own handle and time, and it goes when every holder has closed or timed out (the ✕ still closes it for everybody).Shell modals: every ✕ (toast, modal, confirmation),
MODAL_BACKand everyCONFIRM_BTNcarry a translatable title and aria-label; the schema editor’s export C/JSON switch is two buttons (aria-pressed), switched through the FSM (EV_EXPORT_VIEW); its confirmations pass the keysdelete/cancel(they rendered in English in every locale).gobj-ui 7.25.13 / 7.25.14: in
C_YUI_TREEDB_TOPICS, a write cut by the drop no longer asks for its topic out of session (the adapter logged “cannot route ‘nodes’ — not in session”). A drop marks the view, and the first edge that finds the transport in session reads every open topic table again, once: a node another writer created, changed or deleted during the drop shows. An “up” before the transport is in session, or with no drop before it, reads nothing. A form already answered by the edge is not answered a second time.gui_agent 0.22.79 - 0.22.81 (
C_AGENT_TREEDB_LINKand the Schemas tab). The session closing ANSWERS every routed treedb request in flight (“the connection dropped”; a form stayed busy and the schema editor stuck until the page was reloaded). The deadline of a routed request starts at the controlcenter’s dispatch ack: 60 s, plus the base64 size of its__files__at 128 KiB/s; a late answer is logged as a warning, and a late write that succeeded echoes its node event. Requests are numbered once for the page (with two treedb views mounted an answer could go to the wrong one). The Schemas tab: itssave-schema/saved-schemarounds carry their id and another round’s answer is ignored; a session close settles the Save, the saved-schema round and the comparison; the apply deadline is oneC_TIMERchild, 30 s per step; a Save that answers only “nothing to save” while drafts are marked is said, and a Save that withdrew the saved schema says so (new i18n key). A hosted treedb view hears a session edge after its transport (EV_TRANSPORT_EDGE, a posted event), so a view is never told “connected” while its transport is not in session yet.gui_agent 0.22.87: an
applytimeout decides with the answers that came. One treedb applied means restart, so kill / run / play goes on without the silent owners, and the toast names them and says their state is unknown (the tab of SDK 7.25.4 stopped and left the applied schema for the next unrelated restart). With nothing applied there is no restart, and the toast names the silent owners. New locale keys (en, es). 0.22.88: only the gobj-ui range (^7.25.14).gobj-js 7.25.1:
kw_get_str()returns its default as given, as the C one does. It returnedString(default), so a default of0ornullcame back as the truthy"0"/"null":kwid_get_ids()added the id"0"for a record with no id, and the ievent client named its target service"null"or"0"instead of itswanted_yuno_service(the server then used its main service).KW_CREATEstores a string default, andnullfor any other.gobj-ui 7.25.15,
C_YUI_SCHEMA_EDITOR: a move sent by the host closes the dialog of the screen it left (the column or topic form, the import, the orphans), saying “the view moved: open the dialog again” (new consumer i18n key); a Save there answered “Event NOT DEFINED” and the edit was lost. A confirmation answered on another position is refused: a Yes to “delete this column?” asked ondb/usersand answered after a move todb2/usersdeleted the column ofdb2.gobj-ui 7.25.15,
C_G6_NODES_TREE/C_YUI_TREEDB_GRAPH: a__graphs__write that did not land (an error answer, a refusal by the transport, no session) is written again at the next Save (new input eventEV_GRAPHS_WRITE_REFUSED {topic, graphs_load}); it was recorded as saved before it left. A refusal names the load it belongs to:C_G6_NODES_TREEcounts its loads (eachEV_CLEAR_DATAis one), sendsgraphs_loadin theEV_UPDATE_NODEof__graphs__,C_YUI_TREEDB_GRAPHechoes it back on the three refusal paths, and a refusal from an older load is ignored with a warning (gobj-ui 7.25.15 - 7.25.19 applied a refusal answered after a reload to the fresh load). A host that sends nographs_loadis handled as before.The ranges: gui_treedb 0.17.58 - 0.17.61 took gobj-ui ^7.25.11 - ^7.25.14 (gobj-js stayed ^7.22.2); gui_agent 0.22.89 and gui_treedb 0.17.62 took gobj-ui ^7.25.15 and gobj-js ^7.25.1.
gobj-js 7.25.2 / 7.25.3:
kw_get_list(),kw_get_dict()andkw_get_bool()answer like the C readers. A value of the reader’s type is the answer; an absent key or a value of another type gives the default back (kw_get_list()/kw_get_dict()as given,kw_get_bool()as a boolean).kw_get_list()wrapped ({x: [1, 2]}gave[[1, 2]], a default ofnullgave[null]),kw_get_dict()turned anulldefault into{}, andkw_get_bool()read any value throughBoolean()(the string"false"was true).kw_get_bool()logs a value that is not a boolean even withoutKW_REQUIRED, as C does (“path MUST BE a json boolean”, the kw traced); withKW_WILD_NUMBERit reads a number (0is false), a"true"/"false"string in any case (else its decimal integer, asatoi()does:"0x1F"is false since gobj-js 7.25.5, 7.25.2 - 7.25.4 read it as hex) ornull(false), and a list or a dict isfalseand logged (“path MUST BE a simple json element”), as in C:kw_get_bool(gobj, {on: "false"}, "on", true, kw_flag_t.KW_WILD_NUMBER)isfalse.KW_CREATEstores only a default of the right type,KW_EXTRACTtakes out only a value of the right type,KW_REQUIREDlogs a wrong type.kw_set_subdict_value()passedKW_REQUIREDwhere it meantKW_CREATE, andtrace_json()no longer needswindow(in a worker or node aKW_REQUIREDmiss threw instead of logging). With no gobj,kw_get_list()/kw_get_dict()of gobj-js 7.25.2 - 7.25.4 logged a false “gobj bad type” (fixed in 7.25.5). No caller in the SDK or the project SPAs relied on the old answers; every caller ofkw_get_bool()passes a boolean (__hard_subscription__,__own_event__,options.create).gobj-ui 7.25.16:
C_YUI_SCHEMA_EDITORhears a Refresh during a Save and runs it when the writes end (it started the load at once and the rest of the write queue was never sent, with no message). The Save of a form while a reload is owed keeps what was typed: the load that lands opens the form again on the schemas it read, with only the changed fields put back on top (two new consumer i18n keys).C_G6_NODES_TREE: a refused__graphs__write keeps Save lit until a Save writes it, or until an echo of the__graphs__topic shows it written (gobj-ui 7.25.19; 7.25.16 - 7.25.18 kept it owed after such an echo). The shell dialogs’ default labels are the i18n keysok,yes,no,delete,cancel(they were English literals no locale holds);yui_installasksinstall this app. Every consumer SPA that shows these dialogs carries the keys (yunovatiosgui-commongotok); yunomusica shows none of them.gui_agent 0.22.90: Apply is off while a Save of schemas is unanswered and refused if it arrives anyway (confirmed before the Save answered, the apply dropped the Save’s answers and Save stayed off for good). The Apply dialog closes, and says so, when a Save or a new round changes what it lists; an Apply outside
ST_READYis refused and said; a late answer of an apply step is logged, not dropped; round numbers are counted per page, so a tab opened again cannot take an old tab’s answer. Five new locale keys (en, es). gui_agent 0.22.90 and gui_treedb 0.17.63: the ranges (gobj-ui ^7.25.16, gobj-js ^7.25.2); gui_treedb gets theokkey.gobj-ui 7.25.17, gobj-js 7.25.3, gui_agent 0.22.91, gui_treedb 0.17.64. In the treedb graph, a Save whose
__graphs__write the backend refused lit Save again only in tests: the real transports (C_IEVENT_CLI, gui_agent’sC_AGENT_TREEDB_LINK) answer with the request’s__md_command__and not itsrecord, so the view logged “a refused graphs write names no topic” and the next Save found nothing to write. The topic now travels in__md_command__(graph_topic). In the schema editor, a write in flight also makes its bodyinert: the pointer was already blocked, but the keyboard reached the focused row and sent actionsST_SAVINGdoes not declare (“Event NOT DEFINED”). A write marks the topic of its record, not the topic on screen (an import from the topics screen marked nothing, so no draft chip and no export warning). The reason a load failed changes language; a treedb or topic name in a notice is shown as it is (gobj-ui 7.25.17 translated it: a topicnodesread “Nodos”; fixed in 7.25.19). In gui_agent’s Schemas tab, a Save or asaved-schemaround whose owner never answers ends after 30 s (aC_TIMERper round, like the apply steps) and names the silent owners (new keyssave unanswered,saved schemas unanswered); it stayed open until the session dropped, and Apply stayed off with it. A Save cut by a drop reads the saved schemas again when the session is back, and since 0.22.94 a read that shows the Save landed clears “unsaved changes” (an error answer, or data that is not a list, proves nothing and keeps it, with a warning naming the owner; gui_agent 0.22.94 cleared it); a latesave-schema/saved-schemaanswer is logged as a warning (gui_agent 0.22.91 - 0.22.93 dropped it with no log). The ranges: gobj-ui ^7.25.17, gobj-js ^7.25.3.gobj-ui 7.25.18, gui_agent 0.22.92, gui_treedb 0.17.65. The treedb topics and graph views publish
EV_RECORD_WRITTENwhen one of their writes lands, to tell the host which record was written. Itstreedb_nameandrecordcame from the answer’s command frame, which carries back only the request’s__md_command__, so throughC_IEVENT_CLIand gui_agent’sC_AGENT_TREEDB_LINKthey arrived as""and{}; a link or unlink in the graph also arrived with no topic, because it sent no__md_command__. The event takestreedb_namefrom the view andrecordfrom the node the store answers (the node as written; the deleted node for a delete; the child for a link), and a link or unlink carries the child’s topic and both refs (parent_ref,child_ref). No code read the empty fields, so nothing changes on screen. Every other command-frame read in gobj-ui, gui_agent and gui_treedb reads only what its request puts in__md_command__. The ranges: gobj-ui ^7.25.18.gobj-js 7.25.4:
kw_get_int()andkw_get_real()answer like the C readers.KW_EXTRACTdeleted the value before its type was looked at, so a value that was not a number was lost and the default came back; now only a value the reader answers with is taken out (kw_get_str()had the same ordering and takes out only a string now). A value that is not a number is logged (“path MUST BE a json integer” / “... a json real”) with or withoutKW_REQUIRED, as C does.KW_WILD_NUMBER, which was ignored, reads a boolean (1/0), a string (strtoll()base 0 for int:"0x1F"is 31;parseFloat()for real) ornull(0); a list or a dict answers 0 and logs “path MUST BE a simple json element”.kw_get_int()truncates withMath.trunc()instead ofparseInt(), which read1e-7as 1. No caller relied on the old behaviour. gui_agent 0.22.93 and gui_treedb 0.17.66 take gobj-js^7.25.4.gobj-js 7.25.5, gobj-ui 7.25.19, gui_agent 0.22.94, gui_treedb 0.17.67.
kw_get_str()logs a value that is not a string (“path MUST BE a json str”, the kw traced) with or withoutKW_REQUIRED, as C does, and gives the default back; anullvalue, or a key present withundefined, is never logged, andKW_REQUIREDstill logs a missing path (the gobj-js 7.25.0 of SDK 7.25.4 logged it only withKW_REQUIRED). With no gobj,kw_get_str()no longer logs a false “gobj bad type”, and the three readers use the C wording. In gui_agent’s Schemas tab, Save, Differences and Apply work only inST_READY: a re-discovery left the old tree on screen with the three live, and each click answered “Event NOT DEFINED in state” (SDK 7.25.4 too). The discovery has a 30 s deadline (new keydiscovery unanswered), and one that fails or finds no treedb removes the old tree and says why. A write in a data treedb no longer lights “unsaved changes” (only a write intreedb_system_schemadoes). Since 0.22.95 the three also need the session (toolbar_ready():ST_READYand the agent link connected; the tooltip saysnot connected to an agent), and an Apply confirmed out of session is refused, with a toast, before the tree is taken down: in SDK 7.25.4 the Apply dialog could be confirmed after a drop, took the tree down and then failed to send (0.22.94 gated the three on the state alone). A Save or Differences out of session is said in a toast too. The treedb topics view publishesEV_RECORD_WRITTENfor a delete too, with the node the treedb answered, andcreatedis true for its +New.C_YUI_NODEdeclares its navs’EV_NAV_ITEM_CLOSEandEV_DRAWER_CLOSE_REQUESTED, and every nav event inST_OFF; the topic form declares its json viewers’EV_EXPAND_PATH(it answersEV_SUBTREE_ERROR, new keythis part cannot be loaded here): each was a latent “Event NOT DEFINED in state”. Since gobj-ui 7.25.20EV_SUBTREE_ERRORtakes{path, error | i18n, by_design}:C_YUI_JSONdraws ani18nkey witht()anddata-i18n, so the stub changes language, and logs aby_designrefusal as a WARNING (a free-texterroris drawn as it came and logged as an ERROR). The form andC_YUI_JSON_PAD(collapsed in the source) sendby_designkeys, and the “no session” of the topics, graph and gui_treedb views sends a key (gobj-ui 7.25.19 drew the form’s stub in the language it was clicked in, and logged an ERROR at each click). The toast keyraw json viewer unavailableis in every app locale.C_YUI_UPLOTno longer readsstroke/fillwithkw_get_str()(a function there is not logged). The consumers’kw_get_str()calls were audited: none reads a value that is not a string. gui_agent 0.22.94 and gui_treedb 0.17.67 take gobj-js^7.25.5and gobj-ui^7.25.19.gobj-js 7.25.6, gobj-ui 7.25.20, gui_agent 0.22.95, gui_treedb 0.17.68.
kw_find_path()answers as C (SDK 7.25.4 too): a kw that is not a dict or a list returned0, sokw_get_int()/kw_get_real()answered0andkw_get_bool()answeredfalseinstead of the default, and the log went throughgobj_short_name(null)(“gobj bad type”); from gobj-js 7.25.2 on,kw_get_str()/kw_get_bool()also logged a second line. Anullmiddle segment threw aTypeErrorwhere C gives the default. Now the answer isundefined, logged once (“kw must be list or dict: ‘<path>’”); a missing or null middle segment is logged only when verbose, a scalar one always, as the C recursion does.KW_CREATEno longer writes into a null kw. No caller passed such a kw. In the treedb graph, a__graphs__create echo goes throughapply_graphs_echo(), as the update echo does: in SDK 7.25.4 it rebuilt the saved copy of every topic from the live objects, so the unsaved edits of the other topics counted as saved and the next Save did not write them. The schema editor’s writes check the session first (SDK 7.25.4 too): out of session a write is not sent, the editor leavesST_SAVINGwith “cannot reach the treedb”, and the reconnect reloads; with a directC_IEVENT_CLIit waited inST_SAVINGwith no way out (no deployed consumer was hit). Its refused-write toast gets the key, not a text translated before it. gui_agent 0.22.95 and gui_treedb 0.17.68 take gobj-js^7.25.6and gobj-ui^7.25.20; gui_treedb sends a key for its “no session” stub.gobj-js 7.25.7, gobj-ui 7.25.21, gui_agent 0.22.96, gui_treedb 0.17.69. The raw-json viewer of
C_YUI_TREEDB_TOPICS/C_YUI_TREEDB_GRAPHanswers a drill asked out of session (a warning, no command sent,EV_SUBTREE_ERRORwith the keyno session) or refused by the transport (the refusal text); a refused drill of the whole document is also a toast. In SDK 7.25.4 the stub stayed on “loading” for the life of the viewer, even after the reconnect. gobj-jskw_set_dict_value()logs and answers -1 when a middle segment isnullor a scalar (“segment ‘<k>’ is not a dict or a list: ‘<path>’”), with the kw unchanged, instead of throwing aTypeError(SDK 7.25.4 too): every typed reader withKW_CREATEover such a path answers its default, as C does (gobj-js 7.25.6 said so, and still threw withKW_CREATE). gui_agent’s Differences has a 30 s deadline (diff_deadline, aC_TIMERchild) and a round number (diff_round): one owner that never answered kept the button off for the life of the tab, and an answer of an earlier comparison counted in the next (SDK 7.25.4 too); now the report shows what came, with the silent owners named at its top (new keydifferences unanswered, en and es), and a late answer is logged. Save’s title and aria-label say why it is off (no schema owner in this yuno, the busy reason,not connected to an agent) and change language, as Differences and Apply already did (gui_agent 0.22.95 left Save without it). The gui_agent console answers a click on a__collapsed__stub of an answer (print-tranger expanded=1) withthis part cannot be loaded here(a warning, nothing sent): the console was its viewer’s subscriber (the parent, by the CHILD model) and did not declareEV_EXPAND_PATH, so each click logged “Event NOT DEFINED in state” and the stub stayed on “loading” (SDK 7.25.4 too); type the command again withpath=. gui_agent and gui_treedb take gobj-js^7.25.7and gobj-ui^7.25.21.gobj-js 7.25.8, gui_agent 0.22.97, gui_treedb 0.17.70: a subscription that rewrites the kw gets its own (security; SDK 7.25.4 too). The port of the C fix (see “Security”).
gobj_publish_event()gave every subscriber the same kw, and each subscription’s__local__(kw_pop) and__global__changed it: one subscription forged or stripped the event of every subscriber after it and of the publisher, and a later__filter__was evaluated on the altered kw.C_IEVENT_CLI’smt_inject_event()also wrote its ievent stack,__msg_type__and the removal of__service__into the published kw, so a local subscriber after the transport got the transport’s__md_iev__. Now a subscription with a non-empty__local__/__global__gets a shallow twin (Object.assign({}, kw): its own top level, the nested values shared),__global__goes in as a copy (a receiver cannot change the subscription), and__filter__andmt_publication_filtersee the publisher’s kw; every other subscriber shares the publisher’s kw, at no cost.mt_inject_event()works on a shallow copy with__md_iev__copied deep. No JS consumer sets__global__/__local__(gobj-ui, yunos-js, wattyzer, yunovatios, estadodelaire, hidraulia, yunomusica: 0 hits), so gui_agent 0.22.97 and gui_treedb 0.17.70 only take the fixed runtime (gobj-js^7.25.8); gobj-ui’s devDependency moved to^7.25.8, with no gobj-ui release. The rule is written in docs/doc.yuneta.io/api/js/events.md.gobj-ui 7.25.22, gui_agent 0.22.98, gui_treedb 0.17.71: a raw-json drill in flight when the session drops is answered (SDK 7.25.4 too). gobj-ui 7.25.21 answered a drill that could not go out, not one already sent: the transport answers nothing in flight on a close, a plain reconnect keeps the same view and viewer, and the stub stayed on “loading” for the life of the viewer (gui_treedb, wattyzer, yunovatios; gui_agent’s link answers on the close).
C_YUI_TREEDB_TOPICS/C_YUI_TREEDB_GRAPHkeep the drilled paths in flight and answer each oneEV_SUBTREE_ERROR {path, i18n: "the connection dropped"}on the disconnect edge (EV_TRANSPORT_STATE, or the shell’sEV_CONNECTION_STATEfor the topics view), so a click once the session is back asks again; a failure that lands for a path already answered is a warning. New consumer keythe connection droppedin gui_treedb, wattyzer and yunovatios gui-common (gui_agent had it).
BREAKING¶
A keyless
tranger2_open_list()returns a list flaggedload_failed/load_failed_keyswhen keys could not be read (7.25.4 skipped them and returned the list with no sign of it); more cases now fail a load: a key flagged at startup, and a record whose content cannot be read (it used to reach the callback as NULL).treedb refuses creates of unloaded ids, snapshot ops with a partial
__snaps__,gc-assetsasset rows with a partial asset topic or an active snap (treedb_gc_files()answers NULL and takes nothing; newtreedb_gc_files2()answers a report).The agent’s audit record format changed: read-only commands are minimal,
sourcereplaces__md_iev__(no longer written),__command__is left out when it repeats the command text, and there is nocontent64(its value is<N bytes sha256:HEX>, thesha256sumof the binary; a value that is not base64 becomes<N chars, not base64, sha256 of the text:HEX>), secrets are written as<redacted>, andwrite-ttyis written as burst records{command, date, user, console, writes, bytes, until, source}with nokw.user, each field of a hop,console_purposeand the console name are redacted, and one of them longer than 1024 bytes is written as<N bytes, not scanned, sha256:HEX>. Tools that read the audit must accept both formats (files written before the upgrade keep the old one).yev_loop / gbuffer (API):
gbuffer_t.addris astruct sockaddr_storage(it was astruct sockaddr) with a newaddrlen;gbuffer_setaddr(gbuf, addr, addrlen)takes the length and returnsint(-1, logged, on a NULL gbuf or a bad length; it returnedvoid), andgbuffer_getaddrlen()is new;yev_create_sendmsg_event()takesdst_addrlen;sock_info_t.addris astruct sockaddr_storage(itsaddrlenwas already there). Rebuild every user.yev_create_connect_event()/yev_rearm_connect_event()bind a non-emptysrc_url("host:port","[ipv6]:port","schema://host:port"); a bad one, a failed resolution or a failed bind fails the connect (the in-tree callers pass NULL).mkrdir()returns -1 for a path of PATH_MAX or more, and over a path that exists and is not a directory;rmrdir()/rmrcontentdir()return -1 (logged) for a tree too long or too deep.find_files_with_suffix_array(),walk_dir_array()andget_ordered_filename_array()return -1 with the listing empty when an entry cannot be kept, and the last two also when the root is a directory that cannot be opened (7.25.4: 0; a root that is not a directory already answered -1); those three andwalk_dir_tree()return -1 when areaddir()fails (7.25.4: 0, with the entries not read missing), and a subdirectory that cannot be read fails the walk. A walk callback that returns FALSE stops the whole walk (7.25.4: only its directory); a subdirectory that cannot be opened for a cause other thanEACCES,ENOENT,ENOTDIR,ELOOPfails the walk, as does an entry whoselstat()fails for a cause other thanEACCES/ENOENT(7.25.4 skipped both); a path that does not fit in PATH_MAX, or a tree deeper than 1024 levels, fails the walk (-1, logged). A log file whose mask has a day letter (DD,W,ZZZ) and no year, last written before today, is emptied at the first record afterrotatory_open()(the yuno logs useW); anMM-only mask is judged by its month. Audit: the new secret names (andcookie_domain) are<redacted>, a string past the record’s scan budget (128 MB for all its strings) is written as<N bytes, not scanned, sha256:HEX>, and a record reaches the file at once. The agent’sdir-*commands answer -1 for a directory that cannot be listed (7.25.4: 0 and an empty list).rotatory: a log file renamed by another program is no longer noticed (a removed one still is): rotate yuno logs by copy and truncate, or remove them. The full-disk lines name the file (“rotatory(): stop logging to ‘<path>’ because full disk: ...”, was “rotatory(): stop logging because full disk: ...”), and a new line says when it resumes (“rotatory(): logging to ‘<path>’ again: ...”). The newfile callback (
rotatory_subscribe2newfile()) runs for a new name or a size rotation only, not for the same file opened again. A mask with no date letter (a fixed name) is never emptied; anMM-only mask empties its file only in a new month. Withrotatory_keep_all_old_files(), a size rotation whose rename fails keeps the file (it grows) and retries every minute of the monotonic clock. At a new day the file of the day before is not size-rotated.rotatory:
exit_on_failapplies torotatory_open()only; a later open that fails is printed once and tried again at the next record (7.25.4 exited the process).The agent removes audit files older than
audit_keep_days(default 7; 0 keeps all) at every start and at each new audit file; the first start after the upgrade removes the backlog.gobj-c: the kw a command handler gets holds its own reference of a
gbuffer(the parser increfs it,kw_update_missing()). A handler that took a second reference to work around the old double release now leaks it: take the buffer out of the kw (KW_EXTRACT) and use that reference.tranger2_write_topic_var()/tranger2_write_topic_cols()return -1 when the file cannot be written (7.25.4 ignored the write and returned 0);tranger2_create_topic()returns NULL, and does not open the topic, when atopic_versionchange cannot writetopic_cols.jsonortopic_var.json(7.25.4 ignored the failure and opened it), and returns NULL, leaving nothing on disk or in memory, when any part of a NEW topic cannot be made (7.25.4 returned the half topic); the three md2 flag rewriters return -1 on a non-master; a revive whose store another process holds exits aton_critical_error(default: exit), any other revive failure demotes to replica; C_TRANGER’smasterreads the effective state (it read the configuration);save_json_to_file()'s failed close is a CRITICAL aton_critical_error.New topics carry
marks_tm_unorderedintopic_desc.json, and a file of theirs whosetmgoes back gets a<file>.tm_unorderedmarker beside its md2 (mark-tm-orderwrites them too); newtranger2_mark_tm_order().__system__is projected whole from an installed literal: topics, columns and attributes the literal no longer declares are deleted there (with their instance history); a topic changed without a version raise now shows the literal’s content, with a warning. Seeding with no literal installed comes from the file, withc_schema_version0 unless the file is the literal.apply-schemawrites a new record,saved_schemas/<treedb>.applied.json, and refuses when it cannot write it;save-schemarefuses a draft with no topics, andapply-schemaa saved schema with no topics (withouttreedb_name, that refuses every treedb);save-schemarefuses while a projection is unfinished (new recordsaved_schemas/<treedb>.unfinished.json); a failed create, update or link of the projection leaves it unfinished and retried; a secondopen-treedbis refused, and one whose schema is refused answers -1 (7.25.4: 0 “Treedb opened!”) -- so a refused agent schema stops the agent (exit 0, not relaunched); new answer texts foropen-treedb/close-treedb/delete-treedb(“Treedb opened!” -> “<yuno>: treedb opened: ‘X’”, “Treedb closed!” -> “<yuno>: treedb closed: ‘X’”, “Treedb_name not found” -> “<yuno>: treedb ‘X’ not found”, ...); a client store locked by another process is not reconciled. The log order at open changed (the client tranger’s logs come first).msg2db: the pkey2s of an id that did not load whole whose newest message is in the damage or before it are absent instead of stale (new
msg2db_id_incomplete()); while the damaged file is the current period’s, every new message of such an id is refused (msg2db_append_message()answers NULL).tr_queue / tr2q_mqtt: the
tr_queue_t/tr2_queue_tstructs grew (rebuild their users);trq_load()/tr2q_load()return -1 when the load did not read every pending message;trq_check_backup()/tr2q_check_backup()return -1 while the backup is refused, and when the backup fails (the queue keeps its topic).tranger2_backup_topic()opens the topic again after a failure that follows its close.timeranger2:
tranger2_open_topic()returns NULL when the topic’skeys/cannot be listed (7.25.4 opened it with no keys); a key whose directory cannot be listed is flagged"unlisted"and fails every load until it can be listed; an append into it lists it again first and returns -1 while it cannot.tranger2_list_topic_names()returns NULL (logged) when the store cannot be listed (7.25.4:[], no log); C_TRANGERmark-tm-order all=1and the MQTT broker’slist-queues/clean-queuesthen answer -1.timeranger2: an append into a file flagged unreadable returns -1. A master cuts back an md2 that ends in a torn row (it writes the store at open). An md2 that 7.25.4 wrote after a torn row (also with content after its last row), or whose last two whole rows are not good records, is not cut: its key fails every load, on a master and a replica. A torn md2 is cut back with a WARNING; the other shapes log a CRITICAL (“md2 file of the key ends in a whole row that is not on a row boundary: written by 7.25.4 after a torn row; not cut, repair it by hand” or “md2 file of the key ends in a part of a row after a last whole row that is not valid: not cut, repair it by hand”), and an append into such a file is refused (“Cannot append record, its md2 file ends in a part of a row that must not be cut back: the append is refused”). A replica’s rt_disk update makes no cache cell for an md2 with no whole row yet.
tr_treedb: a write whose save fails is taken back in memory and tells no event;
treedb_autolink(),treedb_clean_node()andtreedb_replace_links()return -1 when the save fails (7.25.4: 0); a forcedtreedb_delete_node()that cannot unlink a child, or delete its key, is refused and changes nothing (7.25.4 deleted the parent); the events of a link, unlink or update are told after the child’s save (7.25.4 told link events before it); newtreedb_update_node_and_links();treedb_autolink()refuses a ref whose hook links into another column than the one the ref arrives in (“fkey reference: its hook does not link into this column”; 7.25.4 linked it through that column);treedb_save_node()returns -1 for a node that no index holds (7.25.4 wrote it back into its deleted key), and for a node whose pkey2 value was changed in place (7.25.4 wrote a new instance and moved the node’s slot);treedb_create_node()indexes the pkey2 value its record holds (7.25.4: the raw kw value); a delete of a key with several instances counts, and withforceunlinks, the children of every instance, and withoutforceis refused while a parent’s hook holds another instance;treedb_delete_instance()takes the instance out of its parents’ hooks and hands its children to the primary; after a load,treedb_get_instance()of the primary’s own pkey2 value returns the primary node itself (7.25.4: a copy without links). A write that saves instances the caller did not ask it to write (the unrefs and children of a forced delete, a refused delete taken back, a failed unlink taken back, a snap shot that clones a tagged primary) leaves the newest record of each key on the instance that wrote it before, writing that record again last when another instance wrote a newer one (untagged, no event told): the next reload takes the same primary as without the write (7.25.4 made the instance saved last the primary).tr_treedb:
treedb_create_node()refuses, for a topic with hooks, an id holding^or ofNAME_MAXbytes or more; a dict hook refuses a node of another topic with an id it holds;decode_parent_ref()/decode_child_ref()return FALSE for a part that does not fit;treedb_delete_node()refuses a node whose child topic did not load whole, and counts every child its hooks hold (a dict hook’s child whose id holds two^was not counted, and the parent was deleted);treedb_delete_instance()answers -1 when a tombstone fails. C_NODEimport-db:data.erroresis keyed by cause ("node exists","cannot create the node (see the log)"), and a create refused for another cause than an existing node is afailurein every mode (skipcounted it asignored) and stops anabortimport.C_TREEDB: every answer of every command starts with the yuno (create-topic “<role^name>: topic ‘<t>’ created in treedb ‘<db>’”, was “Topic created!”; delete-topic “...deleted from treedb ‘<db>’”, was “Topic deleted!”; “<role^name>: treedb ‘<db>’ not found”, was “Treedb_name not found: ‘<db>’”; the -403 answers); a schema whose qualified ids collide (names with dots) is refused at open, save and apply.
C_TREEDB answers carry new fields (
withdrawn,stale,broken,withdrawn_at_open,unfinished_projection,stopped, andmasterandopenedintreedbsrows) andsavedchanged meaning;withdrawn_at_opencan carry the kind"left_by_older_release"; the “nothing to save” answers ofsave-schemacarrydata.store_ahead({}, or the topics the store runs ahead of the file in use, withtopic_version,running_versionandpath);delete-treedbanswersdata: {"treedb_name", "deleted": [ids]}and answers 0 for a treedb with nothing left (it was -1); every answer ofsave-schemacarriesdata.left_by_older_release;EV_OPEN_TREEDB/EV_CLOSE_TREEDBanswer -1 with an ERROR (they did nothing). Thesaved-schema/diff-schemadiff carries a__topics_order__row and a__cols_order__row per topic (a reorder is a difference).save-schemaanddelete-treedbanswer -1 while the__system__tranger is stopped, andcreate-topic,delete-topicandapply-schemaof a treedb while its own tranger is stopped (7.25.4 read themasterof the stopped tranger:save-schemaanswered “nothing to save” anddelete-treedbdeleted nodes).save-schemawrites into__system__the position each node has in the saved schema, places a node whoseordersays nothing where the file in use declares it, and publishes a changed topic past thetopic_versionthe store runs; the tie WARNING is logged when imposing too (with"imposed": 1). C_NODEsnap-contentanswers -1 for a topic that is not of its treedb, andimport-dbof a content that is not json logs a WARNING (Protocol, with thepeername) instead of an ERROR with a stack. C_NODEgc-assetsdatais a report, not a list;instancesanswers -1 on failure;import-dbanswers -1, with a comment, when it aborts or anything fails (it answered 0), and counts a link it cannot make as alink failure;export-dbwith nofilenamenames the file<treedb>-<schema_version>-<date>.trdb.json;link-nodes/unlink-nodestake a child ref split at its first^(an id with^is accepted).The command comments of C_NODE and C_AUTHZ changed: they start with the yuno and name what they did, for example “Node update!” -> “<role^name>: Node update! ‘<id>’ of topic ‘<t>’”, “Nodes linked!” -> “<role^name>: Nodes linked, ‘<topic>^<id>’ to ‘<topic>^<id>^<hook>’”, “Node deleted” -> “<role^name>: Node deleted, ‘<id>’ of topic ‘<t>’”, “Snap deactivated” -> “<role^name>: Snap deactivated, treedb ‘<db>’”, “Snap activated: ‘<name>’” -> “<role^name>: Snap activated: ‘<name>’”, and “Cannot activate snap ‘<name>’: <last log message>” -> “<role^name>: snap not found: ‘<name>’” or “<role^name>: cannot activate snap ‘<name>’ (see the log)”.
ip lists: C_TCP_S refuses at accept a peer in the yuno’s
denied_ips(7.25.4 refused it only at the login of an authenticating gate); C_UDP_S drops the datagrams of a peer indenied_ips, or not inallowed_ipswithonly_allowed_ips(7.25.4 heard every peer);[::1]and[::ffff:127.0.0.x]are exempt like127.0.0.x.is_ip_allowed()/is_ip_denied()look up an IPv6 peer by its ip without the port.C_UDP_S: its output event is
EV_STOPPED(the event table declared the stateST_STOPPED, and nothing was published); a datagram the kernel refuses is dropped (it stopped the server); theEV_RX_DATAgbuffer always carries the peer as its label. C_IEVENT_SRV stops its protocol gobj when it is stopped. New yuno attributeenable_subscription_authz(off: nothing changes); C_NODE’sreadcarries the alias__subscribe_event__(itsauthzslisting shows it), and so does C_TRANGER’s, whoseEV_TRANGER_RECORD_ADDEDisEVF_AUTHZ_SUBSCRIBE. C_UDP_S has a new stat,rxRefusedMsgs, and its refusal WARNINGs come at the first drop of a cause and then at most once a minute, with the new fieldsdropped,rxRefusedMsgsandnext_warning_in_ms(7.25.4 had no refusal); a host that keeps theEV_RX_DATAgbuffer keeps it (the next read takes a new one). C_IEVENT_SRV keeps only__first_shot__of a peer’s__config__(__hard_subscription__,__own_event__,__rename_event_name__and any other key are dropped with a WARNING), and a channel’s close removes its subscriptions by force, hard ones included.C_CHANNEL’smt_enablestarts its tree (gobj_start_tree()), andC_IEVENT_SRV’smt_startstarts its protocol gobj.gobj_unsubscribe_event()no longer counts a hard subscription it leaves in place as removed (it logs a WARNING).ip lists: the
add-/remove-allowed/deniedip commands refuse a text that is not a numeric ip (7.25.4 stored it) and store the canonical form (2001:DB8::1->2001:db8::1); a link-local address inallowed_ipsneeds its interface;remove-denied-ip/remove-allowed-ipof an ip that is not in the list answer -1 (7.25.4: success); the stored lists are normalised at the first start, and an entry that is not an ip is dropped (see “Upgrade steps”). C_IOGATE’s answer to a badchannel_nameregular expression is “<role^name>: channel_name is not a valid regular expression: ‘<re>’” (was “regcomp() failed”).gobj-js 7.25.1 - 7.25.5: the typed
kw_get_*readers answer like the C ones.kw_get_str()returns its default as given (it returnedString(default)) and logs a value that is not a string;kw_get_list()no longer wraps its answer;kw_get_dict()no longer turns anulldefault into{};kw_get_bool(),kw_get_int()andkw_get_real()read a value of another type as C does, and log it;KW_EXTRACTtakes out only a value of the reader’s type. gobj-js 7.25.6:kw_find_path()answersundefinedfor a kw that is not a list or a dict (it answered0, sokw_get_int()/kw_get_real()gave0instead of the default) and for a missing or null middle segment (it threw). gobj-ui 7.25.20:EV_SUBTREE_ERRORtakes{path, error | i18n, by_design}-- a host that sends a translatederrorshould send the key asi18n(a free-texterroris still drawn as it came, and logged as an ERROR); aby_designrefusal is a WARNING. See the JS section.yuno_agent:
find-new-yunosmarks the rows already registered at the new release, andcreate=1skips them (it failed on each); thedir-*failure answer is “<role^name>: cannot list ‘<dir>’, see the log”. MQTT broker:list-queues queue=<name>answers -1 for a queue that cannot be opened or read whole (7.25.4: 0).gobj-c:
kw_set_dict_value()overwrites a key that exists (7.25.4 kept the old value) and returns -1, logged, when it writes nothing (a scalar or null in the middle of the path, an index outside its list, an empty path). Newkw_twin().gobj_publish_event()gives a subscription with a non-empty__local__/__global__its ownkw_twin()of the event: such a subscriber no longer sees, nor changes, what another one changed, and the__filter__of every subscription sees the publisher’s kw (7.25.4 applied each__local__/__global__to the one kw of every subscriber).gobj_subscribe_event()of a subscription that matches a hard one returns the hard one, with a WARNING (7.25.4 made a second one), and a hard subscription over a plain one replaces it. A log call leaveserrnoas it found it, and a failedmkrdir()leaves the cause inerrno;rmrdir()/rmrcontentdir()return -1 on areaddir()error (7.25.4:rmrcontentdir()returned 0).C_IEVENT_SRV: of a peer’s subscription only
__filter__, the allowed__config__keys and the__global__keys that do not start with_(norgbuffer) are kept; the rest is dropped with a WARNING. New attributesmax_subscriptions(default 5000) andmax_subscription_size(default 16384): a subscription beyond them is refused. The withdrawal of a subscription is filtered the same way, and matches one with a__global__.C_IEVENT_SRV/C_IEVENT_CLImt_inject_event()work on a twin of a shared kw. C_GSS_UDP_S: newmax_channels(1024),max_pending_bytes(8 MB) andmax_frame_size(1 MB, the old size of every frame buffer, which now starts at 4 KB); a frame longer thanmax_frame_sizeis delivered cut, without an ERROR.C_TCP: a stop always publishes
EV_STOPPEDonce, also the stop of a client that is already disconnected (7.25.4 published nothing there and kept the connect event); it can come insidegobj_stop(). A write that cannot be created or started drops the connection, with an ERROR (7.25.4 ignored it, and the connection hung).C_TRANGER:
add-recordappends (master-only,write; 7.25.4 answered -1 “Pending to review”); the pushes of a live list carryrt_id, itslist_id;close-rt,close-iterator,close-list,get-pageandget-list-dataof a handle opened by another session answer -403. A handle’s owner is(src, user): the channel of the command and its__username__, so through the agent (command-yuno, one link for every operator) a handle is its user’s, and a relayed command without a user owns nothing it did not open.C_TREEDB:
save-schemapublishes every topic whose saved place is not the file’s (anorderrow inchanges), writes theorderof a node with more than one parent as 9999 (says nothing) and names it indata.places_not_written, and answers -1 withdata.twinswhen two topics of a treedb (or two columns of a topic) have one name;saved-schema’sdraft_changednames the topics whose places the draft shifts; a failedopen-treedbkeeps the saved schema (withdrawn_at_open.saved_schema_version0). tr_treedb:treedb_link_nodes()refuses a topic under a name its treedb already has (“Treedb already has a topic with this name”); an unlink clears the parent’s ref in every instance of the child, and a delete counts (withoutforce) or clears (withforce) the instances of a child that name it and that no hook holds.timeranger2:
tranger2_delete_key()returns -1 when thestat()of the key directory fails with anything butENOENT(7.25.4: 0, and the key was dropped from the cache and announced); a topic whose key listing meets astat()failure (nod_type) does not open.tr2q_check_backup()does not back up while messages are queued (answers 0);tr2q_move_from_queued_to_inflight()answers -1, the message kept queued, when its content cannot be read.yev_loop: the close of an fd takes back the submissions of other events on it that the kernel has not taken; those events complete as
STOPPED,-ECANCELED.yev_loop_destroy()does not free a zero-copy send whose notification has not come after 5 s (a WARNING).rotatory: the size limit is compared in bytes (7.25.4: whole megabytes, so a limit of 8 MB rotated at 9 MB).
yuno_agent:
create-yunoanswers -1 for a release name longer thanNAME_MAX(7.25.4 stored it cut and answered 0).gobj-js 7.25.7:
kw_set_dict_value()answers -1, logged, when a middle segment isnullor a scalar (it threw). gobj-js 7.25.8: the publish twin rule, as in C.C_IEVENT_SRV: the back-metadata stored with a peer’s subscription is the reversed top record of the frame’s ievent stack only (7.25.4 stored the frame’s whole
__md_iev__), and a subscription whose routing is bigger thanmax_subscription_sizeis refused. A session frame without its routing closes the channel (7.25.4 processed it). A frame that repeats a subscription the channel holds leaves it as it is (no secondmt_subscription_added(), no second first shot; 7.25.4 deleted and made it again), and one that overrides a held subscription replaces it. A (un)subscription of an event that is not public is refused and the channel goes on (7.25.4 answered -1 and left the channel deaf); a command or stats request without its command, or with aservicethat is not a string, gets a negative answer. Its peer-frame logs are WARNINGs at most one each 10 s per kind, withsuppressed(7.25.4: a line per frame, most of them ERRORs).gobj-c: “subscription(s) REPEATED, will be deleted and override” and “Hard subscription REPEATED, the one there is kept and returned” carry no stack, and their kw is capped to 256 bytes. A subscription is matched with its kw as it is stored (
__hard_subscription__,__own_event__and__rename_event_name__taken out,__original_event_name__put in with a__global__), ingobj_subscribe_event()andgobj_unsubscribe_event(): a repeated__own_event__/__rename_event_name__subscription replaces the one there, and an unsubscribe with the same kw removes it (7.25.4 made a second one, which could not be withdrawn).C_TCP_S: a refusal by the ip lists is logged at the first refusal of each cause and then at most once a minute per cause, with
refused,refusedConnxsandnext_log_in_ms(7.25.4: one INFO per refused connection); new statrefusedConnxs.C_UDP_S: a
udps://url is refused at start (ERROR, -1; withexitOnError, the default, the yuno exits with 0);use_sslis always false. A read that cannot be started again stops the server, with an ERROR (7.25.4 went on, deaf); a stop with a datagram in flight ends inST_STOPPEDwithEV_STOPPED.C_GSS_UDP_S: a published frame carries its peer’s label and address.
EV_SEND_MESSAGEsends to the gbuffer’s address, or else to the known peer its label names, and answers -1, with an ERROR, for a gbuffer with neither (7.25.4 logged “UDP channel NOT FOUND” for a send by address, and a send by label had no address).MQTT (C_PROT_MQTT2): an ack of the wrong packet (a PUBACK for QoS 2, a PUBCOMP for QoS 1 or before the PUBREC, a PUBREL of the wrong QoS in the client) is a protocol error: a WARNING, the connection closed (an MQTT 5 peer gets DISCONNECT 0x82), and the message kept in flight (7.25.4: an ERROR, and the message removed as delivered). “Message not found in trq_out_msgs” is a WARNING (was an ERROR). The out window of a queue is the lower of the broker’s
max_inflight_messagesand the client’s Receive Maximum, also when the broker’s is 0.list-queues qos=combines with thepending=condition (7.25.4 replaced it).C_TREEDB:
save-schemawrites theorderof a node that hangs from more than one parent as 9999 (says nothing), so a node that stops being shared is placed by the file of the parent that remains.timeranger2 / gobj-c: without
d_type,find_keys_in_disk()askslstat(): a symbolic link inkeys/is not a key (7.25.4 followed it), and anEACCESentry is a key, as withDT_DIR;find_files_with_suffix_array()lists an entry whoselstat()fails withEACCES(7.25.4 skipped it). tr_queue / tr2q_mqtt: a queue without topic takes a topic the tranger has open again, and tries the open again when its cause may be gone (see “Data loss and integrity”).yev_loop (internal, no API change):
get_sqe()clears the fd and theuser_dataof every submission entry it hands out, so an entry of the queue is told by the event that owns it, never by what an earlier operation left in the ring slot.yuno_agent audit: see the audit record bullet above (
user, the hops,console_purposeand the console name redacted).Log texts -- match on the new ones if you alert on them:
timeranger2: “Cannot read last record, md2 file corrupted” is gone (see the md2 bullets above). “Cannot read first/last record of md2 file” are now “Cannot read a record of md2 file, read FAILED” or “..., short read” (with
row:first,last, orend/before lastin the torn-tail check); a truncated md2 or content logs “... short read” instead of “read FAILED”, and a write that stops part way logs “... short write: the file size limit or the disk is full” (withwritten/expected) instead of “... write FAILED”.timeranger2: “Bad data, anystring2json() FAILED.” is now “Bad data, the content of the record is not json”, or “Cannot read the record, this process has not the memory to parse its content (MEM_MAX_BLOCK)”.
timeranger2: “Cannot mark md2 file as unordered, file_id too long” / “..., a reload will misread its time range” are “Cannot mark md2 file, file_id too long” / “Cannot mark md2 file, a reload will misread its time range”.
timeranger2: “next rowids not consecutive” / “previous rowids not consecutive” are “next segment begins before the row just read” / “previous segment ends after the row just read”.
timeranger2: “Master lock NOT retaken after a stop, go on as not master” is three texts, “Master lock NOT retaken after a stop: cannot open the lock file, go on as not master”, “...: another process holds it, ...” and “...: flock() FAILED, ...”.
timeranger2: the ERRORs “what id?” and “Disk already exists” of an rt disk are the WARNINGs “Invalid rt id (empty)” and “rt disk id already in use by the same creator, refused”.
timeranger2, new: “Cannot re-create topic_cols.json for a new topic_version: the topic keeps its version, and is not opened”; a key or a topic that cannot be listed: “key directory cannot be listed when its cache was built: every load of the key says load_failed”, “The history of the key is not whole: its directory could not be listed when its cache was built”, “Cannot append record, its key cannot be listed: its row would follow rows no cell counts”, “Cannot open topic: its keys cannot be listed”, “Cannot list the keys of the topic” (and “, readdir() FAILED”), “New records of the key not read: its directory cannot be listed”, the INFO “key directory listed again: its files are counted, and the key is not flagged”; “Cannot list the topics of the store” (and “, readdir() FAILED”); a backup: “Backup of topic failed: the topic is opened again as it was, not backed up” (WARNING), “Backup of topic failed, and the topic cannot be opened again”, “Cannot back up topic, its directory not found”, the CRITICAL “Backup of topic failed, and the backup cannot be moved back: the data of the topic is in the backup”; tr_queue / tr2q_mqtt: “Queue backup failed: the queue goes on in its topic, not backed up”, “Queue backup failed, and the queue has no topic”, the ERROR “Queue without topic, it cannot be opened” and the INFO “Queue topic taken again”; a create: the CRITICAL (at
on_critical_error, after the removal) “Cannot create topic: it is not whole, what was made is removed”; a faileddelete_key: “Cannot index the key again after its files changed: the filtered iterator is empty”, “Cannot delete key, stat() of its directory FAILED” (the key is not deleted, dropped from the cache nor announced), “Cannot tell the rt_disk feeds that a key was deleted, opendir() of disks/ FAILED”, “Cannot tell every rt_disk feed that a key was deleted, readdir() of disks/ FAILED”; “Cannot list the keys of the topic, stat() FAILED”; the CRITICAL “Cannot append record, its md2 file cannot be opened: its content was cut back”; the ERRORs “md2 file of the key unreadable when its cache was built: every load of the key says load_failed” and “The history of the key is not whole: a md2 file of it could not be read when its cache was built”; tr_queue / tr2q_mqtt: the ERRORs “Queue backup refused: its last load did not read every pending message” and “Queue loaded without some of its messages: its first_rowid is not moved nor saved”; the CRITICAL “Cannot cut back a md2 file that ends in a part of a row: the file is damaged”; the ERRORs “Cannot append record, its file is flagged unreadable: its row would follow rows no cell counts” and “Cannot mark the topic: it is not marked; the markers written stay, and the cells read keep their whole ranges”. “Cannot list the keys of the topic, stat() FAILED” is no longer logged forEACCES(the entry is a key, as withd_type).msg2db, new: the ERRORs “msg2db: a key whose history did not load whole: only the messages newer than the damage are served, a pkey2 whose newest message was not read is ABSENT and its state unknown (msg2db_id_incomplete)”, “msg2db: the damaged file of the key is the file of the current period: every new message of the key is REFUSED until the file is repaired or the period changes” and “msg2db: load_failed_keys holds an item that is not a string: that key is not reloaded”, and “Msg2Db topic NOT FOUND” (
msg2db_id_incomplete()).tr_treedb: new ERRORs “A write that did not reach the disk could not be taken back whole in memory: ...”, “A refused delete cannot put back a child it had unlinked: ...” and “Cannot write again the newest record of a key: at the next open another instance of it is the primary”, and the CRITICAL “Cannot write again the newest record of a key after its clone: at the next reload the photo is its primary” (a snap shot).
tr_treedb: the ERROR “Child node without fkey field” (one per node, at every open) is the WARNING “An fkey column is filled by no hook: its refs link nothing”, once per column per open, with
nodes_with_refs. “tranger2_delete_instance() FAILED” is gone, and so are “wrong array child hook type” and “md_treedb not found” of the delete guard. New: “Wrong reference: a part of it is too long”, “Invalid ‘id’: it holds a ‘^’, the separator of a reference”, “Invalid ‘id’: too long to be part of a reference”, “Cannot link, the dict hook holds a node of another topic with this id”, “A dict hook holds a node of another topic with this id: this link is not loaded”, “Child data not found in dict parent hook: its slot holds a node of another topic”, “Cannot link, the reference the child has is too long”, “Cannot delete node: a topic its hooks hold did not load whole, a child that did not load may hang from it”, “Cannot delete instance, a row of it cannot be tombstoned: the instance stays, its newest rows alive”, “Cannot create node, its id has records on disk that could not be loaded”, “Cannot shoot a snap: snaps did not load whole, the active snap is unknown”, “Cannot activate a snap: snaps did not load whole, the active snap is unknown”, “Cannot deactivate a snap of too many active ones, it stays active on disk”, “Cannot save a node that is being deleted”, “treedb topic loaded WITHOUT the whole history of keys that cannot be read: ...”, “Cannot build the reference of a node: a part of it is too long, or holds a ‘^’”, “Cannot save a node that no index holds: its record would bring back what was deleted”, “Cannot save a node whose pkey2 value changed in place: its record would be another instance”, “Cannot put an instance back into the hook of a parent: its slot holds another node”, “A dict hook holds a node of another topic with this id: the child of a deleted instance is not handed to the primary”, “Cannot hand the children of a node without treedb_name, topic_name or id”, “Cannot look for the parents of a node without treedb_name or id”, “Treedb already has a topic with this name” (withidandsibling_id; the column text now carriessibling_idtoo), “A write taken back cannot save again an instance that stopped naming a parent: on disk it names it no more”. “Cannot delete node: has down links” carrieschildrenandunheld_instances. “Child data not found in list parent hook” / “... dict parent hook”, “Duplicate fkey on load, deduping parent hook” and “delete_primary_node() FAILED” are no longer logged for another instance of the key (a relink or unlink of a sibling instance, an instance held through a list hook, a key only the secondary indexes hold).tr_treedb:
treedb_autolink()'s “update_node, new link: parent node not found” is “fkey reference: parent node not found”, the text oftreedb_replace_links(). “cannot delete asset, a snapshot still links it” is also “cannot delete asset, cannot tell whether a snapshot links it (see the log)” when the snapshots cannot be read.gobj-c:
mkrdir()'s “Not a directory” is “Not a directory: the path exists and is not a directory” (7.25.4 never logged it: its check could never be true). New: “Cannot list directory, no memory for an entry” (find_files_with_suffix_array()), “Cannot list directory tree, no memory for an entry” and “Cannot list directory tree, the directory cannot be opened or read” (walk_dir_array()), “Cannot list directory, readdir() FAILED” (find_files_with_suffix_array()), “Cannot read directory, readdir() FAILED” (walk_dir_tree()), “Cannot close json file, what was written may be lost” (save_json_to_file(), a CRITICAL). The directory walks: the ERRORs “Path too long, the directory cannot be walked”, “Tree too deep, the directory cannot be walked”, “Cannot list directory, stat() FAILED” (find_files_with_suffix_array()) and “stat() FAILED, the directory cannot be walked” (was “stat() FAILED”), the WARNING “Cannot open subdirectory, it is skipped” (7.25.4 logged nothing forEACCES/ENOENTand the ERROR “Cannot open directory” for another cause); the root that cannot be opened is the ERROR “Cannot open directory”.gobj_unsubscribe_event(): the WARNING “Hard subscription not removed, only gobj_unsubscribe_list() with force removes it”;gobj_subscribe_event(): the WARNING “Hard subscription REPEATED, the one there is kept and returned”; it and “subscription(s) REPEATED, will be deleted and override” carry no stack, and theirkwis a compact dump capped to 256 bytes (7.25.4: the whole kw, with a stack). New: “Cannot remove directory, readdir() FAILED” (rmrdir()) and “Cannot remove the content of directory, readdir() FAILED” (rmrcontentdir()).kw_set_dict_value(): new ERRORs “path empty”, “json_object_set() FAILED”, “json_array_set() FAILED”; a scalar middle segment logs “long path” (its “short path” is gone).kw_twin(): “kw to twin must be an object”, “json_copy() FAILED”. A caller’s CRITICAL after a failedmkrdir()(“Cannot create TimeRanger subdir. mkrdir() FAILED”) now names the real errno (it said “Success”).rotatory (stdout and syslog): “_rotatory(): vfprintf() FAILED, <err>” is “_rotatory(): fwrite() FAILED, ‘<path>’, <err>”; “_rotatory(): Cannot rename ‘<path>’ to ‘<old>’, <err>” ends with what follows (“, the file is emptied” or “, the file is kept and grows, the rename is tried again every minute”) and is printed once until “_rotatory(): the size rotation of ‘<path>’ works again”. New: “_rotatory(): writing to ‘<path>’ again”, “_rotatory(): N pieces of ‘<path>’ in one day, the last one is replaced”, “_rotatory(): Cannot empty ‘<path>’, <err>”, “_rotatory(): ‘<path>’ is open again”. “_rotatory(): Cannot open ‘<path>’ file, <err>” (and “Cannot create ‘<path>’ file” / “directory”, “_rotatory_truncate(): Cannot open ...”) is printed once until the file opens again, and no longer ends the process.
C_UDP_S, new: the ERRORs “Cannot send datagram: dropped”, “UDP: read FAILED, the server stops listening”, “UDP: no memory for the next read, the server stops listening”, “UDP: the read cannot be started again, the server stops listening”, “Cannot start the read of the UDP server” and “A secure url (udps://) is not supported by C_UDP_S: there is no DTLS” (the ERROR “Cannot set ‘trace_tls’, ‘crypto’ is not a dict” of C_UDP_S is gone); the WARNINGs “UDP: datagram refused by the kernel, dropped”, “UDP server stopped: the datagrams waiting to be sent are dropped”, “UDP_S: Ip denied, datagram dropped” and “UDP_S: Ip not allowed, datagram dropped” (the last two at the first drop of each cause, then at most once a minute, with
dropped,rxRefusedMsgsandnext_warning_in_ms).C_TCP_S, new: the INFO “TCP_S: Ip denied”. It and “TCP_S: Ip not allowed” are logged at the first refusal of each cause, then at most once a minute per cause, with the new fields
refused,refusedConnxsandnext_log_in_ms(7.25.4: one line per refused connection); themsg2of “TCP_S: Ip not allowed” is the same text (it began with a globe sign).C_IEVENT_SRV, new: the ERROR “No permission to subscribe event”; the INFO “UNSUBSCRIBING event never subscribed, its subscription was refused” for the withdrawal of a refused subscription, and the WARNING “UNSUBSCRIBING event matches no subscription of this channel” for any other unsubscribe that matches nothing (both in place of the ERROR “No subscription found” with a stack); the WARNINGs “SUBSCRIBING config keys a peer may not set, ignored” (with
keys) and “SUBSCRIBING config is not a dict, ignored”; the WARNINGs “SUBSCRIBING keys a peer may not set, ignored” (itskeysname the__local__, any stray key and each dropped__global__`<key>), “SUBSCRIBING refused, bigger than max_subscription_size” (field,size) and “SUBSCRIBING refused, the peer holds max_subscriptions” (once, until the peer is under the cap). The dropped-keys WARNINGs, the size refusal and both UNSUBSCRIBING texts are logged at most once per kind, per channel, per 10 s, withsuppressed(the no-match WARNING dumps at most 256 bytes of the kw, withkw_size); “No permission to subscribe event” is logged at the first refusal of each service and event on a channel. New: the WARNINGs “SUBSCRIBING refused, its routing is bigger than max_subscription_size”, “SUBSCRIBING repeated, the one held is kept”, “SUBSCRIBING overrides one held, it is replaced”, “Frame without its routing (md_iev ievent stack), channel closed” (in place of the ERRORs “md_iev NOT FOUND” / “md_iev stack NOT FOUND”, with a stack) and “Request without the command or stats it asks, refused”, with the answer “<role^name>: request without the command or stats it asks”. “It’s not my role” / “It’s not my name” (ERRORs) are the WARNINGs “It’s not my role, channel closed” / “It’s not my name, channel closed”, with the routing capped. These ERRORs are now WARNINGs at most one each 10 s per kind and channel, withsuppressedandusername: “event ignored, dst_service not authorized for this channel”, “event ignored, service not found”, “SUBSCRIBING event ignored, not PUBLIC or PUBLIC event”, “UNSUBSCRIBING event ignored, not PUBLIC or PUBLIC event” and “Not authorized to request stats of a different service”; the WARNING “Service not found” is limited in the same way. The peer’s strings in them are capped to 128 bytes, and the service named in their answers (“Service not found: ‘<service>’”, ...) to 60.C_GSS_UDP_S, new: the WARNINGs “Too many peers, datagrams of new peers dropped” (on the transition), “Too many bytes in unfinished frames, datagram dropped with the unfinished frame of its peer” and “Frame without end within max_frame_size, delivered cut” (at most once per 10 s, with a count), and the ERROR “EV_SEND_MESSAGE without a peer: no address, and its label names no known peer; dropped”; the ERROR “gbuffer_append() FAILED” of a frame that did not fit, the frame buffer’s “no memory”, and the ERROR “UDP channel NOT FOUND” of a send by address are gone.
C_TCP, new: the ERRORs “Cannot start a write: the connection is dropped” and “Cannot create a write: the connection is dropped”.
C_TRANGER, new: the WARNING “Handle of another session, refused” (
kind,id,src,user), the ERROR “Handle not registered, its owner user cannot be stamped” (an internal check) and the -403 answer “<role^name>: <kind> ‘<id>’ is not yours: another session opened it”;add-recordanswers “<role^name>: record added to topic ‘<t>’, rowid <n>”, “<role^name>: What record? It must be a dict with the pkey of topic ‘<t>’”, “<role^name>: cannot add the record to topic ‘<t>’ (see the log)” and “<role^name>: tranger ‘<name>’ is READ-ONLY, this yuno is not its master: add-record runs on the master” (its ERROR “TODO pending to review” and answer “Pending to review” are gone).MQTT (C_PROT_MQTT2 client): the WARNING “removing an inflight qos2 dup message” is “QoS 2 message received again (dup): it replaces the copy waiting for its PUBREL”; new ERROR “Cannot load the content of a queued message: it stays queued”; a client’s QoS 2 PUBREL no longer logs “QoS mismatch”. “QoS mismatch” of an ack of the wrong packet, “Unexpected message state for QoS 2” and “Message not found in trq_out_msgs” are WARNINGs (
MSGSET_MQTT, withclient_idandpeername), where 7.25.4 logged ERRORs.C_IOGATE: “regcomp() failed” is “<role^name>: channel_name is not a valid regular expression: ‘<re>’”.
C_YUNO ip lists, new: the WARNINGs “ip list entry renamed to the form a peer is looked up by” and “ip list entry dropped, it never matched a peer” (with
entryand the cause), the ERROR “ip list is not a dict”; the answers “<role^name>: ‘<ip>’ stored as ‘<canonical>’”, “<role^name>: ip ‘<ip>’ is not in allowed_ips” / “... is not in denied_ips”, and the refusal of a text that is not an ip, with its cause.MQTT broker, new answers:
list-queues“<role^name>: cannot list the topics of the store, see the log”, “<role^name>: cannot open the queue ‘<name>’, see the log” and “<role^name>: the messages of the queue ‘<name>’ cannot all be read, the list is PARTIAL, see the log”;clean-queues“<role^name>: cannot list the topics of the store, nothing cleaned, see the log”.yuneta_agent, new: the WARNING “Audit: bad md_iev from a peer, written as empty”, the ERRORs “audit scan: a replacement behind the copy” and “Audit: a field of a peer cannot be written” (internal checks: nothing of that text is written); the
dir-*answer “<role^name>: cannot list ‘<dir>’, see the log”;find-new-yunosmarks a row “already registered, pending promotion (deactivate-snap): create-yuno ...”, comments “<role^name>: N yuno(s) already registered at the new release, pending promotion: run deactivate-snap”, and logs the ERROR “yuno release name too long”;create-yunoanswers -1 “<role^name>: Yuno ‘<role>.<name>’: release name too long (‘<binary version>’ + ‘<config version>’)” after that ERROR (7.25.4 stored the name cut, with no log).C_TREEDB: “Topic from C differs from the one in use, but its topic_version is not higher: not applied” is “Topic from C declares other columns than the store runs, without raising its topic_version past it: the store keeps running its own”; “Schema from C takes over the file in use while system holds a draft saved over it: ...” is gone, replaced by “Schema from C withdrew work on the schema at open”; a record write failure logs “Cannot write a record of saved_schemas/”; “Treedb schema not projected in system” is gone (a
delete-treedbof a schema that is not there answers 0). New: the WARNINGs “TreeDB schema ids moved to qualified names only in part: ...” and “Schema from C has the schema_version of the dynamic schema in use but another content: ...” at every open that ties (7.25.4: only whenc_schema_versiondiffered); the ERRORs “EV_OPEN_TREEDB is not implemented: ...” / “EV_CLOSE_TREEDB is not implemented: ...”; C_NODE “Cannot write the export of the treedb”; the WARNING “Schema file in use declares other columns than the store runs, at a topic_version behind the store’s ...” (withpath; it ends “... to keep what runs, edit the topic in system to those columns, then save-schema and apply-schema; to run the file’s, raise its topic_version above running_version (and the schema_version) in the schema from C”); the ERRORs “Topic of a saved schema not found in system, its place is not written”, “Column of a saved schema not found in system, its place is not written” and “Topic of a saved schema not found in system, its topic_version is not written”; the “nothing to save” answers ofsave-schemaname a topic the store runs ahead of the file (“...; the store runs ‘users’ ahead of it with other columns, which the draft cannot say: ...”). The tie WARNING carries"imposed": 1on an imposed open. C_NODEimport-dbof a content that is not json logs the WARNING “frame is not json” (Protocol,peername) in place of the ERROR “json_load_callback() FAILED” with a stack. C_TRANGERmark-tm-orderanswers “<role^name>: cannot open topic ‘<t>’ (see the log)” for a topic that cannot be opened. The ERRORs “Schema refused: two topics of the treedb in system have the same name, unlink or rename one of them” / “Schema refused: two columns of a topic in system have the same name, ...” (withname,first,second), the ERROR “Schema refused: two elements have the same qualified id in system (a name with a dot), rename one of them” (withid,first,second), and thesave-schemaanswer “<role^name>: cannot save the schema of ‘<db>’: the topics ‘<topic1>’ and ‘<topic2>’ have the same name ‘<n>’, and a schema keeps one per name: unlink or rename one of them”; a save comment may end “; N node(s) hang from more than one parent and get no place of their own (oneordercannot say a place in each): see places_not_written”; the INFO “Saved schema kept: the open that installs the schema from C did not open; the next one that installs a schema from C and opens withdraws it”.C_NODE: “Cannot save the node after its links (autolink)” is gone; a create logs “Node created, but its links cannot be saved (autolink): the node stays without them”. An
update-nodeof a node that does not exist logs “Cannot update node: it does not exist”, or withcreate“Cannot update node: it does not exist and it cannot be created (see the previous log)” (7.25.4 logged the last message of the process, or “treedb_create_node failed”). New: the ERRORs “Cannot delete a node on a READ-ONLY replica”, “Cannot link nodes on a READ-ONLY replica” and “Cannot unlink nodes on a READ-ONLY replica” (7.25.4 had only “Cannot write a node on a READ-ONLY replica”).yev_loop: “io_uring_get_sqe() FAILED” is gone. New: the WARNINGs “Submission queue full and the kernel takes nothing: kept for the next cycle” and “Submissions the kernel did not take: submitted again at each cycle”; the ERROR “Submissions not taken by the kernel for many cycles of the loop: their operations wait”; the INFO “Submissions taken by the kernel again”; the CRITICALs “No memory to keep a submission” and “No memory to keep a completion: submission handed over as it is”; the ERROR “No memory to keep a submission: <what did not happen>”; the ERROR “Loop destroyed with events whose completions did not come: freed” (with
cancel_submitted), and before it, when the cancel cannot be submitted, “Submission queue full: the cancel of the events left is NOT submitted, their completions may not come”. A connect with a badsrc_urllogs the ERROR “Bad src_url: cannot bind the connect” (new), “getaddrinfo() src_url FAILED” (unchanged) or “bind() src_url FAILED” (was “bind() FAILED”). “Cannot start event: sendmsg addr NULL” is “Cannot start event: sendmsg addr NULL or bad addr length”. New: the WARNING “An fd is closed with submissions of other events on it that the kernel did not take: taken back, completed as canceled”, the ERROR “...: dropped, the event will not complete”, the WARNING “Loop destroyed with zero-copy sends whose notification did not come: NOT freed, the kernel may still read their gbuffer”, and the ERRORs “Cannot replace the gbuffer of an event with an operation in the kernel” (with a stack) and “io_uring_submit_and_wait_timeout() FAILED”.tr2migrate: “Bad data, json_loadfd() FAILED.” is “Bad data, the content of the record is not json”; a json file it cannot load logs the texts of
load_json_from_file().
Known limitations¶
default: {}placeholders are dropped by save + apply, so arequiredcolumn whose literal really declared'default': {}loses it.getaddrinfo()runs synchronously inside the event loop (connect, source bind, listen): a lookup blocks every gobj of the process, and a dead firstnameservercosts ~6 s per lookup (3 s x 2 per query). The cache of 7.8.2 makes it one lookup per host per TTL.OIDC discovery requires
end_session_endpoint: against Auth0, and some Cognito setups, which do not publish it, discovery fails; settoken_endpointandend_session_endpointexplicitly for those IdPs.Each Live card of a C_TRANGER opens its own rt_disk feed, one inotify instance each: many Live cards can use up
fs.inotify.max_user_instances(128 by default), as they did on one node.An md2 truncated to 0 rows behind the yuno’s back loses its rows, as in 7.25.4 (now with a warning); check the
.jsonsize before repairing.With
enable_subscription_authzoff (the default), theEV_TREEDB_NODE_*feed of a treedb and theEV_TRANGER_RECORD_ADDEDfeed of a C_TRANGER reach any authenticated user, while C_NODE’s and C_TRANGER’s commands askreadwith every gate off. A project yuno that publishes the feed again under its own service (thedb_history*yunos) is checked only if it flags its eventsEVF_AUTHZ_SUBSCRIBEand aliases a permission.
v7.25.4 (2026-09-23)¶
The 2026-09-23 review of the 2026-09-22 treedb/timeranger work¶
A third, read-only review of the fixes of 2026-09-16 and 2026-09-22 (six
reviewers, repros against outputs/lib) found 12 mediums and ~35 lows that
no test caught. All of them are fixed below, each with a test that was red
before where one could be written.
BREAKING
treedb on a replica:
treedb_create_node()is refused up front, andtreedb_store_files()refuses bytes (“Cannot store a file, NO master”).gobj_update_node()on a C_NODE replica answers NULL for any non-volatil update, before anything moves.tranger2_open_rt_disk()answers NULL for an id that a live feed of the topic already uses (any creator) and for an id longer thanNAME_MAX. A refused id from a peer is a WARNING without a stack, and “Cannot open rt” is no longer logged on top of it.tranger2_open_list()of one key answers NULL when its history load fails. Iterators carry two new fields,segments_stampandload_failed.Scans: a
tcondition does not end a scan inside a file marked.unordered, and atmcondition never ends a scan early. Rows that used to be lost are returned now.After
tranger2_stop(), opening a topic again takes the master lock again, or turns the tranger into a non-master if another process holds it.Snapshot guards and
treedb_delete_instance()refuse (-1) when they cannot read.treedb_activate_snap()answers -1 when a save fails,__clear__included, and C_NODEdeactivate-snappasses the -1 on.C_NODE:
delete-nodewithoptions.ignore_snapsaskscreateas well asdelete.EV_TREEDB_UPDATE_NODEis no longer a PUBLIC event: a peer could inject it throughC_IEVENT_CLIand write with no permission. A refusedupdate-nodeno longer carries the topic desc.C_AUTHZ:
EV_ADD_USER, thedisabledwrite ofEV_REJECT_USERandEV_IDP_USER_CREATEDrefuse on a replica (-1, logged).C_TREEDB:
apply-schemaanswersdata(named:{treedb_name, applied, saved_schema_version, in_use_schema_version}; unnamed: a list of rows{treedb_name, result, comment, data}), and an unnamed apply writes nothing if any applicable treedb fails. A C literal newer than the file in use takes over a saved-but-not-applied draft in__system__, with a warning.C_TRANGER:
open-rtrefuses an rt_id longer thanNAME_MAX. Open/close failures answer an explicit comment instead ofgobj_log_last_message().
Security and integrity
C_AUTHZ
ac_reject_user: a rejected user whosedisabledwrite fails (replica, failed save) kept its live sessions. The sessions are dropped from the node read before the write, whatever the write answers.rt_disk feeds:
disks/<id>is keyed by id alone, so a client withreadcould open a second feed with the id of the treedb’s own feed on a replica and remove its directory by closing its own feed. That replica then stopped following the topic, and nothing was logged. One id, one feed.treedb on a replica: an update with a
filecolumn wrote a blob into the master’s store and moved links in memory before refusing. The refusal comes first now. The autolink path ofmt_update_nodechecks its save.tranger2_stop()left the fds it closed infd_opened_files, and anytranger2_topic()afterwards cleared__closed__, so the shutdown closed the same fd numbers again, which could belong to someone else by then. The stop marks them -1; only a real reopen revives the tranger.apply-schemawrites the file in use through a temporary file, fsync and rename (it truncated it in place).save_json_to_file()closes the fd when the write fails.
Correctness
C_TREEDB:
draft_changeddiffs the draft against the SAVED file when one newer than the file in use exists: after a Save, the editor said “unsaved changes” until Apply, or for ever on an imposed treedb. An unnamedapply-schemais all-or-none.prune_schemadrops the meta-schema placeholderdefault: {}.save-schemaguards the__system__master. Explicit failure messages.timeranger2: late rows in a marked md2 file are found in both directions, by the iterator index too.
tranger2_iterator_get_page()re-reads an unfiltered iterator’s segments only when the key’s cache changed.treedb_delete_instance()reads its rows with an iterator (a pkey2 value holding/logged two errors per delete).topic_var.jsonis replaced through a temporary file, so the rowid counter is never missing on disk. NULL file ids are refused in the three remaining sites. A replica’s page that cannot read a deleted key’s rows says so. fs_watcher: a missing ROOT stays an error.tr_treedb: a ref in a column its hook no longer fills is stale and is removed with a warning. Such a node could not be deleted, not even with
force.C_TRANGER: a
no_rtlist is freed after its topic closes (it leaked). The keys watch has its own creator (<gobj>^keys), so a client feed named<it>^__keys__can no longer be closed byget-page. It records only the keys the iterator pages. A multi-key iterator counts its unfiltered parts live at every page, so newest-first sees rows appended after the open.C_NODE:
links_refusedis local tomt_update_node, so a nested update cannot reset it.The
__md_tranger__contract of the append is stated exactly: the record is changed in place (a carried dict is dropped before the dump; the dict is attached when a realtime list takes the record). New testtest_append_md_contract.
New tests: test_stop_reopen, test_append_md_contract,
test_rt_disk_multi_feed (feed hijack), tr_treedb_files 20, tr_treedb_hook_rename 5,
a walk of C_TREEDB’s command table refusing a denied user plus replica
refusals, N11 with a real save over a dict-shaped file,
test_command_delete_user 8-9, test_late_record in both directions,
and the N1 traversal test given a real loop and a reachable victim.
gobj-ui 7.25.5, yunos-js gui_treedb 0.17.57 / gui_agent 0.22.78¶
kernel/js/gobj-ui-> 7.25.5:EV_DRAFTSreplaces the host’s draft marks instead of adding to them.A form write no longer hangs after a websocket drop in ANY host:
yui_shell_set_connection_state()publishesEV_CONNECTION_STATE, andC_YUI_TREEDB_TOPICSlistens to it on its shell. wattyzer and yunovatios did not forwardEV_TRANSPORT_STATE.Pending writes are keyed by topic and serial.
The pkey2 of an update is sent as the record had it, not through a
datetime-localwidget, which loses the seconds.The DST limit of
form_time_value.jsis written down.Wiring tests on a document double.
yunos/js-> gui_agent 0.22.78:Apply is counted per TREEDB (
data.applied, or the row result on a 7.25.3 node) and restarts the yuno only if something was applied.Drafts are gathered fresh per saved-schema round.
A 60 s deadline on each two-hop request, and routing errors of
mt_command_parsercount as failures.Toasts stay re-translatable.
yunos/js-> gui_treedb 0.17.57: placeholders and titles of the Rows options, and the “records were gone” text says what happened.Deployed to every consumer; hidraulia (v1) untouched. Every deployed SPA was loaded with a real login, a websocket session and one forced reconnect. Zero errors, after a pre-existing
C_IEVENT_CLItrace-level lookup in yunomusica was fixed.
C_TREEDB: apply-schema asks create-delete, save-schema write¶
BREAKING (authz). 7.25.3 meant to make
apply-schemaaskcreate-delete(it replaces the schema of a treedb whole, likecreate-topic/delete-topic) and put the permission onsave-schemainstead: an operator with onlywritecould Apply but no longer Save, not even adry_run. Nowsave-schemaaskswriteandapply-schemaaskscreate-delete, as 7.25.3’s notes said.test_c_treedb_system_schemachecks both with a user granted everything butcreate-delete.
yunos-js gui_treedb 0.17.56¶
yunos/js-> gui_treedb 0.17.56: discovery re-runs only for a connection whoseC_NODEservices lackmaster(0.17.55 countedC_TRANGERtoo, which never has one, so every session re-ran the whole discovery), and a finished scan no longer logsC_TIMER^scan_timer_N: GObj NOT RUNNING(agobj_stop()afterclear_timeout(), which already stops it). Deployed toartgins.ytreedb.com.
v7.25.3 (2026-09-23)¶
BREAKING in this release (flagged after the fact, 2026-09-23)¶
These behaviour changes shipped in 7.25.3 without the flag:
tranger2_append_record()no longer attaches__md_tranger__to the caller’s record. Read the out-param instead:md2_record_ex_tgainedg_rowid(the key’s global rowid). The dict is attached only to what a realtime list takes.md2_record_ex_tgrew: any object built against the old header must be rebuilt (the md2 format on disk did not change).treedb_update_node()/gobj_update_node()answer NULL on a replica (saved update) and when the save fails; they used to answer the node.C_TRANGER
open-iterator/open-rt/open-liston an id already open answer -1, not 0.C_NODE
activate-snapanswers result 0 on success (7.25.0-7.25.2 answered the activated snap’s id).C_AUTHZ write commands answer “READ-ONLY replica” up front on a replica.
Permissions of the schema commands: the notes below say
apply-schemaaskscreate-deleteandsave-schemakeepswrite. As shipped, the check landed onsave-schemainstead (awrite-only user could Apply and could not Save). Corrected in the next release.
gobj-js 7.25.0¶
kernel/js/gobj-js-> 7.25.0: the version back in line with the SDK, no runtime change since 7.22.2 (a comment inlib_treedb.js, a test ofkwid_new_dict()). Published and tagged so the documented anchors oflib_treedb.jspoint into a tag again; JS API index repinned.
gobj-ui 7.25.4, yunos-js gui_treedb 0.17.55 / gui_agent 0.22.77¶
The JS half of the 2026-09-22 lows:
kernel/js/gobj-ui-> 7.25.4 (a delete whose rows or cards went while the question was open tells the person; a stringtimecolumn is written as ISO text; new consumer i18n key),yunos/js-> gui_treedb 0.17.55 (pre-0.17.53 connections re-scanned, whole-topic Rows card never persisted, placeholder follows the language) and gui_agent 0.22.77 (late apply answers ignored, Save in flight logged, backticks in change ids, Apply tooltip). Deployed to every consumer. JS API index regenerated.
The lows of the 2026-09-22 review, C side¶
timeranger2: a
topic_desc.jsonthat does not load answers NULL at once (it went on withtopic == NULLthrough seventeen spurious errors), and its load carries no exit bit: a failed READ never exits, and the name may come from a peer (the TOCTOU between the existence check and the load ends in an error, not an exit). A topic opened again aftertranger2_stop()clears__closed__, so the shutdown closes what it opened.C_TRANGER:
open-iterator/open-rt/open-liston an id that is open answer -1 “already open” (an open that opens nothing was a success-shaped 0 with no data); the epoch a multi-key iterator is judged by is one counter for the process, not per gobj (two C_TRANGERs over one tranger could hand the same number to two openings); the parts of a multi-key iterator are opened under their own creator (<gobj>^parts), so a client’s one-key iterator named<id>^<key>no longer collides with them (test intest_c_tranger).treedb: a hook column of type
stringis refused at the__system__write (it was blessed there and refused by every link into it, after the unlink of a replace); arequiredcolumn flaggednowis the clock’s and is not asked of the caller; a snapshot guard that cannot READ the key’s records says so (“cannot tell whether a snapshot holds it”) instead of naming a snapshot;update-node’s refusal names its contract (“some of its links were NOT changed ... that column was left as it was”), and the flag it reads survives an update-node re-entered through a treedb event.C_TREEDB:
prune_schema_node()keeps a column’sdefaultas written,[],{}and null included;apply-schemaaskscreate-delete, ascreate-topic/delete-topicdo (save-schemakeepswrite); the replica guards ofsave-schema/apply-schema/saved-schemaread the tranger’s EFFECTIVEmasterwhen the treedb is open (a master that could not take the store in exclusive opened as a replica); an unnamed command over several treedbs names the ones that FAILED in its comment.C_AUTHZ: every write command (
create-user,update-user,enable-user,disable-user,delete-user,set-user-pwd,set-max-sessions) answers “READ-ONLY replica” up front on a replica;set-max-sessionsanswers the refusal of its update, and the session drop of a rejected user logs the refusal of its volatile update.Docs:
shoot-snapafterdeactivate-snapwaits for a reload (treedb.md);apply-schema’s permission (data.md).
gobj-ui 7.25.3, yunos-js gui_agent 0.22.76¶
kernel/js/gobj-ui-> 7.25.3: the schema editor rebuilds its drafts from the host (EV_DRAFTS, N13).yunos/js-> gui_agent 0.22.76: the Schemas tab readsdraft_changedoff everysaved-schemaanswer and sends it to every editor under its tree; both SPAs on gobj-ui ^7.25.3. JS API index regenerated.
timeranger2: a marked md2 file is not read whole at every wake-up of a follower¶
A file the master marked
.unordered(a late record, M16) was read whole byload_cache_cell_from_disk(), and a follower loads the cell again on every append of the master: each append cost the follower every row of that file, 32 bytes a row (N12 of the 2026-09-22 review). The cell in memory already holds the range of the rows it counted, so only the rows after them are read (widen_cell_from_rows()fromrows + 1) and the two ranges are joined.test_late_record: a marked file that keeps growing serves the new rows, the late one and the old ones alike.
C_TREEDB: saved-schema says which topics the draft changes¶
saved-schemaanswersdraft_changed, the topics whose draft in__system__differs from the file in use, computed with the same diff asave-schemauses (draft_changed_from_rows()). The schema editor kept that mark in the memory of one session only, so a reload of the page, a reconnect or its Refresh lost the chip, the banner and the export warning while__system__still differed from the file (N13); gobj-ui 7.25.3 rebuilds them from this answer, which gui_agent 0.22.76 hands it. Red test intest_c_treedb_system_schema.
gobj-ui 7.25.2, yunos-js gui_agent 0.22.75¶
kernel/js/gobj-ui-> 7.25.2: a topic form’s write is answered every way it ends, the transport closing and a write asked with no session included (N8 of the 2026-09-22 review), and a form busy twice comes back whole (N9).yunos/js-> gui_agent 0.22.75: an apply-schema that one owner refuses after another applied goes on to the restart and says who refused (N10); both SPAs take gobj-ui ^7.25.2. JS API index regenerated.
C_TREEDB: the diff of a save reads a schema file whose topics are a dict¶
save-schema(andsaved-schema,diff-schema) diff the draft against the file in use throughdiff_treedb_schema(), which read the file’s topics withkw_get_list(): a file whosetopicsis a DICT keyed by name -- what a node opened withimpose_c_schemaoff wrote before 7.25.0, fromget_treedb_schema()as it was -- read as no topic at all, so every topic was “changed”,nothing to savenever answered, and everytopic_versionand theschema_versionwere bumped and written again at each save (N11 of the 2026-09-22 review). The topics are read as a list whatever shape the file holds them in (schema_topics_as_list()), a dict’s topic taking its name asid; andschema_topic_version()no longer logs “kw must be list or dict” for a topic a dict-shaped file does not hold. Red test intest_c_treedb_system_schema: the same file in use as a list and as a dict gives the same changes.
C_NODE: activate-snap answers result 0; C_AUTHZ: enable-user answers the refusal¶
activate-snappassed the return oftreedb_activate_snap()through as the command’s result: the id of the activated snap since the 7.25.0 side fix that made the library return it (0 before, unless switching snaps).ycommandtakes its exit code fromresult, and a client testing for 0 read a success as a failure. The result is 0 on success (N6 of the 2026-09-22 review; the agent, which tests>= 0, never noticed).enable-userhanded the return ofgobj_update_node()to the response as it was: a refused update (a replica, any refusal of the treedb) answered result 0, “User enabled”, with no record. It answers -1, asdisable-userdoes since 7.25.0 (N7). Red tests intest_c_node_authzandtest_command_delete_user.And underneath, in the library:
treedb_update_node()answered the node whatevertreedb_save_node()said, and moved the node in memory before the save. On a replica the append is refused (“NO master”) and the update “worked” until the next reload -- pattern 1 of the 2026-09-21 review, which the command-level guards of 7.25.0 covered for C_NODE’s commands and not forgobj_update_node()(C_AUTHZ, the agent). A saved update on a replica is refused BEFORE the memory moves, with “Cannot update node, NO master”; a memory-only update (saveFALSE) goes on; and a failed save answers NULL.
C_TRANGER: a dead id is free again, a backward page counts from the live end, a key born again is a deleted key¶
The three registries (iterators, realtime feeds, stateful lists) drop an entry whose handle went with its topic: at
mt_stop, afterdelete-topic, whenget-pageanswers “closed with its topic”, and whenopen-iterator/open-rt/open-listmeet the id again. A stale entry answered “already open” (result 0, no data) to its own id after a close and reopen of the topic, adelete-topic+create-topicor a stop/start of the service, and everyget-pageafter it -1, until aclose-iteratorby hand (N3 of the 2026-09-22 review; the identity resolution of 7.25.0 closed the use-after-free and kept the entries).get_single_key_page()cuts a backward window with the LIVE row count of the key, the one the page’stotal_rowscarries, where it used the count frozen at the open: once the key grew, “newest first” skipped the newest rows and a page past the old count came back empty (N4). A filtered iterator pages its index, which does not grow.A multi-key iterator (
rkey) keeps ONE rt_mem over the topic,only_mdand fed to nobody, for itskey_deletedcallback: a key deleted and created again between two pages was judged intact by its presence in the topic’s cache and paged with the row count of the dead one, result 0 and no log; it answers -1 “was deleted” now, like a one-key iterator (N5).Tests:
test_c_trangeropens the same ids again after a topic reopen, pages a growing key backward, and deletes and re-creates a key under anrkeyiterator (red before).
timeranger2: an unfiltered iterator pages the key as it is¶
tranger2_iterator_get_page()takes an unfiltered iterator’s segments from the cache again before every page. They were taken at the open only, whiletotal_rowswas the live count: the rows appended since the open were in no segment, so a page past the count of the open came back empty and thepagesit announced could not be read. A filtered iterator keeps its index and its segments, as before.
timeranger2: the id of a disk feed stays inside its topic¶
tranger2_open_rt_disk()refuses anidthat is empty,.,..or holds/, with “Invalid rt id (path metacharacters not allowed)”. The id is the directory<topic>/disks/<id>/, and a followerrmrdir()s it before creating it (and again on close): thert_idofopen-rtand thelist_idof a realtimeopen-listreached it from the wire unconfined, under thereadpermission, sort_id=../../../../<dir>removed that directory as the yuno’s user. Same family as the topic-name rule of 7.25.0, which confined every OTHER name from the wire. A backtick is accepted: an rt id is not a kw path segment, and treedb names its own feeds with them. Red test intest_topic_path_traversal.
C_NODE: update-node with create_only asks for create¶
The
createpermission check ofcmd_update_nodewas keyed onoptions.createalone, whilemt_update_nodeturnsoptions.create_onlyintocreate: a caller withupdateand nocreatecreated nodes with{"create_only": 1}(the GUI sends both, so its path was checked; a rawycommandwas not). Both options ask forcreatenow when the node does not exist. Red test intest_c_node_authz.
timeranger2: the append returns its metadata in the out-param, g_rowid included¶
md2_record_ex_thas a new field,g_rowid: the key’s GLOBAL rowid, files included, the number__md_tranger__stores asg_rowid. Every place that fills the struct fills it (the append,get_md_by_rowid(), the disk feed), so a callback or a queue that copies the struct carries it. It was the one thing the out-param did not return, and the only reason a caller had to read__md_tranger__off the record after the append: treedb did, in two places, and now readsmd_record.g_rowid.tranger2_append_record()no longer adds__md_tranger__to the record for the caller. It adds it only to the record it hands to a realtime list that takes it (one that wants the key and is not anonly_mdfeed), built once, right before the first such callback. It used to decide by the record’srefcount, which says who HOLDS the record, not who reads the dict: the queues,c_trangerandwebstatshold it for other reasons and paid for a dict they never read (the queues then kept it in memory per queued message untilmsg_iev_clean_metadata()stripped it on send). Callers hand over their only reference and read the out-param. Pays off for adb_tracksingestingraw_trackswith no realtime list open: +13.5% appends/s intest_topic_pkey_integer(+21% together with the change below).A
__md_tranger__the record carries in is dropped before the content is written. Nothing stripped it, so a record appended again with the dict inside (webstatsstores the report it keeps each run; anadd-recordofc_trangerwith a record taken from aget-page) wrote the metadata of the EARLIER record into the.jsonas content: eight lying integers per record, invisible on read because the load overwrites the key, and permanent on disk.tranger2_read_record_content()refuses a__t__it cannot name a file for, and a body it cannot read, instead of attaching the metadata to a NULL record; and theg_rowidit attaches is the struct’s, where it used to put the file-relativerowidthere.C_TRANGER’s realtime-only feed (open-rt) stops rebuilding a__md_tranger__over the one the append attaches; the consumer gets the same eight fieldsget-pageemits.
timeranger2: an append names its file once¶
tranger2_append_record()computed the record’s file id (gmtime+strftimeof thefilename_mask) three times: for the.json, for the.md2and for the cache cell. It is computed once and handed to the three;get_topic_wr_fd()takes the file id instead of__t__. The data, the metadata and the cache cell now name the same file by construction (the cell used to recompute it from the masked time). +4.6% appends/s intest_topic_pkey_integer. The write of a user flag andtranger2_delete_instance()name the file once too: they computed it for the existence check and again for the descriptor.mark_file_unordered()logs an error instead of silently truncating a<file_id>.unorderedmarker name that does not fit, and still flags the cell as unordered in memory: a marker that cannot be written is not a file that is ordered, and without the flag every later late record logged the same error again.The three writers that name a record’s file (the append, the write of a user flag, the payload wipe of
tranger2_delete_instance()) refuse whenget_file_id()fails (gmtime()rejected the__t__), instead of going on with an empty file id: the record landed inkeys/<key>/.json+.md2, hidden from every load, with a cache cell whose id was"", and the append answered 0.
timeranger2: an append reuses the integers of the cache it updates¶
tranger2_append_record()updates the key’s cache on every append (the file’s cell and the key’s totals: 10 integers). They were replaced with a newjson_integer()each time -- 10 allocations and 10 frees per append. Now the integer already there is set in place (set_cache_int()); nothing holds a reference to them (readers copy the value or deep-copy the cell). Every writer of those fields goes through it, the range widening of a cell and the key totals included, so the invariant has one implementation. +4.8% appends/s intest_topic_pkey_integer(179.8k -> 188.3k without a realtime list, 158.3k -> 165.8k with one).
webstats: each top client says whether fail2ban banned it¶
webstats reads
fail2ban_log_path(/var/log/fail2ban.log) and its last rotation (.1, or the newest dated one on RHEL) and puts on each row of Top clients and Top offenders"banned": {at, jails, ban_number, ban_time}orfalse.Restore BanandUnbanare not counted;Increase Bangives the ban number and length. Recordversion3 (bannedon the rows, afail2banblock with the bans by jail).falseonly when the log was read: an unreadable log leaves the rows empty and says so in Needs attention. Needs attention also warns when none of the top offenders with 3+ probes was banned -- a jail watching nothing.
tools/fail2ban: escalating bans for the probe jail, and a readable fail2ban.log¶
install-probe-ban-escalation.sh:bantime.incrementonyuneta-nginx-probe(1d, 2d, 4d, capped at 1w) plusdbpurgeage = 30d, so fail2ban remembers the earlier bans. Chosen over the stockrecidivejail, which reads the weekly-rotatedfail2ban.logand misses a scanner that returns after a rotation -- as both returns measured on a node did.fail2ban-client -tbefore and after, the previous state restored on failure. Opt-in:dbpurgeageis server-wide policy.make-fail2ban-log-readable.sh:root:adm 0640andyunetainadm, as Debian ships it; RHEL ships the log0600.
webstats: the top clients are named -- country and organisation, from RDAP¶
The first
whois_rows(10) rows of Top clients and Top offenders now carry who the address is: country, the organisation that holds the network, the network name and its range, looked up over RDAP (the JSON successor of whois, over HTTPS). One service answers for every address: RIPE redirects an address it does not hold to its registry, and the lookup follows it (verified against the five registries and IPv6).New state
ST_LOOKING_UPbetween reading and reporting. One lookup at a time, each by aC_PROT_HTTP_CLbuilt for it and destroyed once itsC_TCPsays it stopped; awhois_timeoutbounds each one. A lookup never stops the mail: a failure leaves{"error": ...}on the row, a WARNING in the log andlookup failed: ...in the mail.The stored days are the cache: an answer younger than
whois_cache_days(30) is taken from them, not asked again. Failures are never cached.Private, loopback and link-local addresses are never asked. Everything in an RDAP answer is HTML-escaped before it reaches the mail.
Record
version2: addswhoison the rows and awhoissummary (cached/looked_up/failed), printed under Sources.New attributes
whois_enabled,rdap_url,whois_rows,whois_cache_days,whois_timeout; new trace levelwhois. The node must reach the registries on port 443.
v7.25.2 (2026-09-22)¶
gobj-c: gbuf2json_from_peer() -- a frame from a peer that is not json is a warning¶
New
gbuf2json_from_peer(gobj, gbuf, peer_gobj): parses json RECEIVED from a peer and, when it is not json, logs one WARNING (MSGSET_PROTOCOL, “frame is not json”, thepeernameof the transport underpeer_gobj, the parser’s error, the length) with a dump capped at 256 bytes and no stack -- the decoder-severity rule.gbuf2json(gbuf, 2), which callers used for that, logs an ERROR with a stack and the whole buffer, and names no peer: an internet scanner on a yunovatios gate left exactly that.Switched:
c_qiogate(the ack of the remote gate -- which also went on with a NULL ack and logged more errors; it returns now),logcenter(a truncated UDP datagram is a warning, not an error with a stack) anddba_postgres.c_nodekeepsgbuf2json(): it reads its own store.Test:
gbuffer/test_gbuffer_guards(json back; bad bytes give one warning and no error, read from a log handler of the test).
C_TREEDB: saved-schema on a treedb with nothing saved no longer logs an error¶
The console asks
saved-schemaof every treedb it shows, and most have no saved schema: the missing file reachedkw_get_int()as NULL and logged “kw must be list or dict” (path: schema_version) on each call -- seen in the agent of wattyzer after 7.25.1.apply-schema,save-schemaandtreedbsread the same files the same way; all five sites go throughschema_version_of(), 0 for a file that is not there. Test:c_treedb_system_schemaaskssaved-schemaandtreedbsbefore any save (red against 7.25.1 on that log).
v7.25.1 (2026-09-22)¶
C_TREEDB: impose_c_schema is configuration, not a persisted attribute (BREAKING)¶
The attribute loses
SDF_PERSIST. The value a yuno runs with is its configuration: itsmain.c('global': {'treedbs.impose_c_schema': false}) or its config file. A value saved by the command outranked the configuration, survived every new binary, and lived nowhere a deploy could see. A file saved by an older release is ignored (the loader reads onlySDF_PERSISTattributes).set-impose-c-schemais removed, and the permissionimpose-c-schemawith it. Kept as an in-memory switch it could never act: the value is read when a treedb opens, a treedb opens when its yuno starts (close-treedbrefuses while the yuno plays), and a restart reads the configuration again. To change the value, change themain.cor the config file and deploy.A yuno that relied on a persisted
set=0opens imposing again (the default is1) until its configuration saysfalse.Per treedb:
dynamic_schema_treedbs(list, configuration, not persistent) names the treedbs that open from their schema file whateverimpose_c_schemasays, which stays the default of every other one:'global': {'treedbs.dynamic_schema_treedbs': ['treedb_wattyzer']}. The yuno’s code still wins.New command
treedbsin C_TREEDB: the treedbs the service opened (treedb_system_schemafirst), withimpose_c_schemaas it applies to each,decided_by(code,dynamic_schema_treedbs,impose_c_schema,system), and the literal, in-use and saved schema versions. Test 9b ofc_treedb_system_schemacovers both.
A treedb_schema_<db>.c is what the schema editor exports: the graph, then the literal¶
treedb_schema_authzs.c,treedb_schema_mqtt_broker.c,treedb_schema_controlcenter.c,treedb_schema_yuneta_agent.c(and the docs-onlytreedb_schema_mqtt_subscriptions.c) hold the graph of the schema as a comment and the literal, nothing else: exactly what gobj-ui 7.25.1’sschema_to_c()exports, so an edit made in the GUI goes back into the source by replacing the file whole. The graph is derived (schema_to_diagram()), and replaces the hand-drawn ones; the notes those carried outside the schema are gone, except controlcenter’s plannedlists/viewer_enginestopics, kept inc_controlcenter.cbeside their drafts. The broker’s alarm msg2db schema moves to its ownmodules/c/mqtt/src/msg2db_schema_alarms.c. No schema changed.New
scripts/schema_diagram.mjs <file>...: rewrites the comment from the literal with the editor’s own code, after checking that the exporter reproduces that literal;--checkexits 1 when a comment is stale.treedb_system_schema.c, the meta-schema, keeps its form.
v7.25.0 (2026-09-21)¶
C_TRANGER: a multi-key iterator holds no iterator per key (M22)¶
The memory half of M22 of the 2026-09-21 review.
open-iterator rkey=...kept one full tranger2 iterator per matching key for the life of the iterator -- keys x files of memory, +147 MB for one whole-topic card over 1000 keys with 60 daily files, and gui_treedb reopened that card on every visit. It counts each key’s rows at open, keeps a number per key and the match conditions, andget-pageopens only the keys its page touches, and closes them. Same answers, sametotal_rows, frozen at open as before. A filtered page rebuilds the index of the keys it reads.A topic closed and opened again still makes such an iterator “gone” (the A4 rule): the topic carries an epoch in memory, and the iterator is valid only on the opening it was made on. A deleted key is found by asking the topic, since no part iterator is there to be marked.
gui_treedb 0.17.54: the whole-topic Rows card is no longer remembered and reopened on every visit; its options offer
keys (regex), filled with.*; a new Rows card starts atfrom rowid = -100.Test:
c_tranger(no tranger2 iterator held after the open, nor after a page; red against the previous library).
treedb: the fkey mark is derived, and no file carries it -- a hook can be renamed¶
M2 and M3 of the 2026-09-21 review.
parse_hooks()marks each child fkey column with the one hook that fills it ("fkey": {parent_topic: hook}), and the loader keeps only the links it names. The mark was written to disk:parse_schema()marks the literal in place, and that literal is what the child’stopic_cols.jsonand the treedb schema file are written from. Renaming a hook raises only the PARENT’stopic_version, so the child reloaded the old mark -- “Only can be one fkey” at every open, and the links made through the new hook dropped at every restart.parse_hooks()clears every mark before it computes them, so a stale one read from a file of an older release is ignored;treedb_open_db()writes the schema file and the topics’ cols without it. In memory (and indescs) the mark is as it was.M3: a ref to a hook that no longer exists is stale. What a renamed or removed hook leaves in every child (
departments^d1^users) hangs from nothing, and unlinking it failed on the missing hook (since 3fea635f3), so the child could be neither relinked, cleaned nor force-deleted.unlink_child_from_parent_ref()now removes such a ref from the child with a warning (“Parent ref names a hook that no longer exists”) and goes on. The children lose that parent: link them again through the new hook.Test:
tr_treedb_hook_rename(red against the previous library on the error, the lost link, both files, and the clean and forced delete of a child that names the old hook).
C_TREEDB: an edit of a schema is a draft; save-schema publishes it, apply-schema puts it in use (BREAKING)¶
M36 of the 2026-09-21 review, the owner’s design. With gobj-ui 7.23.196 (the schema editor stops raising versions and marks drafts) and gui_agent 0.22.74 (Save, the imposed banner, and an Apply dialog with the relaunched yuno and the changes).
The three commands take no
treedb_nametoo: then they act on every treedb opened there with a schema from C and list the answers (apply-schemathen applies only whatcan_apply), which is what a console holding a whole yuno sends.An edit of
__system__moves no version. Every write to acolsortopicsnode used to raisetopic_versionandschema_version(the “a write publishes itself” of M8), so an edit half made was already the schema of the next start. The writes are drafts now;publish_schema_change()is gone, and so is the__schema_publishing__marker the projector set to keep it from answering its own writes.save-schema treedb_name=X [dry_run=1]publishes the draft: it compares it with the schema file IN USE (diff_treedb_schema(), the comparator ofdiff-schema), raises thetopic_versionof each topic that differs and theschema_versionto the one in use + 1 (never lowering a number of__system__), writes them into__system__, and writes the schema tosaved_schemas/X.treedb_schema.jsonunder the__system__tranger -- never over the file in use. The saved schema reads like a literal: empty attributes,_geometryand the projection’s bookkeeping versions are pruned, andtopicsis a list. Idempotent.saved-schemaanswers what was saved, aflat_diffagainst the file in use,impose_c_schemafor that treedb (the code’s force included) andcan_apply.apply-schemacopies it over the file in use: master only, refused when C imposes the treedb’s schema, and only a higherschema_version. It takes effect at the next open.BREAKING: with
impose_c_schemaoff a treedb opens from its schema FILE, not from__system__. The literal is handed totreedb_open_db()withoutimpose, so it is installed only when it is newer; the file wins on ties and when it is ahead, which is whatapply-schemamakes it.__system__is not read at open in either mode. Every in-tree yuno and every project’sdb_history*forcesimpose_c_schema=1from its code, so none of them changes behaviour: for them Save works and Apply is refused.Docs:
YUNO_TREEDB.md(“A treedb never opens from__system__”, the three-step cycle with examples), the C_TREEDB page (the attribute and the three commands), and the C_NODE permission table, which still listedsystem-schemaandtraceas open to anyone (M41 closed that).Test:
c_treedb_system_schema-- edits and creates/deletes are drafts, the whole save/saved/apply cycle (red against the previous library), and three checks rewritten on the new model; its expected log is a table now.
treedb: a now column is stamped by every write, writable or not¶
M5 of the 2026-09-21 review, the owner’s option B.
7.24.0 stamped a
nowcolumn on an update only if it waswritable, and its CHANGELOG said there were twonowcolumns in the tree. There are about twenty more -- the “Update Time” of wattyzer, estadodelaire, hidraulia, yunovatios, the mqtt broker and treedb_authzs -- declared['persistent','time','now'], with nowritable: they stayed frozen at the create.nowis stamped by EVERY write now, persistent or volatile;writableplays no part.The instant a thing was born is a
timecolumn withoutnow. A create with no value gives it the clock and an update leaves it alone -- which is what__assets__.t(when the bytes arrived) needs, so it lostnow(__assets__topic_version 2). No new flag was needed: that is the only column in the tree that wants a birth time.A
nowcolumn is an integer epoch; one of another type is not stamped (it used to be written as""by every stamp). None exists in the tree.tr_treedb.h/lib_treedb.js: the comment ofpersistentsays what the treedb does with it -- stored on disk, readable, NOT writable (unlikeSDF_PERSISTof a gobj attribute, which impliesSDF_WR).Test:
tr_treedb_filescase 19 (red against the previous library: a non-writable and a volatilenowcolumn stayed where the create put them).
treedb: a rowid id is not handed out again under an active snap¶
M4 of the 2026-09-21 review, the owner’s option A.
The seed of the rowid counter read the treedb’s id index, which, with a snap active, holds only what the snap loaded. In a store with no
last_rowid_idyet, a create without id got the id of a node created after the shot -- on disk, not in the index -- andexist_primary_node()asks the same index, so nothing refused it. The real case was__graphs__. The seed reads the keys of the topic, tranger2’s cache (every key on disk).treedb_activate_snap()returned the tag of the snap it REPLACED (0 when none was active), not the tag of the one it activated, as its header says. Callers only test< 0, so nothing broke; theresultof C_NODE’sactivate-snapcarries the right tag now.Test:
tr_treedb_rowidcase 7 (red against the previous library: the create got8, the id of the node the snap did not load).
treedb: force unlinks the children, ignore_snaps deletes what a snap holds (BREAKING)¶
M11 and M12 of the 2026-09-21 review, the owner’s option A.
forcemeant two things, and every caller that wanted one got both.treedb_delete_node()read it as “unlink the children” AND as “delete a node a snapshot still holds”. The agent’sdelete-yunoand gobj-ui’s topic table force every delete for the first, so no snapshot guard ever fired for them: a release a snap froze was deleted withoutforce, and the table broke rollbacks with an ordinary delete. They are two options now:forceunlinks the children,ignore_snapsdeletes what a snap holds.treedb_delete_instance()takesignore_snapstoo (forcewas only the snapshot override there, and does nothing now).BREAKING: a
delete-node force=1(ortreedb_delete_node()withforce) of a node a snap holds is refused now, “cannot delete node, a snapshot still holds it”. Passignore_snapsas well, or delete the snap first.The agent keeps its contract. Its
force=1ondelete-yuno,delete-binary,delete-config,delete-realmanddelete-public-servicestill means “even if a snap holds it”, and it passesignore_snapsfor that.delete-yunodrops its own guard on the tag in memory, which is 0 for anything saved after the shot; the treedb’s guard reads the records. Sodelete-yuno yuno_release=<frozen>withoutforceis refused now, as it always said it would be.gobj-ui’s table keeps sending
force(for the children) and now shows the snapshot refusal.Tests:
tr_treedb_snap_cloneandtr_treedb_delete_instancepin both halves (red against the previous library).
The loose ends of the 2026-09-21 review: M15, M26, M31, M41, M42¶
M15 -- commands that answered success for what they did not do.
The agent’s
update-binaryandupdate-configanswered0when the record was not updated (gobj_update_node()returned NULL), andsync-binariescounted the binary installed;delete-realmanswered0whatevergobj_delete_node()said. They answer-1and why now.C_AUTHZ
create-user/update-userwith a role that cannot be linked wrote the user and answered “User created/updated”. A role is checked first -- a refroles^<role id>^usersto a role that exists -- and the user is not written otherwise; it keeps the roles it had.treedb: a DICT hook took the newest instance of a child, so a
delete_instanceof it left it hooked and a forced delete of the parent SAVED it back to disk -- the primary after a reload (the agent’s binaries and configurations). A dict hook keeps the child’s primary instance now, as an array hook keeps the one it has.treedb: a child hooked by several instances of one parent was unlinked from ONE; the others kept it while its fkey no longer named them, and then refused every unlink and a forced delete until a reload. An unlink takes the child out of every instance’s hook.
M41 -- C_NODE: three commands answered anyone.
system-schemaand theset-link-eventsdisplay askread,trace(process-wide) asksupdate.test_c_node_authzwalks C_NODE’s command table instead of a hand-written list, so a command added without a permission fails it.M42 -- tests.
test_c_node_authzopens the same treedb as a REPLICA and checks that every write answers READ-ONLY and changes nothing; gobj-js’kwid.test.jspins the 7.21.0 fix it did not (a record withoutidis logged).M26, M31 -- gobj-ui 7.23.195. A writable time column keeps its seconds across a save; record text in the treedb table is a text node, never HTML.
Tests:
command_delete_user(the role cases),tr_treedb_update_instance(a dict hook over a pkey2 child, and an unlink from one instance; red against the previous library, the resurrection included),c_node_authz. The agent’s three answers have no test: there is no test bench for agent commands.
treedb + C_TRANGER: a schema write is checked and published whole; a session takes its lists¶
M6, M7, M8 and M21 of the 2026-09-21 review (TODO.md).
create-topicwrote a topic before checking its columns (M6) and answered “Topic created!” for one with no columns or noidcolumn, which the next open could not load.treedb_create_topic()now validates the columns first, asparse_schema()does at an open, and refuses with “Topic refused: bad columns” and the reason: nothing is written. A failedtranger2_create_topic()is no longer ignored either.A column written to
__system__skipped two rules an open applies (M7): afilecolumn must be a stringfkey, and ahook/fkeycolumn cannot be both nor be of any type but dict, list or string. Stored, such a column lost its whole topic at the next open. They are refused at the write now, with “Column definition refused”.“A schema write publishes itself” was false for two doors (M8): a column created with its link (
autolink/refs) and a DELETE raised no version, so the change was stored and never reached the running treedb. Create, link and delete published like update did -- superseded, in this same release, by M36 below: no write publishes any more,save-schemadoes.A live
open-listbelonged to nobody (M21): no owner was stamped and no reaper walked the lists, so one opened by a session that then died went on collecting every append in memory until the yuno stopped. It is the session’s now, like its iterators: closed with the session’sEV_ON_CLOSE, kept while the session lives.Tests:
tr_treedb_schema_parse(M6),c_treedb_system_schema(M7, M8) andc_tranger(M21), each red against the previous library.
timeranger2: a topic name stays inside its database, and a replica cannot delete a topic¶
A1, A2 and A3 of the 2026-09-21 review (TODO.md).
A topic name from the wire reached the filesystem unchecked — the only test was “not empty”. With a
readpermission on a C_TRANGER, a name such as../other_db/usersread another database whole; withdeleteit removed a topic there; withcreateit planted one. Create, open, delete, backup,tranger2_topic_path()andtranger2_write_topic_var()/_cols()now refuse an empty name,.,.., and any name holding/or`, with “Invalid topic name (path metacharacters not allowed)”. A leading.stays legal, unlike a key: MQTT queues are<client_id>-IN/-OUT, and the broker accepts a client id such as.foo.One
readcommand could stop the yuno. A name that is a directory but not a topic (..,<topic>/keys, any stray directory) went toload_persistent_json()as a critical, and withon_critical_error=2that is anexit(0)the watcher does not relaunch — reachable on the agent, controlcenter and broker throughtranger_system_schemaandtranger_authz.tranger2_open_topic()now answersNULLfor a directory withouttopic_desc.json: “Not a topic: topic_desc.json not found”.delete-topicon a REPLICA removed the master’s topic.tranger2_delete_topic()was the one destructive call with nomasterguard; it andtranger2_backup_topic()have one now.treedb_delete_topic()refuses before it closes the topic in memory, and C_TREEDB’screate-topic,delete-topicanddelete-treedbanswer a replica with the READ-ONLY response C_NODE already gives.New test
tests/c/timeranger2/test_topic_path_traversal.c(red against 7.24.1: the traversal deleted the other database’s topic, and the replica deleted the master’s).
C_TRANGER: deleting a key no longer stops the master through an open iterator¶
A5 of the 2026-09-21 review (TODO.md), its first half.
An iterator open on a key that is then deleted. A filtered iterator pages over its own row index, and
tranger2_delete_key()removes the files that index points into; the nextget-pageopened a file that was gone, a critical, and with C_TRANGER’s defaulton_critical_error=2anexit(0)of the master that the watcher does not relaunch. An unfiltered iterator answered a short page with the oldtotal_rows, result 0 and no log. One gui_treedb session was enough: its whole-topic Rows card stayed open across a delete-key. Every iterator C_TRANGER opens now registers timeranger2’skey_deletedcallback, which only MARKS it (it runs insidetranger2_delete_key()'s walk of those iterators); its nextget-pagecloses it and answers “iterator ‘<id>’ closed, its key ‘<key>’ was deleted: open it again”. That covers every deleter of the tranger, not only this service’sdelete-key. gui_treedb 0.17.52 re-opens its whole-topic Rows card on the delete-key answer.test_c_trangernow runs withon_critical_errorWITHOUT the exit bit: a critical in the middle of it used to be anexit(0), which ctest reads as a pass.
timeranger2: a failed READ never exits the process¶
A5 of the 2026-09-21 review, its second half (the decision: a read that fails has written nothing, so it is no reason to leave).
get_topic_rd_fd()no longer applieson_critical_error. It still logs “Cannot open file to read” as a critical, and the caller answers an error. It was the one read path that could exit, and only on a master: with the default2, anexit(0)nothing relaunches.A read no longer creates a file.
get_md_record_for_wr()— the read beforewrite_user_flag/set_*_flag/delete_instance, and the whole oftranger2_read_user_flag()— went through the write fd, which on a master CREATES a missing md2; then the read failed withon_critical_error. It now answers “Record metadata file not found” first, and its three criticals (lseek, short read,__t__mismatch) do not exit: they return before writing.What still exits, on purpose: the writes, and above all a short write of an md2 row (
tranger2_append_record), because continuing would misalign every later append of that file. Loading a topic’s own files at open also keepson_critical_error.New test
tests/c/timeranger2/test_read_never_exits.c, run withLOG_OPT_EXIT_NEGATIVEso that an exit is a FAIL for ctest (anexit(0)reads as a pass).
C_TRANGER: a handle is judged by its identity, not by its topic’s name¶
A4 and M22 of the 2026-09-21 review.
A topic closed and opened again made its stale handles look alive. C_TRANGER kept the POINTER of every iterator / rt / list it opened for a client and judged it alive by asking whether a topic of that NAME was open.
delete-topic+create-topic, a stop/start of the service, a backup: the name is back, the handles are not, and the nextget-page— or the session closing, which reaps its handles by itself — dereferenced freed memory (SIGSEGV intranger2_close_iterator). Every use now asks the tranger for the handle by(topic, kind, id, creator), aftertranger2_topic_is_open()(the by-id getters OPEN a closed topic). The pointer is kept only for ano_rtlist, which is not the topic’s. The parts of a multi-key iterator keep their id and are resolved the same way.tranger2_get_iterator_by_id()is one hash lookup. The topic keepsiterators_by_id—{creator: {id: iterator}}— beside itsiteratorsarray, maintained by open and close. It walked the array, and every open calls it to refuse a duplicate, so a multi-key iterator of N keys opened in O(N²) — and resolving by identity at each page would have made every page that too. The memory side of M22 (one full iterator per key underrkey=.*) is NOT addressed here: seeTODO.md. The index holds the same iterator objects as the array, soprint-tranger(C_TRANGER) now shows each open iterator twice, underiteratorsand underiterators_by_id; the snapshot tests of timeranger2 ignore the index, whichtest_iterator_indexchecks on its own.Tests: an ABA section in
test_c_tranger(red against the previous library: SIGSEGV), andtests/c/timeranger2/test_iterator_index.cfor the index.
treedb: a link that cannot be made no longer orphans a node, and +New no longer overwrites¶
M1 —
treedb_replace_links()unlinked first and linked after. A new parent that could not be linked (it does not exist, its hook does not link that column, the node itself, a cycle) left the child with NO parent, on disk,EV_TREEDB_NODE_UNLINKEDpublished and noLINKED;update-nodeanswered “Node update!”. Reachable from the gobj-ui form (autolinkalways), and from C_AUTHZupdate-user, which lost every role. A column is now replaced whole or not at all: every new link is checked (link_can_be_made()) before any old one is undone. The record is still saved, andupdate-nodeanswers -1: “node ‘x’ saved, but its links were NOT changed”.M28 —
update-nodegetsoptions.create_only: a NEW node, and one that exists is refused (“Node already exists”);createalone is an upsert. gobj-ui 7.23.194’s +New sends it: a taken id used to overwrite the record and, withautolinkand empty selects, unlink it.Tests:
test_c_node_link_events8b (the only parent replaced by one that does not exist, red: UNLINKED published and the link lost), 8c (the command, red: success) and 8d (create_only, red: overwritten and unlinked).
timeranger2: block 7 of the 2026-09-21 review (the cache cell)¶
M16 — after a late record, time-range queries hid records. A cell’s
[fr_t, to_t]was rebuilt from the FIRST and LAST md2 rows, and a late record (a__t__below what its file holds) is the last row with a lower time: a follower at once, and the master after a reload, dropped the file from any query past that time. The 7.21.0 entry of c46c820a0 said “a reload says the same”; it did not. The master now drops an empty marker beside the md2,<file>.unordered, when a record arrives below the file’sto_t, and a load reads a marked file WHOLE for its range (only that one: a year of daily files read whole would be a gigabyte per key at startup). A follower merges the ranges it knew with the ones it reads instead of replacing them.M17 — every append cost O(md2 files of the key).
find_cache_cell()walked the cells from the first one, with twosnprintfofNAME_MAXeach, to find the one the record went into — which is the last one, or after it, almost always. It looks at the last cell first, and takes its base from the key’s total. Measured with the review’s program, appends to the last file: 1 file 3.8 us, 365 files 3.9 us, 3650 files 3.9 us (they were 3.7 / 32 / 318 us).M18 — a follower with two disk feeds on one key lost records of two files. The watermark was one per (feed, key), and a batch touching two files of the key reseeded the second feed’s mark on the second file before it had read the first — a late record and a current one, or just a rotation while the follower was busy: that feed lost both. One mark per (feed, key, FILE) now, never seeded over another file’s. Verified with the review’s two-process program (master + a SIGSTOPped follower): the second feed got
W@3alone, it getsW C E X Ynow, rotation case included.New test
tests/c/timeranger2/test_late_record.c(red: the follower and the reloaded master did not serve the record inside the range). M18 needs two processes to show, so its check stays out of ctest; the in-process case is in the same test and stays green.
C_TRANGER + gui_treedb 0.17.53: block 4 of the 2026-09-21 review¶
M19 — closing the LAST Live card of a session reaped its paging iterators, and its Rows cards answered “Iterator not found” from then on.
mt_subscription_deletedstill closes the subscriber’s feeds, and its iterators only when nothing else will: a session is watched, and itsEV_ON_CLOSEtakes them when it dies. gui_treedb re-opens, once, a card whose iterator is gone.M20 — “newest first” served the oldest page.
open-iterator backward=1did nothing, and on an UNFILTERED one-key iteratorget-page backward=1kept the window counted from the start and only reversed it (the library’s contract, pinned bytest_topic_pkey_integer_iterator5, and left alone). C_TRANGER now pages a one-key iterator as it already paged a multi-key one: the window is taken from the END, read forward and reversed, whether the key is filtered or not; and aget-pagethat does not say takes the direction given at the open. gui_treedb sends the direction on every page.M35 — gui_treedb’s replica detection did nothing (fixed in gui_treedb 0.17.53 alone:
treedb-infocarries its service in__md_command__, and storing the scanned services keepsmaster).Tests:
test_c_tranger(a get-page of an iterator opened backward, red: the oldest row; a live session paging after its last unsubscribe, red: “Iterator not found”); gui_treedbtreedb_info.test.js,rows_page.test.js.
gobj-ui 7.23.193: block 3 of the 2026-09-21 review¶
The kernel/js/gobj-ui submodule moves to 7.23.193 (its CHANGELOG.md has the
detail): a row delete crosses the confirm dialog by the row’s id and no longer
by its position, which could delete another record (A6); a refused Save keeps
the form open on what was typed (M25, new attr form_waits_for_answer with
EV_WRITE_DONE / EV_WRITE_REFUSED); a card of the graph follows its UPDATED
again (M33); a __graphs__ echo no longer marks other topics’ unsaved layout
as saved (M32); and M24, M27, M29, M30.
Block 2 of the 2026-09-21 review: disable-user, the agent deletes, snaps¶
A7 —
disable-usernever dropped the user’s live sessions, and used a freed node. C_AUTHZ handed the NODE toEV_REJECT_USER, which readsusername(the node keys onid): the lookup failed, the sessions stayed authenticated (disabledis only read at login), and the event freed the node the response then handed out. It passes{username}now, asdelete-useralready did, and a failed update is answered as one.A8 — four agent deletes handed a borrowed list element to
gobj_delete_node(), which owns its kw:delete-public-service,delete-realm,delete-binary,delete-config. The element was freed inside its list and the list’s decref then wrote into freed memory — silent, which is why dailydelete-binarynever crashed. They pass akw_incref()now, andgobj.hsaysownedon both parameters.M14 — every
update-nodecommand WITHOUToptionslogged an ERROR with a stack (“kw must be list or dict”, from asking the absent options forcreate). It is the form the docs use with ycommand.A9 —
shoot-snapwhile a snap is active restored the whole treedb to it. Every key took the clone branch, the clone became the newest record, and the deactivation reloaded the photo over everything written since.treedb_shoot_snap()refuses while a snap is active, and while the treedb is still loaded from one.M10 — a replica could shoot, activate and deactivate a snap: the shot ended in the critical “Cannot save record tag” (an
exit(0)withon_critical_error=2), and deactivate answered success after changing only memory. The library refuses on a replica and C_NODE answers READ-ONLY.M9 — the two snapshot guards of the deletes opened when they could not read the records. They refuse now;
forcestill overrides.Tests:
command_delete_user(disable-user, red: a logged error),c_node_link_events(update-node without options, red: the ERROR),tr_treedb_snap_clone(a shot refused while a snap is active, red: the live payload came back as the photo’s; and a replica, red: the critical). A8 and M9 have no red test: A8 corrupts freed memory without a trace, and M9 needstranger2_open_list()to fail, which a clean test cannot make.
v7.24.1 (2026-09-20)¶
gbmem: the leak audit does not follow what it writes, and says which ref it caught¶
print_track_mem()walkeddl_busy_memto the end while its own logging appended to that same list. A handler that holds one block per line turns that walk into an endless one, each line feeding the next. It now takes the end of the list BEFORE it logs anything and stops there: the report is what was busy at that instant and nothing the report itself allocates. Measured while chasingdb_history_ce, this guard did NOT change that yuno’s count (1207 before and after), so what it reports there was already busy when the walk began — the guard closes the hazard, it does not explain that case, and the comment in the code says so.check_failed_list()logs thereftoo. Filtering by SIZE catches every allocation of that size — 95.648 of them in one startup of a yuno that loads a database — and without the ref there is no way to tell which of them is the one the audit reported. With it, the catch that matters is found by its ref and its stack is the allocation site.A catch WINDOW, and the bytes of what leaked — the two halves that turn a report into a cause, both read from the environment by any yuno:
YUNETA_TRACK_MEM=<ref_min>-<ref_max>[:<size>,...]logs a stack for every allocation inside the window, andYUNETA_TRACK_MEM_DUMP=1prints 64 printable bytes of each leaked block. The window is what makesmemory_check_list[]usable at all: it needs the exact ref or the exact size, and a ref cannot be prepared in advance — it moves a few hundred between two runs of the same yuno. The dump prints BYTES and does not cast the block tojson_t: a block of a given size is not necessarily the jansson struct it looks like, and the blocks that name a leak are the text ones anyway. The report’s header line prints the window it used. Recipe inDEBUGGING.md§11.7.
v7.24.0 (2026-09-20)¶
The treedb GUI round (gobj-ui 7.23.186-7.23.192)¶
A card of the graph says its id AND the instance it is. The label replaced the id with the first secondary key whenever the id column was flagged
rowid/uuid/qualified, so the three utility yunos of an agent read7.23.0-1three times and never said which yuno each one was. It isid · pkey2now, and the flags decide nothing.A Save of the graph writes only the topics that changed. It wrote one
__graphs__record per loaded topic, whatever had moved: one card dragged on a five-topic treedb appended five records, four identical to the ones under them — in an append-only store. The comparison is against what the backend holds, not against the G6 history, because the Save button also lights for changes G6 does not record.A link is drawn in the colour of the two ports it joins (the child topic’s), instead of one neutral grey for the whole graph; the tree relation keeps its width as its mark. The graph zooms out past 20% (floor
0.02). The schema has its own glyph, a draughtsman’s compass, instead of thehexagon-nodesof the data graph it sat beside.The schema can be read as json, and it is the STORED one: the new
schema-filebelow.A column of a topic table has a ceiling (
max_col_width, 420px): thedescriptionofconfigurationsholds a paragraph per row, andfitDataFillgave that column the whole viewport.Peer floors:
maplibre-gl ^6.10.0,tabulator-tables ^6.5.3. No API moved in either.
C_NODE: schema-file, the schema as it is STORED¶
A new command answers the
<treedb>.treedb_schema.jsonthat sits beside the topics, read from disk and whole. It is not whatdescsanswers:descsis the schema the treedb is USING — one desc per topic, cols as a LIST, hooks resolved — while the file keys its cols by name and carries theschema_versionand eachtopic_version, which is the document somebody editing a schema literal compares against. Nor is it the literal the yuno was compiled with: when the store holds a newer version, the file is the one that won.Asked for by the treedb GUI’s schema json button, which until now showed the runtime
descs. An older backend answers “command not found” and the viewer shows that.
treedb: a snapshot freezes the ARRANGEMENT of the treedb too¶
__graphs__is now part of the photo. It holds how the treedb was arranged — one record per topic, written by the graph view — and that is as much what the store looked like as the records are. Until nowtreedb_shoot_snap()skipped every topic whose name starts with__(“Ignore meta-tables”), andtreedb_open_db()opened__graphs__with nouser_flagfilter, so an activated snap gave back the records of the shot drawn with whatever layout was in use at that moment. Both halves changed: the shot tags__graphs__like any other topic (clone included, when an earlier snap already tagged the record), and the open filters it by the activated tag.The other two meta-topics stay out, and for reasons that are not the same.
__snaps__cannot tag itself.__assets__is held by a snap another way —assets_held_by_snaps()walks the links of the records the snap froze — because its blobs are shared by every treedb of the tranger.A snap shot before anything was arranged holds no layout, so activating it leaves
__graphs__empty and the graph comes back to its automatic layout, which is what that photo looked like.test_tr_treedb_snapphase 10 pins it: a layout is saved, the snap is shot, the layout is changed, and the activation reads back the frozen one while the deactivation brings the live one back.Upgrade note: a snap shot before this holds no layout. Its records were tagged when
__graphs__was skipped, so activating it now leaves that index empty and the graph comes back to its automatic layout — the records of the photo are right, the arrangement is simply not in it. Shoot a new snap (the name of the old one is taken until it is deleted) if the arrangement is to be part of it.
treedb: a now column is stamped by every write, not only by the create¶
The clock wrote a
nowcolumn once and never again.normalize_node_field_value()ignores the value it is handed for a column flaggednowand writes the clock, buttreedb_update_node()normalizes only the fields the kw CARRIES -- and no kw ever carries one, since the point of the flag is that nobody writes it. So__graphs__.timefroze at the instant of the first save: a treedb layout saved four times said, four times, that it was saved the first time. The true instant was never lost (the md2tof each record has it), but the column said something else than its own flag.The gate is
writable, and it is what tells the two kinds ofnowapart.__graphs__.timeis writable: it says when the layout was last saved, and an update stamps it.__assets__.tis not: it says when the BYTES arrived, and a rename of the asset -- which IS an update of the asset node -- must leave it where it is. They are the only twonowcolumns in the tree.test_tr_treedb_filescase 19 pins both halves.
Docs: the behaviour of snapshots and of links, written down¶
YUNO_TREEDB.md§3.9 is now the whole snapshot model: whatshoot-snapwrites (one tag per key, in place, on the live record; a clone when an earlier snap already tagged it), what an activated snap reads (the primary index is filtered by the tag, the pkey2 indexes are NOT, and that pair is what lets a node go back and forward between versions), what happens to a write made while a snap is active (tag 0: the photo does not change, and the write joins the primary index after the deactivation), what a snap protects from (both delete guards, neither reading the tag in memory), and what the AGENT adds on top (the promotion re-append ofpromote_highest_release_yunos(), which is why ayunoskey gets one more record per upgrade cycle). §2.8 keeps the timeranger2 half (theuser_flagis the tag) and points there.§3.7 (and §4.2) say the link rule as it now is: a link writes the CHILD, and only when the child’s fkey moved; a link that only fills the parent’s hook, or that was already made, writes nothing and publishes nothing.
The API pages of
treedb_link_nodes(),treedb_delete_instance()andtreedb_save_node()carry the same rules.
treedb: an instance a snapshot froze is not deleted either¶
treedb_delete_instance()asked the tag the node carried in MEMORY. A save is untagged, so an instance updated after a snap was shot carries tag 0 while the record the snap froze is still under it -- and the delete tombstones every md2 row of that (id, pkey2 value), the frozen one included. It now walks the records of the key, keeps the ones of this instance (the pkey2 value is a FIELD, so the walk reads the content and not only the metadata) and refuses when one of them carries the tag of a snap that exists: “cannot delete instance, a snapshot still holds it”.forceoverrides, as it did. It is the twin of the guardtreedb_delete_node()has.test_tr_treedb_delete_instancepins it. This closes point 2 of “tr_treedb snaps” inTODO.md.
treedb: a link that was already there saves nothing and publishes nothing¶
An idempotent link used to append a record and announce a change.
_link_nodes()warns when the parent ref is already in the child’s fkey (“Parent ref already in child fkey, skipping duplicate”) or the child is already in the parent’s hook, and then went on to fire the treedb callback (EV_TREEDB_NODE_LINKED/EV_TREEDB_NODE_UPDATED) and save the child: a record identical to the one under it, and subscribers told of an update that was not one. It now reports whether anything moved, andtreedb_link_nodes()returns without saving or publishing when nothing did. And the two sides are not worth the same: the parent’s hook lives in MEMORY, the child’s fkey is what reaches the disk, so the save follows the CHILD alone. A link that only fills a hook -- everycreate-yunoof a second instance of a yuno: the instance inherits the fkey of the one before it (the ref names the id both share) and its own hook is empty -- appended a record tobinariesand one toconfigurationsevery time. Pinned bytest_tr_treedb_link_events(the same link twice) and bytest_tr_treedb_update_instance(a new instance of the parent linking a child that already names it; its schema grows apartschild topic).
print-role: the CLI and the runtime command answer the same fields¶
--print-roledroppedyuneta_versionsilently. Itsjson_packformat declared six pairs and was handed seven, so jansson stopped at the sixth and the framework version never reached the output.The runtime
print-rolecommand (C_YUNO) did not answerdate, the build datetime the CLI has always printed; it is read from theappDateattr. Both now answer role, name, alias, version, date, description, yuneta_version, tags, required_services, public_services and service_descriptor, in that order.
The build date of a yuno says its zone: ISO 8601 UTC (CLI 0.19.3)¶
datein--print-role,--versionand the agent’sbinariestopic is now2026-09-18T16:13:49Z. It was__DATE__ " " __TIME__(Sep 18 2026 16:13:49): the LOCAL time of the machine that compiled it, with no way to tell which zone.tools/cmake/project.cmakecompiles every yuno underTZ=UTC(CMAKE_C_COMPILER_LAUNCHER env TZ=UTC, so the 132main.cstay as they are), andentry_point.cpublishes the datetime as ISO 8601 with itsZ. It needs ayunetas initto take effect. Binaries already installed keep the old string until they are rebuilt. The--versionbuffer is sized for name, version and datetime (it wasNAME_MAX, and the compiler now proves the overflow possible).yunetas sync-binaries(CLI 0.19.3) reads both forms in its fallback date compare.
treedb: a write made while a snap is activated is untagged (BREAKING for who relied on it)¶
A record takes a snap’s tag exactly once, from
shoot-snap.treedb_save_node()andtreedb_create_node()took the tag of the ACTIVATED snap (08ba69dcb, 7.23.0), so what was written during a rollback went into the photo being looked at. On wattyzer, with snap “18-sep” active,install-binary auth_bff 7.23.0appended a record tagged 1, and the snap then held two records of the key. Every write is now tag 0. With a snap active, the primary index is still the snap’s records, and the pkey2 indexes hold every other instance. A write made meanwhile reaches the primary index after the snap is deactivated. Meta-topics (__snaps__,__graphs__,__assets__) are no longer tagged while a snap is active either.test_tr_treedb_snap_clonepins it: without the fix the update and the create under snap A are tagged 1, and A no longer shows its shot content.
The Developer window is readable (gobj-ui 7.23.184-7.23.185)¶
Its stylesheet was fixed pixels between 9 and 13, controls included. Now it uses rem: controls at 1rem with a finger’s padding, traffic and log text at 0.9375rem, only secondary text under 0.9rem. 7.23.185 indents the expanded JSON four characters per level (
4ch, it was 16px, about two). Consumers moved to^7.23.185: gui_treedb 0.17.51, gui_agent 0.22.73, yunovatios gui-central/gui-controlador and wattyzer.
treedb: a collapsed view with metadata is no longer a “pure node”¶
node_collapsed_view()withwith_metadatamarked the view with the old key__pure_node__: false. The 2024 rename topure_nodemissed it, because the key sat alone on its own line. The deep-copied metadata kept the node’spure_node: true, so a view (hooks collapsed, a copy) passed everypure_nodeguard, andtreedb_save_node()appended it as a record. The view now sayspure_node: falseand the guards refuse it. Visible in anynodesanswer readwith_metadata(gobj-ui’s per-table Raw JSON showed both keys).test_tr_treedb_immutablepins it: without the fix it fails on “treedb_save_node() took the view”.The same for
gobj_node_tree()withwith_metadata(C_NODE’smt_node_tree). It returnsjson_deep_copy(node), a copy of the whole subtree, and every__md_treedb__in it still saidpure_node: true(13 of 13 in the__system__treedb tree). The copy is now markedpure_node: falseat every level. No caller in the tree asks for it with metadata yet, so the bug was latent;test_c_treedb_system_schemapins it.get-nodeandnodesgo throughnode_collapsed_view(), so the fix above already covers them. So doesexport_treedb: an export madewith_metadatanow writespure_node: false.
Each treedb topic table has its own Raw JSON (gobj-ui 7.23.183)¶
A button between Columns and Export shows the table’s records as JSON, each with its
__md_treedb__metadata. The table asks its host (EV_REQUEST_JSON, a new output event of the hosted child, declared by its one hostC_YUI_TREEDB_TOPICS), which readsnodeswithwith_metadata. Consumers moved to^7.23.183: gui_treedb 0.17.49, gui_agent 0.22.71, yunovatios gui-central/gui-controlador and wattyzer.
Toolbars: the common buttons keep one order (gobj-ui 7.23.182, gui_treedb 0.17.48, gui_agent 0.22.70)¶
Buttons that belong to one view go first, then the common block Refresh, Columns, Export, then Close at the right. gobj-ui’s treedb topic toolbar is now Search, Schema · Refresh, Columns, Export (Schema used to sit between Refresh and Columns). The tranger cards of gui_treedb are Options, Share · Refresh, Columns, Export · Close (Rows) and Pause, Clear, Share · Columns, Export · Close (Live).
v7.23.0 (2026-09-19)¶
upgrade-yunos shoots no snap for nothing (CLI 0.19.2), tranger cards show every column (gui_treedb 0.17.47)¶
yunetas upgrade-yunospreviewsfind-new-yunosBEFORE the rollback snap. With nothing new it stops without a snap. Before, it shotpre-upgrade-<date>first and then answered “Nothing to do”. That snap tagged every current record and cloned the ones an earlier snap had tagged.deploying-yunos.mddescribes the new order.gui_treedb 0.17.46-0.17.47: a record’s
uflagnames the snap that tagged it (read from__snaps__) andsflagnames its bits. The card always shows every column; the Record/Metadata/All selector is gone.
Rows of every key of a topic (C_TRANGER open-iterator rkey, gui_treedb 0.17.43)¶
open-iteratortakesrkeyin place ofkey. It opens one iterator on every key that the PCRE2 regex matches, and lays them end to end in key order, the same ordertr2listprints a topic in.get-pagepages over that concatenation (backwardcounts from its end), and each record names its key in__md_tranger__.key. The match conditions apply to each key. A multi-key iterator is registered, reaped and closed like a one-key iterator, and it is dropped the same way when its topic closes under it.list-keysandopen-iteratornow share one key matcher.gui_treedb: a “Rows topic” button beside “Live topic” opens that view for the whole topic, with a
keycolumn and header sort over the loaded page. The Keys picker’s page size also offers All.
gobj-ui 7.23.181: a JSON viewer given its document at create expands¶
C_YUI_JSONcreated withjson_datastarts inST_READY. The attr filled and drew the tree, but the FSM stayed inST_EMPTY, where expand, collapse, search and copy are not declared -- the first click on a>answered “Event NOT DEFINED in state”. Found on the treedb GUI’s record viewer.gui_treedb 0.17.42: a tranger record opens in a real window on desktop -- a
C_YUI_WINDOWthat moves, resizes, maximizes and remembers its geometry, like the raw-tranger viewer -- instead of a fixed dialog. Mobile keeps the adaptive sheet.
The Developer window: TRAFFIC and TRACES are two feeds (gobj-ui 7.23.176-7.23.180)¶
Found using the window on the deployed treedb GUI, with Traffic ticked alone
to read what was going to the backend -- and getting a list of event names
beside a browser console showing the four payloads.
7.23.176 -- the two feeds stop sharing one selector, which steered the wrong one: the four view modes rewrote the TRAFFIC (
Name onlyleft a message as its event name and nothing else) while the TRACE lines ignored them altogether.VIEWis the traffic’s now and says only how much room its payload takes (Collapsed/Expanded); the payload is always there. What shapes a trace moved to the TRACES row besideSimple mach: a newPayloadchip for thejsonlines a trace dumps, which used to vanish as a side effect of the traffic view being set to names. Both feeds can be on at once, so neither control may borrow the other’s. The console mirror of the traffic obeys the same filter and the same view as the window -- it was called BEFORE the filter and never read the view, so it printed lines the window had just hidden. That is the promise 7.23.33 made for the framework LOGS, which the traffic half had never kept.7.23.177 -- a payload no longer MOVES when the pointer passes over it (the nested indent was a
:hoverrule), and a folded object says its first FIELDS instead of how many it has:{header: "id", fillspace: 18, …}where it said{5}, which for a schema of twelve columns was twelve identical{5}.7.23.178 -- no tooltips over the log. Four
titleattributes popped a box over what was being read, three of them repeating what the screen already said.7.23.179 -- the message’s source is dim TEXT in the entry’s header, the one thing the removed tooltip said that is written nowhere else; it matters most in an app browsing several backends. Plus the first test of
yui_dev.js, whose import also guards a trap the file carries: its stylesheet is a template literal, so ONE backtick in a CSS comment stops the module from parsing.7.23.180 -- the log is painted with ink, not with opacity. Every role was one grey dimmed by a different amount, and opacity blends text TOWARDS the background, so the more a line mattered the less of it was left: the preview measured 4.39:1 and the source 3.78:1, both under the 4.5 floor. Eight tokens now, one per role, measured against the entry’s own background in both schemes; a string has a colour of its own for the first time, and a folded object’s preview is tokenized so it reads with the same ink as the row it previews.
Consumers: yunos/js gui_treedb 0.17.37-0.17.41 / gui_agent 0.22.64-0.22.68,
with the five new i18n keys (collapsed, payload, traffic payload folded,
traffic payload laid out, show the payload of the traces). Deployed to
artgins.ytreedb.com, artgins
impose_c_schema now projects the schema into __system__ too¶
__system__ is the only place a schema can be ASKED for -- from ytreedb,
from gui_agent, from any node command. A treedb that only ever opened with
impose_c_schema on (the default) had no projection at all, so the schema it
runs could be read from its binary and nowhere else.
Opening with impose still does not READ __system__ -- the treedb opens
from the schema in C -- but the MASTER now writes it in the two cases where
nothing of anybody’s is lost:
Projection in __system__ | What happens |
|---|---|
| none for this treedb | seeded from the schema in C |
schema_version lower | re-made from the schema in C |
schema_version equal or higher | left as it is |
The third row is the one that matters: a write to __system__ publishes
itself by raising the version, so a dynamic edit is never overwritten by the
projection, whatever impose does to the schema file. It stays readable with
diff-schema and comes back by turning the flag off, exactly as before.
Only the master writes __system__, which the projector did not check
before: a replica reads the treedb from disk as it is at that moment and
reconciles nothing, the migration of legacy ids included. It used to attempt
the write on an ordinary open too, where a non-master treedb keeps it in
memory, never reaches disk and logs nothing.
Inside a projection that IS being re-made, impose applies at topic level as
it does on disk: a topic is written because it DIFFERS, not because its
topic_version is higher. Under the ordinary rule the projection would
describe a topic the store no longer holds -- the one case impose exists to
repair. An identical topic is not re-appended.
The log line of that open says “system not read” where it used to say
“system ignored”. treedb_open_db()'s options parameter also
documents "impose" in its prototype now, not only in the block above it.
Tests: c_treedb_system_schema test 13 pins the three cases -- seeded,
re-made, left alone -- and test 14 the master rule: it takes the master’s
C_TREEDB down, opens the same store again as a replica, and checks that it
READS the projection the master left and moves nothing when a literal far
ahead is imposed on top. Tests 9 and 10 already pinned that an edit in
__system__ survives an imposed open.
treedb: a snapshot freezes what it shot (BREAKING for who relied on the latest snap following)¶
The design question the 2026-09-15 review left. A save inherited the snap
tag the node carried in memory: after shoot-snap S every later update was
written INTO S, so activate-snap S answered the updated content and only
the creations made after the shot were reverted. The latest snap never froze;
a snap only did when the next one was shot. The agent’s rollback -- shoot
before a deploy, activate if it goes wrong -- reverted the new rows and kept
every change to a row that already existed, and nothing said so.
current_snap_tag() was there for this and nobody called it. Now
treedb_save_node() and treedb_create_node() tag a record with the snap
that is ACTIVATED, 0 when none is: in normal operation every write is tagged
0, so a snap holds exactly what was live when it was shot; an edit made inside
an activated snap stays inside it. treedb_shoot_snap() is unchanged.
Two things follow from the tag no longer riding on updates:
The delete guard asks the records. A delete erases the whole key, and the guard refused a node whose tag in memory was one -- which after an update it no longer is. It now asks whether any record of the key carries the tag of a snap that exists (the tag in memory answers first, then a metadata-only walk of the key), and refuses with “cannot delete node, a snapshot still holds it”.
forceoverrides, as before.The asset gc holds what a snap froze for as long as the snap exists. An asset a node named when a snap was shot stays held after the node moves on, until that snap’s row is deleted -- which is what a snapshot means. It used to be released by the move, because the move itself was written into the snap.
Tests: tr_treedb_snap_clone (an update after the clone freezes neither
snap; a node a snap holds is not deleted, updated or not) and
tr_treedb_files test 13, rewritten to the new semantics. Against the
previous library both fail on the frozen-snap assertions.
A sweep of the minors the post-implementation audit left¶
Nothing here changes a happy path.
delete-treedbrefuses the system schema by name (“‘<name>’ is the system schema, it cannot be deleted”), asclose-treedb/create-topic/delete-topicalready did. It fell through to a “not found” error, since the system schema is not projected intreedbs. Its answers no longer come fromgobj_log_last_message(): a failure says “not projected in__system__, or one of its nodes refused the delete (see the log)”, anddelete_client_treedb_schema()logs the missing projection itself.mt_treedbsincrefs its kw withkw_incref(), the pair of theKW_DECREFit releases with. Test 12 ofc_treedb_system_schemacovers the refusal.A follower no longer logs a stack trace for a directory that vanished before it could be watched. The master signals a deleted key with a directory that appears and vanishes (7.21.0), and on a follower in another process the
rmdircan land between theis_directory()test and theinotify_add_watch();ENOENTthere is now a warning without stack trace (“Directory gone before it could be watched”), the parent’sIN_DELETEfollows.helpers.c’s file lister assembles itsstatpaths withbuild_path()in both branches;testing.ccarries the ArtGins copyright its 7.22.0 helper earned; two comments intr_treedb.cstop describing the hook+fkey column the parser refuses since 7.21.0.yunos/js: its CHANGELOG files gui_agent 0.22.62 / gui_treedb 0.17.36 as released, which they are.yuno_agent: the SPAsgui_agentandgui_treedbcan open a session on the agent.ac_on_openaccepts a fixed list of client roles (ycommand,ycli, ...) and looks up any other role among the agent’s OWN yunos, so an SPA presenting itsyuno_rolewas not found there. The list addsgui_agentandgui_treedband dropsyuneta_gui, a role no client in the repo presents any more.yunos/js-> gui_agent 0.22.63: “For TreeDB” copies the node’s agent as a connection too. Its config has no__top_url__, so the port is read from itswss://gate (agent_secure_port, 1993), the host falls back to the node’s name, and the service is itsC_AGENTone (agent). The agent’s certificate is self-signed, so a browser reaches it only once that certificate is trusted.
Traces: a saved scope replaces main()'s defaults, and the global no-trace is a command (C and JS)¶
A yuno’s main() sets trace defaults before it creates the yuno -- above all
gobj_set_global_no_trace("timer_periodic", TRUE) -- and C_YUNO restores what
the user persisted with the trace commands. The restore only ADDED levels, so a
default the user turned off came back at every restart, and the global no-trace
could not be turned off at all: there was no command for it (DEBUGGING.md said
so, “by design”). And save_global_trace() deleted the __global_trace__ key
when its last level went, so “none” was not something a user could keep.
A saved scope REPLACES what is in force (
set_user_gclass_traces()/set_user_gclass_no_traces()): the global trace scope, the global no-trace scope and each gclass scope are cleared before their saved levels are set. A scope never saved keepsmain()'s default.A scope is saved WHOLE, from the levels in force, an empty one as
[](save_global_trace(), the newsave_global_no_trace(), andsave_user_trace()/save_user_no_trace()for a gclass). A gobj-name key (reset-all-traces gobj=) keeps its level-by-level list.new commands
set-global-no-trace/get-global-no-trace, persisted inno_trace_levels.__global_no_trace__.new API
gobj_get_global_trace_no_level()andgobj_get_gclass_trace_level2()(a gclass’s own levels, without the global ones).--global-traceis applied again right after the yuno is created, so what the command line asks wins over a persisted global scope.The same in
c_esp_yuno.c(not built here: no ESP-IDF toolchain on this machine).
Migration note: a trace_levels / no_trace_levels saved by an older release
holds level-by-level lists, and those now replace main()'s defaults for their
scope too. A gclass no-trace list that misses a default of main() loses it;
fix it with set-gclass-no-trace, or clear the attr with
remove-persistent-attrs.
The JS side moved in the same round: gobj-js 7.22.0 gives the JS C_YUNO the C
yuno’s trace_levels / no_trace_levels and its trace commands with this same
rule, removes the yuno attrs (tracing, trace_timer, trace_inter_event,
trace_creation, trace_start_stop, trace_subscriptions, trace_i18n,
no_poll) the old dev panel wrote, and decides every trace by its bit
(C_IEVENT_CLI’s traffic is its own level ievents). gobj-ui 7.23.172’s
Developer window (7.23.173 fixes a gclass NOT FOUND its Traffic chip logged in an
app with no websocket) sends those commands and never analyses messages: its
“Periodic” filter -- which counted signatures and hid a whole burst of commands
as “recurring” -- is gone, and Periodic is now the timer_periodic level, i.e.
EV_TIMEOUT_PERIODIC. Mute and No poll are gone. gobj-ui v1 1.0.4 (npm
legacy) ports its dev panel to the commands; the mains of hidraulia,
estadodelaire and yunomusica stop passing the removed attrs, and the
setup_locale() of six apps stops reading trace_i18n from the yuno.
Found using the deployed window, same round: gobj-js 7.22.1 fixes
current_timestamp(), which wrote UTC time followed by the local offset (two
hours wrong at +0200); gobj-ui 7.23.174 keeps in the window what arrived before
it was opened, leaves payloads out of Name only / Compact, and gives FIND a
clear button. Then gobj-js 7.22.2 makes the console and the window show the same lines:
a kw is dumped only with ev_kw (a publication and a subscription printed it
under machine alone), and a log sink installed late is handed the last 600
lines written before it; gobj-ui 7.23.175 stamps each row with the time the
line was written.
gobj-ui 7.23.171: the treedb views say the keys of a topic¶
kernel/js/gobj-ui -> 7.23.171, yunos/js and every consumer on
^7.23.171, deployed by the round to the eight hosts. A topic with pkey2s
keeps several instances under one id, and tkey says where the time of a
record comes from. None of the three treedb views said either one:
C_YUI_TREEDB_SCHEMAdrew a pkey2 field with*, like any other required field, although the.cliterals mark it(2). It now draws(2)in bold like the pkey, and(t)on the tkey field.The topic-info panel had no
pkey2srow and hid an emptytkey. It now shows both (append time when there is no tkey),systemby flag name, andpkey/pkey2/tkeyin the key cell of the column table.The topic cards have a
pkey2sline.
(t) is new to the notation, so the legends of the three literals that carry
one (treedb_schema_authzs.c, treedb_schema_yuneta_agent.c,
treedb_schema_controlcenter.c) gained the line. It is a comment only. JS API
doc links repinned to 7.23.171.
gobj-ui 7.23.170: the form’s Save sends the pkey2 back¶
kernel/js/gobj-ui -> 7.23.170, yunos/js and every consumer on
^7.23.170, deployed by the round to the eight hosts. Found auditing 7.23.168
(A8 of the 2026-09-15 review): the rule for what a form writes back --
writable, fkey, file, or the pkey -- left out the SECONDARY keys.
yunos.yuno_release is persistent, required and not writable, so the
update-node went out with no pkey2 and C_NODE resolved it to the PRIMARY
instance: right by chance while the topic table lists primaries, wrong from a
form opened on a row of instances, or on any topic whose pkey2 column is not
writable (configurations.version, public_services). A pkey2 names the
instance the update is for, and it goes back for the same reason the pkey
does. The rule lives in treedb_write_plan.js now, pure and tested with
pkey2s as a list and as the bare string of a C literal. JS API doc links
repinned to 7.23.170.
C_TRANGER: a paging session is watched once¶
Found auditing 7.22.0’s “a session that only PAGES no longer leaks its
iterators”. watch_owner() subscribed C_TRANGER to the session’s
EV_ON_CLOSE on EVERY open-iterator / open-rt, on the strength of a
comment that said gobj_subscribe_event() returns the subscription already
there. It does not: given the same (event, filter, subscriber) again it logs
“subscription(s) REPEATED, will be deleted and override” with a stack trace,
deletes it and creates it anew. So every handle of a session after its first
-- a gui_treedb tab with two Rows cards -- was a warning and a stack trace in
the yuno’s log. The reaping itself was right. It now asks
gobj_find_subscriptions() first and subscribes only when nothing is there.
Test: c_tranger gains a paging session -- a real C_IEVENT_SRV, created and
never started -- that opens two iterators (watched once, no log), and then
closes: both iterators are reaped, the watch goes, and the session’s parent
hears the close once. Against the previous library the second open logs the
warning. That test is also the one 7.22.0 shipped without.
treedb: an unlink from a parent the child does not hang from is refused¶
Found auditing 7.21.0’s “a link into a single-valued fkey moves the child”.
That fix made _unlink_nodes() clear a string reference only when it names
the parent being unlinked, and log otherwise -- and then it went on: it
published EV_TREEDB_NODE_UNLINKED for a link that did not exist and saved
the child, a record identical to the previous one. Reachable from the wire
with C_NODE’s unlink-nodes. The array and dict shapes of an fkey had the
same fall-through since the first version, and the dict HOOK side removed an
absent key in silence, so a mismatched unlink there reached the event with
no log at all.
The check is now made once, before anything is touched, on the child’s
reference -- the half of a link that is persisted -- and it decides: a child
that does not name that parent is not unlinked from it. The call answers -1
with “Cannot unlink, the child does not hang from that parent” (in place of
the three “Parent ref not found in … child data”), the hook and the child
are left as they were, no event is published and nothing is saved. A link
that IS there behaves as before.
Test: tr_treedb_relink, whose mismatch case now counts the UNLINKED events
and the child’s g_rowid. Against the previous library it fails on all
three: the call was not refused, one event fired, one record was appended.
treedb: the rowid counter survives a topic_version change¶
Found auditing 7.21.0’s fix (“a rowid id is never handed out twice”). The
counter it introduced, last_rowid_id, lives in the topic’s topic_var.json,
and tranger2_create_topic() REMOVES that file when the schema’s
topic_version goes up (or down while a treedb imposes its schema), so a key
the new schema no longer carries goes away with it. The counter went with it
too. At the next create get_next_rowid_id() seeded itself again from the
ids still alive, and if the highest id had been deleted it was handed out
again -- the one thing the fix forbade, because a snap’s id rides the records
it tagged. The agent’s yunos, configurations and public_services are
rowid topics whose topic_version moves between releases.
It showed only across a RESTART: in the same process the topic stays open in
the tranger and keeps the counter in memory, so a close and reopen of the
treedb alone could not see it, and the existing test did not. The counter is
now read before the file is removed and written back after the file is
re-created; every other key still follows the schema.
Test: tr_treedb_rowid gains a case that deletes the highest id, shuts the
tranger down, starts it again with the version raised, and creates. Against
the previous library it hands out 4, the deleted id, for 7.
v7.22.0 (2026-09-16)¶
JS: gobj-js 7.21.0, gobj-ui 7.23.169, and the yunos that ride them¶
gobj-js 7.21.0 closes the kwid_* review: kwid_find_one_record() crashed
with a TypeError when there was no data (it read .length off the null
kwid_collect() answers for a kw that is neither a list nor a dict, and one
caller hands it the data of a command answer, which is missing exactly when
the command found nothing); kwid_new_dict() dropped a record without id in
silence where the C twin logs it; kwid_new_list() is ported from C with its
semantics. 19 new cases in tests/kwid.test.js, which nothing covered before.
The version jumped 7.16.6 → 7.21.0 and skipped four: the first two indices are the SDK’s, and this package had drifted again.
gobj-ui 7.23.169 closes the nine gobj-ui findings of the 2026-09-15 treedb
review — every action crosses the FSM now (the Op-column pencil, the confirm
dialogs, the kws that carried G6 event objects, the writes run from DOM
callbacks), the table’s search no longer matches the COUNT of a hook,
ac_unselect_rows() reads its own attr, the toolbar and the search box carry
title/aria-label, the form dialog’s title re-translates, and the first
graph.render() guards against its own view being gone.
yunos/js: gui_treedb 0.17.36 — every deferral is a posted event, the scan
watchdog is a C_TIMER child, the connections table’s row actions are reachable
from the keyboard, and a REPLICA opens without its write buttons (the discovery
asks treedb-info per C_NODE service and stores it, because the library reads
readonly once, when it draws a topic’s toolbar).
tools: audit-agents.sh says when an agent runs a binary that is gone¶
Every node runs yuneta_agent plus yuneta_agent22 as a deliberate redundancy,
and a spare left behind on an old binary is invisible until the day it is
needed — five days on four nodes, running the version-comparison bug 7.12.0
fixed. install.sh restarts the spare after a package upgrade, so the runtime
nodes are covered; a node that BUILDS FROM SOURCE has no install.sh.
tools/agent/audit-agents.sh is read-only and exits 2 if an agent is not
running, 1 if one is stale, 0 otherwise, so it drops into a cron unchanged. The
test is one line and it is EXACT, not a heuristic: Linux refuses to write into a
binary that is being executed (ETXTBSY), so cp over a running agent fails
outright and install / mv / a package succeed only by unlinking first —
which leaves the running process holding an inode with no name, and the kernel
marks its /proc/<pid>/exe (deleted). Every replacement that can happen
while the agent runs is one the kernel marks, so there is no in-place case to
miss.
C_TRANGER: a session that only PAGES no longer leaks its iterators¶
mt_subscription_deleted() reaps the realtime feeds and iterators a SUBSCRIBER
leaked, which covers every client with a Live card. A client that only PAGES --
open-iterator + get-page, no open-rt -- subscribes to nothing, so nothing
ever told the service its session had died: its iterators lived until
mt_stop, one per Rows card per dead session. gui_treedb browsing without a
Live card did exactly that, and a filtered iterator costs more than an
empty handle -- it holds its row index, one rowid per matching record, so a
card over a wide time range pinned a proportional array.
The session is watched directly now. C_IEVENT_SRV publishes EV_ON_CLOSE
when it goes, so watch_owner() subscribes C_TRANGER to it the first time that
session opens a handle -- a subscription of OURS to IT, which asks nothing of
the client and needs no new command -- and ac_on_close reaps by the same
src_gobj stamp the subscriber path already used. The two reaping loops are
one function now, reap_handles_of(), called from both.
Guarded by the gclass NAME of src (C_IEVENT_SRV), because src is not
always a session: a local caller is some gobj of this yuno that does not
outlive us anyway, and the one thing that must NOT be watched is
C_IEVENT_CLI, whose mt_subscription_added forwards every explicit
subscription to the remote peer. A gobj_has_output_event() test cannot tell
the two apart -- both publish EV_ON_CLOSE. (This paragraph said
gobj_has_output_event() when 7.22.0 was cut; the code never did.)
C_TREEDB: delete-treedb refuses a treedb that is OPEN¶
The command deletes the SCHEMA of a treedb -- its projection in __system__
-- and it did not ask whether the treedb was running. It deleted it under one,
without a word: the treedb went on answering from the copy it holds in memory,
its schema no longer existed anywhere, and the damage landed at the NEXT
open-treedb -- the C_TRANGER service of the old one is still alive under its
name, so the create collides and the open dies with an internal “tranger
client NULL” that names nothing an operator can act on. The store on disk is
orphaned by then: data with no schema to read it by.
It refuses now, naming the way out (close-treedb, or the yuno’s own
lifecycle with pause-yuno + play-yuno) -- the guard its sibling close-treedb
already had, for the same reason. force does not lift it: force already
means “yes, delete the schema” here.
The // TODO falla, hay que revisar that sat on the delete itself is gone, and
so is the diagnosis TODO.md carried, which was wrong on both counts: the
parent IS deletable before its children (force unlinks them itself), and the
collapsed view a tree hands mt_delete_node is only read for its id, which
then re-resolves the pure node. What that path does with a CLOSED treedb was
right all along.
Test: c_treedb_system_schema gains test 12 -- a treedb of two topics and
three columns, deleted twice. Closed, nothing of the projection may remain;
open, the command must refuse and change nothing. Red against the previous
library on both halves.
mqtt: a protocol warning says WHO sent the packet¶
137 of the 142 decoder warnings of c_prot_mqtt2.c named a malformed or
hostile packet without naming its sender, which is not actionable on a broker
with a thousand sessions. They carry peername now, read through a peer_of()
helper that answers "" when the bottom gobj is already gone -- a late error
is logged after the transport is torn down.
The 71 gobj_log_error of that file were left alone on purpose. The scope is
the one CLAUDE.md sets for decoder severity -- “could a remote peer trigger
this with bad bytes?” -- and an internal invariant is not the peer’s doing: a
field naming a peer for a fault that is ours reads as an accusation.
c_prot_http_cl.c was read as part of the same sweep and has nothing to
migrate: its seven logs are config errors or internal ones, and the only one
that looks like a decoder is building our own request. It carries the url
now, which is what an operator needs of a bad request -- a peername there
would name the server we chose.
testing: a log assertion that does not have to know the order¶
set_expected_results() matches every captured log against the head of
the expected list, so the list states the exact messages, in the exact order,
the exact number of times. That is the right assertion for a test that drives
one sequence, and it stays the default for all 137 tests.
It is the wrong assertion when two independent gobjs log the same thing and
nothing orders them against each other. c_tcp2/test2 is the case: the client
and the accepted server side both log "Connected" / "Disconnected", and the
driver calls set_yuno_must_die() from inside one side’s close callback, which
logs "Exit to die" synchronously and shuts the yuno down. The other side’s
last log is swallowed -- or is not, when the two close completions land in the
same io_uring batch. So the count moved and not only the order, and no
ordered list could be right: the test failed on a busy box naming a message
that was perfectly correct.
New set_expected_results_unordered() reads the same list as a WHITELIST: a
log matches any entry, a match does not consume the entry, every entry must
still be matched at least once, and anything matching no entry fails. It gives
up “in this order, this many times” and keeps “these things happened, and
nothing else did” -- which is what that test’s own comment always claimed to
check. Documented in docs/doc.yuneta.io/test_suite.md, with the bar for
reaching for it.
c_tcp2/test2 uses it; nothing else changed. Verified both ways: a message
that happens and is not whitelisted fails, and a whitelisted message that never
happens fails.
v7.21.0 (2026-09-16)¶
C_TREEDB: mt_treedbs answers the list its contract promises¶
Last finding of the 2026-09-15 review (M-C8). A framework method answers what
its contract says, and gobj_treedbs() says “a list with treedb names”.
C_TREEDB’s mt_treedbs answered a msg_iev_build_response envelope when
the user had no read: a dict, which a caller reads as a treedb named
result. It now answers NULL, and says why in the log (MSGSET_AUTH) --
a NULL with no message would be a silent error. The allowed path already
answered a list; C_NODE’s mt_treedbs always did.
Latent: no in-tree caller reaches C_TREEDB’s mt_treedbs (the only
gobj_treedbs() call is C_NODE’s treedbs command, on a C_NODE, and
C_TREEDB publishes no treedbs command).
Test: c_treedb_system_schema asks gobj_treedbs() for a list and for the
refusal. Against the previous library the refusal case fails.
C_NODE: every command that reads or writes the treedb asks for a permission¶
From the 2026-09-15 review. C_NODE checks a permission inside each command
handler, with or without the global enable_command_authz gate. Only nodes
(read) and the node writes did. Every other read answered anyone the
routing gate let in, and so did four writes. It matters only in a yuno
with an authz checker (C_AUTHZ), which is the case of the agent, the
controlcenter, the logcenter, mqtt_broker and every gate and database of the
projects.
read:node,instances,pkey2s,parents,children,jtree,hooks,links,snaps,snap-content,print-tranger,export-db,treedbs,treedb-info,topics,desc,descs. Someone who can list (nodes) already had it.update:link-nodesandunlink-nodes(a link writes the child’s fkey),activate-snap,deactivate-snap. Link and unlink asked for nothing.create:shoot-snap(a new row of__snaps__).createANDupdate:import-db, which asked for nothing.update-nodewithcreate=1(the upsert the SPAs create with) now asks forcreatetoo, but only when the node does not exist yet. Before, it created underupdatealone. Anupdate-only user still saves existing nodes through it.
help, authzs, system-schema (the compiled meta-schema) and trace (a
framework switch, which belongs to the global gate) ask for nothing, as
before. A refusal answers -403 No permission to '<permission>' in service '<treedb>', as nodes always did.
Test: c_node_authz (new). It uses a checker that grants by user name.
Against the previous library it fails for every command listed above.
timeranger2: an append to an earlier file goes to that file’s place¶
From the 2026-09-15 review (M-T2). A key’s cache has one cell per md2 file,
in the order the load gives them (by file name). A global rowid is a position
in that order. The append path looked only at the LAST cell, so a record whose
__t__ belongs to an earlier file got a second cell of that file at the end.
Triggers: append-record __t__=, tr2migrate, tr2q_mqtt, or a clock stepping
back across the filename_mask.
The master served the wrong record. With A (day 1), B (day 2) and then C (day 1), an iterator served A, B, A. C could not be read, and the first record of the old file was served twice.
A follower counted the old file twice (4 rows for 3), re-published the whole file to its feed, and served A, B, A, C.
Only a reload was right (A, C, B).
Now the record goes to its file’s cell wherever that cell is. A new file’s
cell goes to its place in the order. The global rowid handed out is the
record’s place, the same number a reload gives. For an append to the last file
it is the total, as it always was. One consequence cannot be avoided: a
record for an earlier file moves every record of the later files one place up
(B goes from 2 to 3). A reload always did this. Now memory and disk agree at
once.
Test: test_out_of_order_append (new). Against the previous library it fails
every check, on the master and on the follower.
timeranger2: a follower hears a deleted key once per feed; a feed can close itself from its callback¶
From the 2026-09-15 review (M-T1, M-T3). Reproduced with a master and a
follower that had inotify armed. test_delete_key_propagation never armed it:
it put yev_loop in the config, which tranger2_startup() overwrites with its
parameter.
A delete reaches each rt_disk feed exactly once. A key with records since the feeds opened fired 6 times on each feed that wanted it. There were three causes. The master mirrors the key’s directory into EVERY feed’s directory. Each removal arrived twice, as
IN_DELETE_SELFof the directory and asIN_DELETEof its parent. And each arrival fired every feed of the topic, not the feed whose directory fired. Nowfs_watcherreports a subdirectory’s deletion only through its parent, in a recursive watch. The root still reports its own deletion. The follower fires only the feed that owns the watcher, asupdate_key_by_hard_link()already does for records. The root’s own deletion is no longer read as a key.A key with no records since the feeds opened is heard too. No feed had its directory, so the delete fired nothing, and the follower’s cache kept the dead key: every later read of it failed. The master now creates and removes the key’s directory where it does not exist. A key directory that appears and vanishes means the key was deleted.
A feed may close itself from its own callback.
fs_stop_watcher_event()called from inside the watcher’s callback freed thefs_eventwhileyev_callback()was still walking the batch with it. The stop is now deferred to the end of the walk, and the rest of the batch is dropped.
No runtime code registers key_deleted today. The stale follower cache
affected every follower.
Test: test_delete_key_propagation. Its rt_disk case now arms inotify, and it
has a new case with a real follower. Against the previous library it fails in
every case: 0 fires in process, 6/0/6 and 0/0/0 on the follower, and a
self-closing feed that neither fired nor closed.
treedb: an update cannot change a pkey2 value; the snapshot clone moves the node’s metadata¶
Two findings of the 2026-09-15 treedb review. Both were reproduced against the installed library before the fix, and each now has a test.
treedb_update_node()refuses a kw that changes a pkey2 value (“An update cannot change a pkey2 value, create the instance”). Nothing is touched or saved. A pkey2 value names an INSTANCE. On disk the new value became a second instance beside the old one. In memory both slots went on holding the live node, so it was listed twice with the new content. Worse,treedb_delete_instance()of the OLD value read the value from that node and tombstoned the rows of the NEW one: after a reload the node was back on its old content, and the update was lost. C_NODE’s lookup already refuses a different value, but it skips an EMPTY one, andrequiredonly asks for the key. Soupdate-node id=X version=""reached the primary from the wire. A kw that carries the same value (as C_NODE sends it) is an ordinary update. A new instance is atreedb_create_node(). Test:tr_treedb_update_instance.After
treedb_shoot_snap()clones a record, the node in memory is on the clone. A record already tagged by an earlier snap gets a clone with the new tag. The clone is the newest record, so a reload makes it the primary, but memory stayed on the original with the OLD tag. The next saves inherited that tag: the previous snap followed the updates while the new one stayed frozen, and a restart swapped them. The clone also went to disk without the immutable bit, so an immutable node lost its protection at the next reload. The clone andtreedb_save_node()now share one append, which moves the metadata and stamps the immutable bit again. The clone still publishes noEV_TREEDB_NODE_UPDATED. Test:tr_treedb_snap_clone(new).tr_treedb_fileshad the old behaviour written into it: its gc case expected the EARLIER snap to follow the move and the later one to keep the asset. It now expects the earlier snap to keep what the node held when it was shot, so deleting that snap frees the asset.
treedb: a cycle in one hook is refused, and every cycle is freed at close¶
A hook holds the child NODE, so a cycle of links is a cycle of json
references. It was accepted, it survived a reload, and it leaked at every
treedb_close_db(): measured at 15 676 bytes for two nodes, and nothing
without the cycle. A cycle in ONE hook (a under b and b under a through
departments) also sent children recursive=1 and jtree into endless
recursion. Neither walk had a visited set or a depth limit, so a client that
could link and read could overflow the yuno’s stack.
_link_nodes()refuses a link that would hang a node from its own descendant through the same hook (“Cannot link, the link would close a cycle in the hook”), before anything moves. A tree is a tree. Only a hook that links its own topic can close one. A cycle through TWO hooks (a department managing one it contains) is data, and is still accepted.treedb_close_db()empties every hook of every node, in theidindex and in the pkey2 indexes, before it frees them. So a cycle no longer leaks, whether it came through two hooks or from a store written before this check.children recursive=1andjtreekeep the PATH they walk and do not follow a node that is already above (“Cycle in the hook, node not followed again”). They keep the path and not a set of visited nodes, so a node that hangs from two parents through an array fkey still shows under both.
Test: tr_treedb_relink.
treedb views: the form writes back only what goes back; a JSON drill keeps its path (gobj-ui 7.23.168, gui_treedb 0.17.34)¶
kernel/js/gobj-ui -> 7.23.168, and the same range in the consumers
(npm run deploy-round). From the 2026-09-15 treedb review (A8, A9, A10):
The form’s Save no longer writes the read-only fields back. It published the whole kw, read-only fields included, as an
update-nodewithautolink, and the backend writes any column it is handed. A non-writabletimecolumn, drawn as adatetime-localwith no seconds, moved back up to 59 s on every save of any other field. One rule,col_goes_back_to_treedb(), now filters both writes of the form.A drill of the raw JSON viewer opens its branch.
print-tranger path=<path>sent the path at the top of the kw, andC_IEVENT_CLIhands back only the__md_command__frame, so the subtree replaced the whole document. The two gobj-ui treedb views and gui_treedb’s Tranger view send__md_command__: {path}.with_copy_button/with_paste_buttonset tofalseno longer break edition mode.
C_NODE link-nodes / unlink-nodes name a missing child; msg2db has a test in the suite¶
A missing child is named as a child. When the child did not exist,
link-nodesandunlink-nodesanswered “Parent not found” and dropped the parent node thatgobj_get_node()had handed them without releasing it. They now answer “Child not found: ‘<topic>^<id>’” and release it.tests/c/tr_msg2dbis registered. It was kept out because msg2db “leaked” 8 tracked blocks per open/close. The leak was the test’s own doing. A master tranger watches its/disksdirectory with inotify, andtranger2_shutdown()cancels that watcher asynchronously. The test never gave the loop the turns to free it, so the memory was still allocated at the leak check. With the loop drained after each shutdown, astest_rt_disk_multi_feeddoes, it passes clean, and msg2db has a test in the suite for the first time.
treedb: two hooks on one fkey are refused; C_NODE’s links / hooks / delete_node answer right¶
An fkey answers to one hook, and
parse_schema()now says so.parse_hooks()marks each fkey with the one hook that fills it (field.fkey = {topic: hook}), and the loader keeps only the links that mark names. A second hook on the same fkey replaced the mark in silence, and at load the links of the first were dropped with no log. The check “Only can be one fkey” was there since the first version, but it askedkw_has_word()of a dict, which is true only for atruevalue, so it never fired. It now compares with the mark already there. The same hook met again is not a second one, because a schema is parsed at validation and again at open. No schema in the SDK or in any project has two hooks on one fkey (all 15 were checked).linksandhookswith no topic answer one key per topic. The loop keyed every topic by the EMPTY topic name it was asked for, so the answer was{"": <the last topic's>}.delete_nodeon a topic that does not exist answers -1. It answered 0, and callers such as the agent’sdelete-*commands (if(gobj_delete_node(…)<0)) read that as a successful delete.
Tests: tr_treedb_schema_parse (two hooks refused, one hook parsed twice
clean) and c_node_link_events.
timeranger2: a topic opens whole on a filesystem without d_type¶
On XFS with ftype=0, NFS, FUSE or overlay, readdir() gives no d_type,
and find_keys_in_disk() asks the inode with stat(). It joined the entry to
the topic’s directory instead of its keys/. <topic>/<key> does not exist,
so no key counted, and the topic opened with an empty cache over files that
were all there: reads answered 0 rows, and the first append started a cell
{rows:1} over a file that already held N, so the wrong record was served from
then on. Dead code on ext4 and xfs with ftype=1, which is why nothing saw it.
It now joins keys/ (with build_path()). The #else branches of this
function and of the file lister in helpers.c used variables they never
declared, which compiled only because Linux defines DT_DIR. Test:
tests/c/tr_dt_unknown, which links with -Wl,--wrap=readdir to hide
d_type from a fully static binary.
C_TREEDB acts only on the treedbs it opened; C_TRANGER’s delete-topic asks for force again¶
close-treedb, create-topic and delete-topic found their target with a
global gobj_find_service(treedb_name) and acted on whatever it returned.
close-treedb treedb_name=treedb_system_schema (with the yuno paused, or
force=1) stopped and destroyed C_TREEDB’s own __system__ treedb, whose
handles live in its private data. The next open-treedb, diff-schema or
mt_stop then used freed memory. Any other service name was destroyed the same
way, and create-topic / delete-topic read the tranger attr of whatever
they found. Now all three refuse a target that is not a C_NODE opened by
this C_TREEDB (its parent), and they refuse the __system__ treedb: “‘<name>’
is not a treedb opened by this service”. Every caller in the tree closes on the
service it opened with, so none is affected.
C_TRANGER’s delete-topic deleted a topic WITH records without force.
It counted the records with tranger2_topic_size(topic, NULL), which passes
the topic where the tranger goes and no topic name. That lookup failed (“Cannot
open topic”), the count came back 0, and the guard “topic with records, you
must force to delete” never fired, so any topic was deleted on the first call.
It also read force without KW_WILD_NUMBER, and the command line delivers
force=1 as an integer (“path MUST BE a json boolean”). The second bug was
hidden by the first. Both are fixed: a topic with records is refused without
force, and force=1 from the command line is honoured. delete-key, two
functions below, always did both right.
Tests: c_treedb_system_schema (the refusals, with force) and c_tranger
(delete-topic refused without force, done with an integer force).
treedb: a column is hook or fkey, never both (BREAKING for such schemas)¶
A column flagged ['hook', 'fkey'] was accepted, and its fkey side was never
written to disk. convert_node2tranger() wrote the column with the shape of a
hook, so every link that stored its reference there was gone after a reload,
and nothing said so. The loader had a second defect in the same place: it keyed
the parent’s hook dict by the wrong reference. No schema in the SDK or in any
project uses the pair, and gobj-ui’s schema editor already refused it (“the
treedb writes one half of a link, never both”).
parse_schema() now refuses such a column (“A column cannot be both ‘hook’
and ‘fkey’”). That is the validation every yuno runs before it opens its
treedb, and open-treedb runs it too. treedb_create_topic() refuses the
topic (“Topic refused: a column is both ‘hook’ and ‘fkey’”), because
create-topic is a live command and a parse failure there only logs. A node
that is both a child and a parent carries two columns: the hook, and the fkey.
The test schemas that used the pair now do exactly that (tr_treedb: a new
departments.manager fkey behind the managers hook). The rule is in
CLAUDE.md and YUNO_TREEDB.md. Test: tests/c/tr_treedb_schema_parse.
treedb: a link into a single-valued fkey moves the child¶
A string column with the fkey flag holds one parent, so a new link
REPLACES the old one. _link_nodes() wrote the new reference and left the
child in the hook of the parent it hung from before, a phantom child. A delete
of that parent without force was refused (“has down links”). With
force it “unlinked” the phantom, and _unlink_nodes() emptied the child’s
string without checking which parent it named. That cleared the reference to
the NEW parent and saved it, so after a reload the child hung from nobody.
Now a link first unlinks the child from the parent its string names. That also
publishes EV_TREEDB_NODE_UNLINKED for it, so a re-link emits UNLINKED + LINKED
where it used to emit only LINKED. An unlink only clears a string reference that
names that parent. Otherwise it logs “Parent ref not found in string child
data” and leaves the child alone. Also fixed: _unlink_nodes() went on to a
switch on a NULL value after “field not found in the node”, and
_link_nodes() printed a const char * with %j. C_NODE’s seed-link guard is
unchanged: it refuses to overwrite a seed’s link before any of this runs. Test:
tests/c/tr_treedb_relink.
treedb: a rowid id is never handed out twice¶
A topic whose id column carries rowid hands out the id when the create
sends none. It was tranger2_topic_size() + 1, and that function sums the
RECORDS of every key: deleting a node lowered the total, so the next id landed
on one that already existed, and every update raised it, so the ids skipped.
__snaps__is such a topic. Shoot two snaps, delete the row of the first (the documented way to release the assets a snap holds) and shoot again: the new snap asked for the id of the second, the create was refused as “Node already exists”, andtreedb_shoot_snap()logged a critical.C_TREEDBruns its tranger withexit_on_error2, so the yuno exited.In a topic with
pkey2san existing id with a DIFFERENT secondary key is not refused: it adds an instance. The agent’syunosis such a topic, andcreate-yunosends no id, so a yuno created after adelete-yunocould become an instance of an unrelated yuno. The agent’sconfigurationsandpublic_services, and the controlcenter’sservices, arerowidtopics too.
The id is now one past every id the topic ever handed out, kept as
last_rowid_id in the topic’s topic_var.json and raised past any numeric id
already in the index, which also seeds it in an existing store: nothing to
migrate. An id is never reused, because a snap’s id rides the records it
tagged: a new snap with a deleted snap’s id would inherit them. Ids no longer
skip on updates, so new ids in an existing store continue from the highest one
present. Test: tests/c/tr_treedb_rowid.
treedb: a column that declares no flag no longer crashes the yuno¶
Every user column is validated against the cols topic of
treedb_system_schema, where flag is an enum that is not required. A
column without flag reached check_desc_field() with its value NULL, skipped
the required test, and the enum branch switched on json_typeof(NULL):
SIGSEGV. It was reachable from C_TREEDB’s create-topic (which passes the
cols it receives), from open-treedb with a schema, and from any C schema
literal. An absent value that is not required now has nothing to check, as the
branch for the basic json types already said; an unknown flag is still refused.
Test: tests/c/tr_treedb_schema_parse.
tranger2_str2system_flag() maps each name to its own bit¶
idx_in_list() counts from 0 and the function applied a count from 1, so
every name landed on the bit of the one before it: sf_string_key gave 0,
sf_rowid_key gave sf_string_key, sf_int_key gave sf_rowid_key, and
sf_t_ms / sf_tm_ms gave bits nobody reads. It went unseen because every
caller in the tree passes "sf_string_key" and tranger2_create_topic() falls
back to a string key when the topic has a pkey. The one real path was
C_TRANGER’s create-topic system_flag=sf_int_key, which created a
rowid-key topic, and a sf_t_ms that was silently dropped.
Only NEW topics are affected: an existing topic reads its system_flag from
its own topic_desc.json. A topic created by hand with that command keeps the
flag it was created with. An unknown name is now logged (“Unknown system_flag
name, ignored”) instead of being dropped in silence. Test:
tests/c/timeranger2/test_str2system_flag.
audit-sshd.sh runs clean over stdin¶
ssh node 'sudo bash -s' < tools/sshd/audit-sshd.sh audits a node without
copying anything onto it. In that mode BASH_SOURCE[0] is unset, so
7.20.0-4’s script printed “BASH_SOURCE[0]: unbound variable” and its last
line named the caller’s working directory instead of the scripts’ one. It now
names the fixed install path, /yuneta/development/yunetas/tools/sshd. The
audit itself was never affected.
treedb file columns are seen, not only named (gobj-ui 7.23.159, 7.23.161, 7.23.166, 7.23.167)¶
kernel/js/gobj-ui -> 7.23.159, and the same range in the consumers. The form
and the table showed a shortened sha256 and nothing else.
Form: a preview under the file control -- a picked file at once, from the
Fileitself (ablob:url, nothing read before save), the stored asset once the host has fetched it.Table: a cell that names an asset is a link; a click opens a popup with every asset the cell names, its type and size, and an “open in a new tab” link. Before, an array column showed its joined ids as one id.
Kinds: image, video and audio in place; a PDF in the browser’s own viewer; any other type a card with the open link.
The bytes come from
C_ASSETSwithget-asset: newC_YUI_TREEDB_TOPICSattrassets_service(default"assets"), eventsEV_REQUEST_ASSET(output of the table view) andEV_SET_FILE_PREVIEW(input ofC_YUI_FORM).yui_asset.jsexportsyui_asset_kind,yui_asset_href,yui_asset_file_answer,yui_asset_open_linkandyui_asset_release.7.23.161: a click on a file cell opened the record too, over the photo.
7.23.166: the Content-Security-Policy that
vite-plugin-yuneta-html.jswrites had nomedia-src, so the browser refused a video or an audio. It now carriesmedia-src 'self' data: blob:.7.23.167: a long file name pushed the form wider than its dialog (the row’s flex items measured their minimum with the whole name).
Consumer keys
show file,open in a new tab,asset not available.
A hook opens the rows it links (gobj-ui 7.23.162, 7.23.163)¶
kernel/js/gobj-ui -> 7.23.163, and the same range in the consumers. The
[N] of a hook cell opened a popup of every child id (5,675 for one device
type), with no height, no scroll and no way out but a click inside it. It now
opens the CHILD topic’s table filtered to the rows whose fkey names the row,
with a chip and its clear button; a hook with several child topics asks which
one. New events EV_OPEN_LINKED (output of the table view, declared by
C_YUI_TREEDB_TOPICS), EV_FILTER_BY_PARENT, EV_CLEAR_PARENT_FILTER,
EV_CHOOSE_LINKED. In 7.23.163 the table loads a hook as its count
(hook_size), not as the ids of its children. Consumer keys filtered by,
clear filter, choose a topic, show linked records.
The graphs take the wheel over their cards (gobj-ui 7.23.160, 7.23.164)¶
kernel/js/gobj-ui -> 7.23.164, and the same range in the consumers. The
treedb schema graph (C_YUI_TREEDB_SCHEMA) gets the family’s toolbar camera
cluster and wheel (it scrolls; Ctrl + wheel zooms). G6’s HTML nodes do not pass
the wheel on, so over a card it did nothing; new yui_graph_forward_wheel()
hands it to the canvas, used by the schema graph and, with a card selector, by
the treedb graph (C_G6_NODES_TREE).
Back from a topic returns to the landing it came from (gobj-ui 7.23.165)¶
kernel/js/gobj-ui -> 7.23.165, and the same range in the consumers. After a
click on a node of the schema landing, Back left the schema on screen under
the cards url, and the landing toggle stopped working. Back now navigates to
the landing’s own route.
7.20.0-4¶
A packaging revision, not a new version: since 7.20.0 the tree under
kernel/, modules/, utils/ and yunos/ has changed only in two JavaScript
submodule pointers and one document, none of which the packages carry. What
changed is the packaging and tools/. The packages are rebuilt as
yuneta-agent-7.20.0-4 and attached to the existing 7.20.0 tag.
sshd leaves the packages: tools/sshd/ scripts, run by hand¶
7.20.0-3 shipped /etc/ssh/sshd_config.d/10-yuneta-ssh-flood.conf in both
packages, and the .rpm enabled fail2ban’s sshd jail. That is right on a
node we operate and wrong on a node of another company: how sshd behaves is
the operating-system policy of whoever runs the node, and a package that
installs an agent has no business changing it. Neither package touches sshd
any more. The same measures, and a few more, are scripts under tools/sshd/
(shipped in both packages at /yuneta/development/yunetas/tools/sshd/), run
by an operator who decided to:
audit-sshd.sh [--all]-- read-only. Readssshd -T(what sshd runs with) and reports FAIL / WARN / INFO: password and root login, weak algorithms and host keys, file permissions, the flood settings, drops in the last hour of the journal, fail2ban. Exit status 2 / 1 / 0.install-sshd-flood-guard.sh [--check|--remove]-- the drop-in, now withPerSourceMaxStartups 10too (OpenSSH 8.5+).sshd -tbefore and after, the previous state back if sshd rejects it; reloads, never restarts; readssshd -Tto confirm the values are in effect, and warns when password login is on.install-fail2ban-sshd-jail.sh [--check|--remove]-- thesshdjail withbackend = systemd; does nothing where the distribution already runs one.fail2ban-client -tbefore any reload.stop-sshd.sh [MINUTES|--no-restart]-- stops sshd and arms a transient systemd timer that starts it again (default 60 minutes). Refuses whensshd -tfails, and arms the timer BEFORE stopping: if it cannot be armed, nothing is stopped.start-sshd.sh-- from the provider’s console: starts sshd and cancels the timer.
With sshd stopped on Debian, sshd -t fails for a missing /run/sshd
(systemd removes it with the unit); every script creates it, as Debian’s init
script does, and retries. Verified on wattyzer (OpenSSH 10.0p2): audit 0 FAIL /
0 WARN, the flood guard installed and in effect.
Upgrading from 7.20.0-3: the .rpm erases the drop-in and the jail (an
edited one is kept as .rpmsave); dpkg leaves the drop-in in place as an
obsolete conffile. Nodes where the drop-in was put by hand are not affected.
7.20.0-3¶
A packaging revision, not a new version: since 7.20.0 the tree under
kernel/, modules/, utils/ and yunos/ has changed only in two JavaScript
submodule pointers and one document, none of which the packages carry. The
packages are rebuilt as yuneta-agent-7.20.0-3 and attached to the existing
7.20.0 tag.
Added: sshd stays reachable under a connection flood¶
Both packages install /etc/ssh/sshd_config.d/10-yuneta-ssh-flood.conf with
LoginGraceTime 20 and MaxStartups 50:30:200 (stock: 120 s and 10:30:100).
On 2026-09-13 an Internet-wide password botnet, 2,000-3,000 attempts an hour
per node, filled sshd’s slots for unauthenticated connections, and past 10 of
them sshd dropped new connections at random: up to 1,464 drops an hour on the
yunovatios controller, the operator’s logins and the deploy tools’ rsync
among them. Password login was already off, so the botnet could not get in;
it only took the slots. With these two lines, applied by hand on three nodes
that day, the drops went to zero.
The postinst /
%postvalidates withsshd -tbefore it reloads, and reloads, never restarts. If the check fails because of this file, the file is set aside as.disabledand sshd is not touched; if it fails without it too, the node’s own configuration is broken and is left alone. A reload with a bad configuration is harmless, but the next restart would not come up, and a node without sshd is a node nobody can reach.A conffile (
%config(noreplace)in the.rpm, 0600 like the distribution’s own files there): local edits survive upgrades. sshd keeps the first value it reads, so a node overrides it with a lower number.It relieves the symptom. Port 22 stays open to the world; the fix for that is a source allowlist, and in the end the sealed node (
yunos/c/yuno_agent/NODE_SEALING.md).
Added: fail2ban watches sshd on RHEL/Rocky¶
The .rpm installs /etc/fail2ban/jail.d/yuneta-sshd.conf, [sshd] enabled
with backend = systemd. Debian enables that jail in its own
defaults-debian.conf; RHEL ships none enabled, so yunovatios-central had
fail2ban running and nothing watching sshd. fail2ban-server requires
python3-systemd on EL9. %post runs fail2ban-client -t before reloading
fail2ban, with the same set-aside rule: a jail fail2ban cannot configure
takes the whole server down, every other jail with it. It does not stop a
botnet of thousands of addresses; it stops the one address that hammers.
7.20.0-2¶
A packaging revision, not a new version: the tree under kernel/, modules/,
utils/ and yunos/ is the same one 7.20.0 was cut from. The packages are
rebuilt as yuneta-agent-7.20.0-2 and attached to the existing 7.20.0 tag.
Fixed: the nightly log rotation reopens nginx on RHEL/Rocky too¶
The postrotate of /etc/logrotate.d/yuneta (the same drop-in in the .deb
and the .rpm) guarded its USR1 to the web server’s master with
if kill -0 "$pid" 2>/dev/null. On RHEL/Rocky logrotate runs in the SELinux
domain logrotate_t, which may send the master USR1 but not the null signal:
avc: denied { signull } ... scontext=logrotate_t tcontext=unconfined_service_t tclass=process, and the policy does not audit
it (it shows only with semodule -DB). So the guard failed, the signal was
skipped EVERY night without a word, logrotate.service ended successfully,
and nginx went on writing to the rotated access.log.1 until the next restart
of the web server. On yunovatios-central every rotated file ended at a
deploy, never at midnight.
The guard now asks /proc/<pid>/comm whether the pid is a live nginx
(stricter than kill -0: a recycled pid of another program no longer
passes), and when there is no live master it says so in the logrotate output
instead of skipping in silence. Verified inside logrotate_t, with the
hardening of logrotate.service: the new drop-in reopens the logs. The
drop-in is a conffile; a node that never edited it gets the new one on the
next package.
v7.20.0 (2026-09-12)¶
Treedb: set-link-events, the link events switched at run time¶
with_link_events could only be set when a treedb’s C_NODE was created, so
a client that follows the graph (the treedb graph of gobj-ui, any frontend)
had to live with what the yuno was configured with: the parent’s
EV_TREEDB_NODE_UPDATED on every link and unlink, from which it cannot tell
which child moved, so it re-reads.
C_NODE has a new command, set-link-events, on the treedb service itself --
the same service a frontend already asks for nodes and descs:
ycommand -c 'command-yuno id=<id> service=<treedb> command=set-link-events set=1'set=1: a link/unlink publishesEV_TREEDB_NODE_LINKED/EV_TREEDB_NODE_UNLINKEDwith the relationship (hook_name,parent_topic_name,parent_id,child_topic_name,child_id).set=0: it publishes the parent’sEV_TREEDB_NODE_UPDATED, what the v1 SPAs (estadodelaire, hidraulia) read.No
set: the current value.
It acts at once, on the open treedb. It is either/or for EVERY subscriber of
that treedb, not per client. Setting it needs the update permission. It is
not persistent: on the next start the treedb has the configured
with_link_events again (C_TREEDB’s attribute, copied to each treedb it
opens). The change is logged (“with_link_events changed”, with the user).
Treedb: an autolink update moves only the links that change¶
C_NODE’s update-node with autolink used to run treedb_clean_node()
(unlink EVERY link of the node) and then treedb_autolink() (link again
from the record). Two consequences:
Every link was announced as broken and remade. A link the update did not change still published
EV_TREEDB_NODE_UNLINKEDand thenEV_TREEDB_NODE_LINKED(or, withwith_link_eventsoff, twoEV_TREEDB_NODE_UPDATEDof the parent), and a subscriber saw the node orphaned in between.One bad ref cost every link.
treedb_autolink()stopped at the first ref it could not link (a parent that does not exist), after the clean had already removed all of them. The record was then saved in that state.
The new treedb_replace_links() (tr_treedb.h) compares each fkey column of
the node with the same column of the record: a ref that is no longer named is
unlinked, a new ref is linked, and a ref in both is not touched (no event, no
write). A link that cannot be made or removed is logged and skipped, and the
others go on. update-node still saves the record: a link can be repaired
later, a lost record cannot.
A ref is now refused when its hook does not link the topic into the column
where the ref arrived (“fkey reference: its hook does not link into this
column”). _link_nodes() links by the hook the ref names, so such a ref used
to be linked into ANOTHER column, without an error. A missing parent logs
“fkey reference: parent node not found”.
No change: a fkey column that the record does not carry is still an EMPTY
column for an autolink, and its links are removed (with the same “fkey
empty” warning). treedb_clean_node() and treedb_autolink() stay as they
were. test_c_node_link_events covers the four cases.
Treedb schema versions: published by whoever changes the schema, never invented¶
The model, as it always was meant to be: a schema is changed either from the
C literal (raise the changed topic’s topic_version, and the treedb’s
schema_version when the runtime must use it) or dynamically from an editor
(gui_agent, ytreedb), which raises both on save. The runtime takes a version
only if it is HIGHER than the one on disk. A literal behind a dynamic schema
stays behind: the schema is now changed dynamically, and a new installation
that must carry those changes takes them into the literal.
C_TREEDB’s projector had broken both halves since 2026-08-12 (98aa51bb8,
5889972f7), and it is restored:
Reconciliation compares the literal with the stored
schema_version, not withc_schema_version. A literal that is not higher is not applied; if it is behind, the log says “TreeDB schema from C is behind the schema in use, not applied”. A new literal used to overwrite a dynamic schema.Numbers are the literal’s, as they are. The projector published under
max(stored, literal) + 1, for EVERY topic of every re-projection. Because the treedb opens from the projection, those numbers reached the store’stopic_var.json/topic_cols.json: a store drifted from its literal although nobody had edited anything (the local agent store ended with all five topics and the schema one number ahead).A topic is published by its own version. Inside a projection a topic is written only if it is new or the literal raised its
topic_version, and of its columns only those that are new or changed (diff-schema’s comparison). A topic the literal changed without raising its version is left as it is, with “Topic from C differs from the one in use, but its topic_version is not higher: not applied”.A meta-schema change re-projects nothing: the schema in use may be a dynamic one. Only the structural move of rowid ids to qualified ones still runs, for a store written with an older meta-schema.
test_c_treedb_system_schema follows the model: versions land verbatim, an
untouched topic keeps its version, a literal behind an edit is not applied and
one ahead of it is, and a column changed without raising its topic is not
published (and diff-schema reports it). Stores that drifted under the old
rule stay ahead of their literal until it passes them, or are reinstalled.
Persistent attributes: a failed save is no longer reported as saved¶
dbsimple.c (the Linux persistence behind gobj_save_persistent_attrs()):
db_save_persistent_attrs()anddb_remove_persistent_attrs()returned 0 whatever happened; they now return what the write returns. A failed write was logged (“Cannot save device json database”) but its caller was told it worked, so a command that checks —set-impose-c-schema,remove-persistent-attrs— answered success for a value that was not saved. Every other caller ignores the result, so nothing else changes.load_json()logs a file that exists but cannot be read (“Cannot load device json database”, with the parser’s error and line). It returned NULL in silence, and the next save then rewrote the file with only the attributes being saved, dropping the rest without a word.db_load_persistent_attrs()no longer leakskeyswhen nothing is saved.
The ESP32 persistence (esp_persistent.c) has the same defects and more; it
is recorded in TODO.md, not changed.
impose_c_schema: the code takes the schema back¶
A new attribute of C_TREEDB (SDF_RD|SDF_PERSIST, default 1). With it on,
open-treedb opens every treedb with its schema from C and neither reads nor
writes __system__; against the disk, a stored schema_version or
topic_version lower than the literal’s takes the literal, an equal one is
kept, and a HIGHER one is overwritten with it — a dynamic change being
reverted. __system__ keeps the changes, for diff-schema or to take them
back. Turn it off to let the schema be changed dynamically (gui_agent,
ytreedb); turn it on again to impose the code — somebody lost that permission,
or the system was broken or changed by mistake — and restart the yuno.
set-impose-c-schemashows it (noset) or changes it (set=1|0), under a new permissionimpose-c-schema, and logs “impose_c_schema changed” with the user. It acts at the next open of the treedbs. Because the value persists, the deploy config (treedbs.impose_c_schema) only gives its first value.The yuno’s code can force it:
open-treedbtakes a parameterimpose_c_schema=1that imposes the schema from C whatever the attribute says. A persistentset=0survives restarts and new binaries, so without it a binary could not impose its schema again; with it, imposing the law is deploying a binary that forces it. The log says “impose_c_schema forced by the code of the yuno, over the attribute”, andset-impose-c-schemalists the forced treedbs inforced_by_code. Every treedb of the SDK is forced: the agent,controlcenterandmqtt_brokerpass it toopen-treedb, andC_AUTHZgives it to theC_NODEoftreedb_authzs, which does not go throughopen-treedb. So do thedb_history*yunos of the projects.treedb_open_db()option"impose"(with"persistent", master only) makes the schema passed win over a newer one on disk, at the treedb and at each topic; timeranger2 rewritestopic_cols.json/topic_var.jsonfor a topic whose stored version is higher while the treedb is imposing.C_NODEgets the matchingimpose_c_schemaattribute, set byC_TREEDB.It restores
use_internal_schema, removed on 2026-08-12 (dd9e3c003), under a name that says what it does — and it now does it: the old option opened with the literal, but a newer persisted schema file still won, so it could not revert anything. The project yunos that still passuse_internal_schemaare not affected (TODO.md).Default
1is a behaviour change on upgrade: every treedb opens from its literal again, as before 2026-08-12. A node where the schema was changed dynamically must turn the flag off to keep those changes in use; a store that drifted ahead of its literal under the old+1rule is brought back to the literal’s numbers.
test_c_treedb_system_schema gets a realm of its own (inside the directory it
wipes) so a persistent attribute can be saved, opens __system__ with the flag
off, and checks imposing: the disk comes back to the literal’s schema and
columns while __system__ keeps the edits. treedb_schema_fidelity opens with
the flag off too.
The agent’s schema marks realms as its main topic¶
treedb_schema_yuneta_agent.c: realms carries 'main_topic': true, published
the way any change to a topic is: its topic_version goes 7 -> 8 and the
treedb’s schema_version 23 -> 24. Nothing changes in what a viewer draws:
realms is the only topic of the agent’s treedb hooked to itself, so the graph
already deduced it. The mark is there as the reference example of the key.
YUNO_TREEDB.md §3.2 / §3.11 and the treedb_open_db() notes show it.
The treedb graph’s find box only looks (gobj-ui 7.23.158)¶
kernel/js/gobj-ui -> 7.23.158, and the same range in the consumers. The find
box searched the whole treedb and unfolded a page of the matches it found, so
every keystroke re-laid the graph out and moved the camera, and the groups it
opened stayed open when the box was cleared. Now it lights the matching cards
ON SCREEN and changes nothing else; emptying the box leaves the graph as it
was. Opening up to a topic stays with the legend’s focus button. The term stays
live across unfolds and refreshes, Enter / Shift+Enter centre the
next / previous lit card (k/N matches), and the count adds what it did not
light (+N not shown, +M in hidden topics). Consumer keys not shown,
find on screen.
The treedb graph: a layout stepper, one icon shape per meaning, elbow edges (gobj-ui 7.23.157)¶
kernel/js/gobj-ui -> 7.23.157, and the same range in the consumers:
A layout change holds the view and repaints the minimap (7.23.145). A new layout moved every card and the camera stayed on the old coordinates -- an empty grid -- while the minimap, which G6 repaints only on draw events that a layout never emits, went on showing the arrangement before. The change now runs through the fold reconcile (the node the reader was looking at stays on its pixel) and the minimap is repainted after every layout.
A layout stepper (7.23.146): up and down beside the layout select, the layout before or after the current one without opening the list. Consumer keys
previous layout,next layout.One icon shape per meaning (7.23.147): a chevron opens or closes one thing or scrolls, a shafted arrow steps a sequence, a tree box steps one level of the tree, a double chevron acts on all of it. The treedb graph’s toolbar had four
chevron-downs with four meanings. The rule is in the JS GUI conventions ofCLAUDE.md.Elbow edges, as an option (7.23.149-7.23.152): a toggle draws every edge as mxGraph drew a tree -- straight out of the port, along the channel between the rows, straight in; the children of one hook share one bus. A reciprocal pair is drawn apart, a line that would cross a card is routed round it by an orthogonal search of our own (
treedb_elbow.js), and the ports count as obstacles. Consumer keyelbow edges.Tried and reverted in 7.23.157, after looking at them on a real treedb: a
compact treelayout (contour packing, then stacked leaves), staggered elbow turns, curves that went round the cards, and a stacked radial. The graph draws as it did at 7.23.152; thecompact-treekey is no longer used.gobj-ui’s
deploy-roundreports a consumer whose build FAILED as failed (it read backOK: a build refused inprebuildleavesdist/as it was), and its--checkflags a consumer BEHIND the version or NOT REBUILT since its install. wattyzer and both yunovatios GUIs had stayed two rounds behind that way, on ten JSON-viewer keys of 7.23.144 they did not define.
The JSON viewer compares two documents, and keeps them (gobj-ui 7.23.144)¶
kernel/js/gobj-ui -> 7.23.144: C_YUI_JSON_PAD gets a second pane on demand
and a compare that shows the differences of the two documents -- one row per
id of the flat form (json2flat / flat_diff), added / removed / changed --
in place of the viewers. The pasted texts and the layout are kept in
localStorage (storage_key), so the pad opens as it was left. gui_agent
0.22.61 and gui_treedb 0.17.31 take it with the new keys; the gobj-ui demo
(demo.yuneta.io, niyamaka.com) opens it from its top toolbar too.
gui_treedb 0.17.32: its About no longer lists the connections -- the Diagnostics table keeps the deployment identity and the session.
gui_agent 0.22.62: “For TreeDB” names a production yuno after its production node. A local copy carrying the production names resolved to the same url, and the export kept the first row -- the development machine, first by host.
gui_treedb 0.17.33: Connections has one row per connection (no service rows) and its checkbox marks a connection; the Topics / Graphs pickers list only the connections connected or marked, each with every service it discovered, and the transport asks for all of them in its identity card.
A JSON viewer in every SPA, node lists by name (gobj-ui 7.23.143)¶
kernel/js/gobj-ui -> 7.23.143, and the same range in the consumers: a JSON
viewer in the account menu of every SPA -- paste JSON from outside and read it
with the library’s viewer (setup_json_pad / C_YUI_JSON_PAD) -- and a treedb
topic table that opens sorted by id. yunos-js: the node lists of gui_agent are
alphabetical, and gui_treedb’s About gains the Diagnostics table gui_agent has.
Then, on the key:value rule (the key of a data record is id, the rest is
value): gui_agent’s node rows carry their key in id and its trees sort by it
(0.22.59); the “For TreeDB” connections document gives every record an id
(<node>^<yuno_id>) and a role^name label, and gui_treedb keeps that id on
import and recognises a connection by it (gui_agent 0.22.60, gui_treedb
0.17.30).
create-config says a __version__ must be a string¶
A config carrying "__version__": 1 was refused with “Configuration version
is required” -- for a file that plainly has one: kw_get_str() gives the
default for a number (and logs “path MUST BE a json str”, which names the
cause). The answer now says “Configuration version is required, as a
string”, with the yuno identity in front. Found when yunetas sync-configs
picked up a data file of yunovatios’ batches (renamed there with the _
prefix of the other data files).
v7.19.0 (2026-09-10)¶
The main topic of a treedb: hierarchical, and markable in the schema (gobj-ui 7.23.142)¶
treedb: a schema topic can carry
main_topic: true-- the topic the tree of the treedb hangs from, for a viewer. Only a topic hooked to itself can be it, and one per treedb:treedb_open_db()logs either mistake and ignores the mark. It is schema metadata, stamped in memory on every open (notopic_version, nothing in the store’stopic_var).tranger2_topic_desc()returnssystem_topicandmain_topictoo, so thedesc/descscommands carry both marks to a viewer. Neither is readable fromcols, and until now neither reached one.The meta-schema (
treedb_system_schema) goes to 18: thetopicstopic (version 8) gains amain_topiccolumn, and the projector ofC_TREEDBwrites it. Raising the meta-schema re-projects every treedb of a node on its next start -- the designed path, nothing to do by hand.kernel/js/gobj-ui-> 7.23.142: the treedb graph’s main topic can only be a hierarchical topic (the reader’s pick, else the schema’s mark, else the one reaching the most others -- the agent’s treedb is drawn fromrealmsnow, not fromyunos); it opens on the first level; the schema editor edits and checks the mark.
Back to the treedb graph with its focus (gobj-ui 7.23.140)¶
kernel/js/gobj-ui -> 7.23.140, and the same range in the five consumers.
The topics view’s graph button returns to the graph as it was left, its
focused topic included, through a new shell helper,
yui_shell_last_route_under() (a page-lifetime mirror of the routes visited).
A node’s child spec can now declare remember_position, so a config can make
a treedb node’s tabs point at where each workspace was left -- yunovatios
turns it on for its treedb nodes.
The treedb topic cards lose the graph icon and gain the topic’s shape (gobj-ui 7.23.139)¶
kernel/js/gobj-ui -> 7.23.139, and the same range in the five consumers. A
card’s graph icon always entered the graph focused on that topic, over what
the reader had left there; the toolbar’s graph button is the way in now.
The card shows instead the topic’s version, its number of columns, the topics
it hangs from and the ones that hang from it -- from the desc, no request.
The treedb topics view opens the graph with no focus (gobj-ui 7.23.138)¶
kernel/js/gobj-ui -> 7.23.138, and the same range in the five consumers. A
graph button left of raw json opens the whole treedb as a graph with no
topic highlighted -- the only way in used to be a card’s graph icon, which
always focuses that card’s topic. Derived from the host’s card route
template, so gui_treedb, gui_agent and yunovatios get it with no change.
The treedb graph: a toolbar in three parts, and never a blank view (gobj-ui 7.23.137)¶
kernel/js/gobj-ui -> 7.23.137, and the same range in the five consumers.
The toolbar reads left to right as build (layout, operation mode), show (fold pair, node views) and ask (find, refresh, raw json).
fix: the graph could open blank, the tree only in the minimap. The saved camera held the first root of the tree at its pixel even when both were off screen, so a layout that moved slightly left nothing in view. The camera is now saved by the node nearest the middle of the viewport, an off-screen saved camera is not restored, and any placement that leaves no node in view is fitted.
A treedb topic keeps its colour (gobj-ui 7.23.136)¶
kernel/js/gobj-ui -> 7.23.136, and the same range in the five consumers.
The topic palette of the treedb graph and of the schema diagram goes by
alphabetical order of the topic names instead of the backend’s order, which
varies from load to load, so two topics no longer swap colours. Every topic
whose colour nobody chose may change colour once with this release; a chosen
colour is saved and not touched.
The treedb graph legend keeps its order (gobj-ui 7.23.135)¶
kernel/js/gobj-ui -> 7.23.135, and the same range in the five consumers.
The legend’s topic chips are in alphabetical order, always: starring a topic
as the main one no longer moves the strip, and the backend’s order -- not the
same from one load to the next -- no longer decides where a chip sits.
The treedb graph: reset to default, from every menu (gobj-ui 7.23.134)¶
kernel/js/gobj-ui -> 7.23.134, and the same range in the five consumers.
Node, port and edge menus each offer three resets -- this one, its kind (topic nodes, topic ports, edges of the same type) and all. A reset forgets the saved value, the element’s own and the defaults of its scope, so the library’s default comes back: a figure saved as a square returns to the circle. The node reset shows in every view; no reset moves a node. It replaces
reset sizes/reset topic sizes.fix: a Save froze the default style of every edge between two topics. Any line width but 2 was saved, and the default for such an edge is 1.6, so each Save wrote them all down with their theme’s colour (no more re-theming after a reload, no topic default reaching them). An edge is saved now only where it differs from what it inherits.
reset all edgesclears what the old Save froze.Six new i18n keys, added in gui_agent, gui_treedb, wattyzer and yunovatios.
The treedb graph: one outline, no browser menu, a way back (gobj-ui 7.23.133)¶
kernel/js/gobj-ui -> 7.23.133, and the same range in the five consumers.
The
shapeview wears the outline of the other two. A figure was drawn with a line width of 2 and never less, while the card and the pill wear 1. It wears 1 now, or the width somebody chose, and it is a circle by default instead of a square. The node popover saves a figure only when it was changed, so applying a colour no longer freezes the default one on the node.The browser’s menu never opens over the graph. G6’s context-menu plugin cancels the event
@antv/gsynthesises frompointerdown, not the DOM’scontextmenu, so a right click on an edge or on the canvas opened the browser’s menu. The container cancels it; a popover’s form field keeps it.An edge has a menu in edition mode (
edge properties,unlink), and a port out of edition gets its node’s menu instead of an empty one.The port menu gains the way back:
reset port,reset topic ports,reset all ports. New i18n keys, added in gui_agent, gui_treedb, wattyzer and yunovatios.
The rule, written down (gobj-ui 7.23.132)¶
The sweep of the last sections became a norm, in the place each reader
already looks: CLAUDE.md here (JS GUI conventions), gobj-ui’s README
(Conventions) and CLAUDE.md, and the CLAUDE.md of yunos/js, wattyzer and
yunovatios.
EVERY control carries a
titleAND anaria-label, and both are translatable. No exceptions.
A control is any input, select, textarea, button or anything that
behaves as one; all four attributes are written where the control is built. It
is a floor, not a preference — the whole ecosystem was measured against it, on
the deployed page, and the target is zero controls without a name.
The rule carries with it the list of what LOOKS like a name and is not: a
<label> beside the control (Bulma’s field), a <label for=x> over a
control that carries only name=x, a placeholder, and the visible text when
it hides on mobile or says the STATE rather than the action — plus the one
shape that IS enough, a <label> that wraps its control, and the one that is
worse than nothing, a LITERAL aria-label beside a visible i18n label,
which overrides the translated text for a reader.
And the two things no attribute reaches: what a WIDGET draws for itself, named
after the render and again on every rebuild; and an <option>, which is text
like any other. Ending with the check, which is the part a rule usually
forgets to say: not a grep — dump title/aria-label from the DEPLOYED DOM
and switch language.
The frontend view, and everything that was left (gobj-ui 7.23.131)¶
kernel/js/gobj-ui -> 7.23.131, and the same range in the five consumers.
The last windows nobody had dumped: the frontend view, the site map, about, preferences, the five login screens, the overlay layer (a toast and each of the three confirm shapes), and the yunovatios controlador, the one SPA that had never been read at all.
One defect, and it is a shape the sweep had not met: an <option>. The
frontend view’s eight layout labels were written into its LAYOUTS table and
put straight on the node, so a language change renamed the select
(“disposición”) and left its options reading “Vertical compact”, “Lanes
vertical”, “Dagre (top → bottom)”. An <option> is text like any other: it
carries data-i18n now, with value kept explicit — a translated option with
no value tells the FSM to enter a layout called “Vertical compacta”, which is
the trap 7.23.13 already paid for once. And the keys are SHARED with the
JSON graph’s own picker (vertical tree, dagre top-down, dagre left-right): the two graphs sit side by side in the same console, so the same
layout is called the same thing in both — the toolbar-vocabulary rule, applied
to a select.
In wattyzer’s login, the three quick-control aria-labels went through
t() with no key on the element, and paint_i18n made up for it with a
hand-written list of three selectors. It works, and it is a list somebody has
to remember to extend the next time a button lands in that header — so the keys
go on the elements and the repaint walks the attribute, the way the other three
SPAs already did.
Read back: zero controls without a name in any of them. What still does not change language is data (family names, a place, a fake schema’s fields), words that are the same in both languages (“No”, “I18n”, “Topics”, “Editor”) and readouts (“19 px”).
The Developer window read in English, and the guard could not see it (gobj-ui 7.23.130)¶
kernel/js/gobj-ui -> 7.23.130, and the vocabulary added in yunos/js (both
SPAs), with accent fixes in wattyzer and the yunovatios GUIs.
The dump pointed at the Developer window. In gui_agent and gui_treedb,
22 of its 32 elements never changed language: the trace chips, the group
labels, the view segments, the output titles. The library’s markup was already
right — it has carried the keys on the DOM since 7.23.113. What was missing
was the vocabulary: 26 keys neither yuno defined, so i18next answered each
with the key itself and the whole window read in lower-case English beside a
Spanish shell.
And it is a blind spot of validate-locales — a wide one. TRACE_DEFS,
the grp / mk_view / mk_expand / mk_dir / mk_out helpers and the
OUT_TITLES lookup all pass their key as a VARIABLE, so a scan of t("…")
sees none of the 46 keys this window asks for and reports OK with the
window entirely untranslated. The library lists them above TRACE_DEFS now,
for a consumer to copy — and the first version of that list was itself
incomplete (33 of 46), which the next dump said out loud: collecting by hand
misses the helpers you forgot you wrote. The list is a convenience; the dump is
the check.
Along the way, four Spanish words that had lost their accent — Creacion, Automata, Trafico, Periodico — and three “Solo” that need one. They came from wattyzer’s bundle, where these translations were first written, and travelled into every consumer that copied them.
Read back on all five deployed sites: 1 of 32, and it is “I18n”, the same word in both languages.
The demo had no guard, and it is the app that mounts everything (gobj-ui 7.23.128 - 7.23.129)¶
kernel/js/gobj-ui -> 7.23.129, and the same range in the five consumers.
The dump run against demo.yuneta.io. It is the richest surface in the
ecosystem -- every gadget of the library is mounted in one page -- and the only
consumer with no validate-locales. It came back with 61 controls with
no name and 40+ names that never change language, and most of what it found
lives in the LIBRARY, so it was true in all six apps at once.
Sixteen literal aria-labels on the library’s own widgets (7.23.129):
the pager’s back and discard, a window’s minimize / maximize / close, the
dock’s close, the breadcrumb nav, the wizard’s back and next, the file field’s
choose / remove, the toast’s ✕, the modal’s back and ✕, the confirm’s ✕. Every
one of them is icon-only, so the aria-label IS the name — and every one
was frozen English everywhere. The demo says it in one line: the wizard’s back
button read “Atrás” and announced itself as “back”.
Two more shapes from the same dump: Tom Select HIDES the <select> it is
given and draws its own box in front of it, so naming the original named the
element nobody can reach (7.23.128); and the topic table’s row icons — edit
and delete — had no name at all, written as an HTML string inside a formatter,
the one place a sweep of createElement2 specs cannot look. Three search boxes
had a placeholder as their only name, which vanishes the moment something
is typed; one of them composed it (t('coordinates') + '...'), so no key could
reach it.
In the demo: a validate-locales of its own, adapted to its inverted
convention (English is the source, so keys are English prose and only the es
bundle is checked), wired into its build. It found 100+ keys the library asks
for and the demo never translated, and three duplicates — one of which was a
real COLLISION: window is the library’s label for the dev window’s output
chip AND was the demo’s noun for its demo windows; the later one won, so the
chip read “ventana” in lower case. Its table was also a Tabulator with no
locale at all, which is worth fixing where a consumer copies the recipe from.
Two blind spots of the guard closed on the way: a key can arrive through a
local ALIAS of t() — yui_tabulator_i18n.js asks for its whole chrome
through tr(key, default), so a scan of t( saw none of it and the paginator
sat in English with the guard reporting OK — and a key of the demo can carry
an escaped quote, where [^"]+ handed back half a key and demanded a
translation that was already there.
Read back at the end: zero controls without a name across sixteen chapters, and what still does not change language is the fake schema’s field names and the graph legend’s topics — data, not prose.
The same dump on the two yunos (gobj-ui 7.23.123 - 7.23.127)¶
kernel/js/gobj-ui -> 7.23.127, yunos/js -> both SPAs on it, and the same
range in wattyzer and the two yunovatios GUIs.
The deployed-DOM dump of the previous section, run against gui_agent and
gui_treedb. Zero controls without a name in either -- the shapes the
sweep chased are gone from this layer -- and two defects that only a dump
finds, both of them about a name that is right once and wrong afterwards.
The shell declaration held its keys in Title-Case English.
app_config.json carries i18n keys as DATA, and nine of them in each app were
written as English prose. i18next answers an unknown key with the key itself,
so four toolbar buttons of each app announced themselves as “Toggle
language”, “Account menu”, “Agent Console home” and “Toggle dark theme”
in a Spanish session, and never changed. They are lower-case keys now -- the
convention the locale files state in their own header -- and both locales
define them. validate-locales reads the file in both yunos as well now, and
collects a value only when it LOOKS like a key: one that does not is a literal
the config carries on purpose (ES/EN on the language button, the way 19 px
is a readout and not prose).
Tabulator’s row-selection checkbox said “Select Row” in every language, and getting that right took four attempts, each one of which the dump refuted:
7.23.123renames it after the render, the way the header filters already were. The formatter hard-codes the label with no locale key and no option.7.23.124: the rename was a silent no-op -- it reached the table’s node withtable.getElement()and returned quietly when that was not a function, and on a Tabulator INSTANCE it never is: only its Column and Row COMPONENTS carrygetElement(), the table itself carries.element.7.23.125: an empty table renders no body, sorenderCompletenever fires -- and the one box such a table does draw is the header’s. Eight tables were on screen in the schemas window with six boxes and not one row.7.23.126:yui_tabulator_relocalize()named the boxes BEFOREsetLocale(), and the re-render that call triggers rebuilds the header, so both names were written onto a header about to be discarded.7.23.127, which is the one that holds: the header is rebuilt far more often than it is built, so the naming hangs offcolumnsLoaded. A language change runssetColumns(), the column chooser runs it, and each rebuild draws a blank box again.renderCompletestays for the BODY boxes, which a sort or a page redraws on their own.
That last chain is the lesson of the section: a name put on a widget’s DOM is not a fact, it is a race with the next render, and the only way to know which one won is to read the page after the renders that matter -- a language change above all.
A name for every control, read back from the deployed DOM (gobj-js 7.16.6, gobj-ui 7.23.121 - 7.23.122)¶
kernel/js/gobj-js -> 7.16.6, kernel/js/gobj-ui -> 7.23.122, yunos/js ->
both SPAs on those two, and the same in wattyzer and the two yunovatios GUIs.
The tooltip sweep of the previous section ended on the windows nobody had
looked at yet -- devices, and whatever was left. The method changed for this
last stretch, and that is the part worth keeping: a static sweep names the
SHAPES of the defect, but only the deployed DOM says which controls a
reader can actually name. Dump every input/select/textarea/button of a
window, resolve its accessible name the way a reader does (aria-label, then
label[for], then a wrapping <label>, then the text, then title, then
placeholder), switch language, and diff. Two things fall out of that dump
that no source scan produces: a control with NO name, and a name that does not
change when the language does.
In the framework, one line and it explains a class of bugs:
refresh_language() looks its four attributes up with querySelectorAll,
which searches DESCENDANTS and never returns the node it is called on. So a
caller that hands it the very element carrying the key got the children
translated and that element’s own attribute left in the source language --
invisible in English, where the key IS the text. It bites where a widget is
built lazily and translated as a unit: the toolbar’s language dropdown opened
as role="menu" aria-label="select language" beside its own trigger reading
“Elegir idioma”.
In the library (7.23.122): yui_toolbar()'s scroll arrows shipped their
raw i18n KEY, waiting for the host to repaint them -- but a view’s toolbar is
REBUILT on every action it carries, and a host that translates its tree once
at mount never sees the new arrows. A deployed Spanish map offered “scroll
left”. They go through t() where they are built now, and keep their keys so
a language change still reaches them.
In wattyzer: 24 controls with no name at all, and 17 literal
aria-labels. A Bulma field puts the <label> BESIDE the control, with no
for and no wrapping, so it names the box for the eye and for nothing else;
the name goes on the control itself, from the label’s own KEY, so the two
cannot drift. Fifteen of the literals sat on buttons that ALSO carry a visible
i18n label, where the literal overrides the translated text for a reader. Two
more findings came from the dump alone: a language switch never reached an
OPEN modal (every modal of that app is appended to the body, and the switch
repainted only the shell’s container), and the brand announced itself as
“wattyzer go to root” in both languages -- app_config.json holds i18n keys
in DATA and nothing asked for them, so validate-locales reads it now, the
way the yunovatios copy already did.
In yunovatios: the two census selects of the device card and the
select-all checkbox of the controllers view had no name; and maplibre labels
its own chrome in English, never through the app’s i18next, so zoom, compass,
fullscreen, geolocate and the attribution toggle all read English on a Spanish
map. The map takes locale: yui_maplibre_locale(t) at construction --
maplibre reads it once, when each control builds its DOM -- and
yui_maplibre_relocalize() on a language change, which is the half a locale
alone cannot do. Same open-dialog fix as wattyzer’s.
Read back at the end: zero controls without a name across eight wattyzer windows and seven of central; what still does not change language there is treedb DATA (family names, a place), which is right.
A contrast sweep of the whole GUI, measured (gobj-ui 7.23.84 - 7.23.120)¶
kernel/js/gobj-ui -> 7.23.120, yunos/js -> both SPAs on ^7.23.120,
and the same range in wattyzer and the two yunovatios GUIs. It began as a review
of the treedb graph round and became a sweep of every surface the library
draws, with getComputedStyle in a browser and not an eye: a ratio ends
an argument, “it looks washed out” does not.
The review of the graph round turned up five things (7.23.84): the
legend’s chip title carried no i18n key while the strip is only redrawn on
EV_LEGEND_STATE; the engine took the language change from a raw i18next
listener instead of the shell’s event; the node, edge and port popovers left
their PREVIEW standing when dismissed with anything but cancel, and the
next Save wrote it; the wheel’s two keys named different gestures, so Cmd +
wheel did nothing on a Mac; and a port was hit by a radius although it can be
a square or a diamond. Then 7.23.85: the colour applied to a CARD lived on
the live style alone, which on an html card sits under the html -- the first
repaint rebuilt it from the topic’s colour and nothing ever reached
__graphs__.
The measured part, surface by surface:
Toolbars (
7.23.86,7.23.87): one row held 16px, 18px and 24px glyphs, becauseemis measured against the button --1.5emis 24px in a plain one, 18 in anis-smallchip, and a labelled button kept the inherited 16. The size lives once now, inrem, inyui_toolbar.js, and every view toolbar of the library speaks it.Brand colours as INK (
7.23.89): measured against the scheme background, every raw--bulma-<name>fails the 4.5:1 of small text in one scheme or the other --link3.51 in dark,warning1.75 in light. Eleven sites moved to-on-scheme; a FILL keeps the raw token and pairs with-invert.The graph’s own drawings (
7.23.90,7.23.91,7.23.92): a card writes on a tint of its own topic colour, where the topic name read 2.47:1; a port is a knob half on that card, and its rim was the same colour darkened 20%; the minimap’s blocks were the colour of its own paper (1.61 on white); the focus amber was one colour for both schemes (2.15 on white).The form and the tables (
7.23.94): a placeholder is TEXT and Bulma paints it at 30% alpha (2.47 / 1.79) -- it is what says what a box is for.has-text-grey, one mid grey for both schemes, is gone from the library.The shell (
7.23.95,7.23.96): blue ink on a blue tint in the two places the shell does it, and two dimmed states that WCAG exempts and a reader does not.The chart (
7.23.97) was not a contrast question but a THEME one:c_yui_uplot.jshad not one reference to a scheme, so uPlot drew a chart for a white page -- axis ink at 1.16:1 on the dark ground, a grid at 1.04, and series namedblueandorange, one of which was always a rumour. And a map label is not text on a page: what is behind it is a tile, so it has a halo now.
Seven options of C_G6_NODES_TREE that no host could reach were forwarded on
the way (7.23.93), because the view that creates the engine is what makes
an option exist.
And the last thing it left open, closed (7.23.98): connected /
disconnected was told apart by green vs red alone, in the map’s circles as in
its labels -- the one pair colour blindness does not read, and in greyscale
two discs are the same disc. A device that is down carries an exclamation
mark, and a cluster with something down carries it inside its own count
(4 !). It rides in the count’s text-field and not in a second symbol
layer on the same point: that is a placement question, and the badge lost it
-- with the filter removed as well, which is how the placement was told from
the expression. Verified on both cases, the negative one included: with all
four devices connected the cluster reads 4 in green.
One more the sweep did not reach (7.23.99): the legend’s star for a
main topic the reader CHOSE. Its button carries pressed_state, which
inverts the ground, while the glyph kept the gold set INLINE -- and an
inline colour beats the ink the pressed rule sets with the fill. That one
star sat on the scheme’s own text colour: 1.06:1 in dark, ~2:1 in light,
while the four stars merely OFFERING to become the main one were bright.
It survived the sweep because only a CHOSEN main is drawn pressed, and
main_topic is a per-treedb per-user preference -- the same strip reads
right wherever nobody has picked one.
Making the glyph inherit the pressed ink fixed the ratio and cost the
COLOUR, which is what the star says -- so 7.23.100 stops drawing that
star pressed at all: it keeps its gold and wears the ring the focused
chip already wears (box-shadow: inset, because a chip in a has-addons
group clips a shadow outside its box). The strip now says the two things
separately: gold star = this is the main topic, ring around it = a reader
CHOSE it, and pressing hands the choice back to the graph.
And asking why that star was not simply wearing Bulma’s .button.is-active
turned up the answer to a different question (7.23.101): the library had
three spellings for “this toggle is ON”. The reason was where the rule
lived -- pressed_state sat in lib_graph.css, which reaches an app only
if the app mounts a graph, because a stylesheet rides its JS import. So the
JSON viewer’s view switch had grown a --bulma-link fill of its own (a
STATE wearing a colour that names a kind of ACTION), the gclass viewer’s two
switches wore is-active with no rule behind them anywhere (1.3:1), and
yunovatios kept a copy of the rule under a third name, is-pressed, with
the reason written in its app.css. The rule and set_pressed_state() move
to yui_toolbar.js / .css, where a toolbar’s toggle belongs
(lib_graph.js re-exports the helper); every one of those sites now says it
the same way. Bulma’s own is-active, measured, is a 10-point lightness
shift -- 1.27:1 light, 1.33:1 dark, against 9.44:1 / 8.46:1 -- and it shares
its declaration with :active, the look of a finger down right now.
7.23.102 finishes it in C_YUI_PERIOD, the last place a STATE wore a
colour: the granularity in use and the picked calendar cell were an
is-link fill and are pressed now. The overflow granularities keep
is-active and that is the RIGHT call -- those are dropdown-item links
in a menu, where it is Bulma’s own mark for the current item; the button
standing in for them looks pressed like any other button. ARIA follows the
element rather than the look: aria-pressed on the strip’s buttons,
aria-current on the menu item, aria-current="date" on the picked cell
-- the widget carried none of the three.
And then the unification turned out to have unified one thing too many
(7.23.103): a TOGGLE and a SELECTOR are not the same control, and the
pressed pill was answering a question only one of them asks. A toggle is
on or off -- no which one -- and keeps the neutral pill. A selector
picks ONE of N, and there the eye reads a row of identical grey buttons
by HUE, not by comparing shades, so the chosen segment is FILLED with the
link colour (selected_state, beside the pressed pair in the toolbar
module). Measured against the button’s own ground the fill is 5.14:1 in
light and 3.53:1 in dark -- less luminance than the pill (9.44 / 8.46)
and a different hue, which is what says this one, of these. Filled: the
graph’s three node views, the JSON viewer’s three, the gclass viewer’s
two switches, the period’s granularities and its picked calendar cell,
and the map’s three modes in yunovatios. Still pressed, because they are
toggles: node labels, the legend’s loose and focus buttons, the graph’s
anchor and its selection mode, the alarms view’s active only.
And the graph’s camera belongs to the reader (7.23.104). A refresh, a
change of node mode and a new main topic each asked for a full refit, and a
refit throws away the one thing the reader had decided: those three rebuild
the CONTENT, and content moving is no reason to move the reader. They hold a
NODE at its pixel now, with the pair the folds already used
(yui_graph_viewport_of / yui_graph_place_at) -- which keeps the zoom,
because it only translates, and survives a relayout that moves everything.
It also survives a RELOAD: C_G6_NODES_TREE takes a camera attr and
publishes EV_CAMERA_CHANGED {zoom, x, y} when a move settles (700 ms of the
browser’s timer -- a wheel notch, a pinch and a drag each fire
aftertransform many times, and what is worth saving is where the gesture
ENDED), and C_YUI_TREEDB_GRAPH persists it under the view’s name like
main_topic.
It took two more to actually do it, and both are worth keeping. 7.23.104
brought the ZOOM back and put the graph somewhere else, because it saved
G6’s getPosition() and replayed it with translateTo() — which are each
other’s inverse only at zoom 1, a trap this library had already paid for
once and written down (yui_graph_place_at() exists for it). A camera is
the zoom, a NODE and the viewport pixel that node was on, restored with
that helper — the same thing the folds keep across a rebuild. Then
7.23.106: the VIEW rebuilt the payload as {zoom, x, y} on its way to
the store, dropping the node, so the engine refused its own camera back.
Measured on the deployed gui_treedb, with a pan so the framing is nobody’s
default: the same card at x=337 y=152 before and after F5, at 121%. The
lesson for the next one is the checking, not the arithmetic — 7.23.104 was
called verified on a zoom readout alone, and the zoom was the half that
worked.
And the graph that looked out of focus was one missing border
(7.23.107). Two reports, one cause: a card, a pill and a figure are one
record drawn at three sizes, and they were speaking two colour languages.
The card is a TINT of the topic colour with the vivid colour around it; the
figure of shape mode was the raw colour filled, ringed by
getStrokeColor() — the same colour darkened 20%, which measures 1.58:1
against its own fill. Not a border, a ramp; and twenty-five of them in a
fan is what reads as a blurred picture. The figure wears the card’s paint
now: 4.67:1 fill to rim, 12.08:1 rim to canvas, at lineWidth: 2 — checked
in the rendered pixels, where the rim is a 2px band with one antialiased
pixel either side. The colour was never lost, by the way: the figure’s fill
measured exactly the legend chip’s RGB. What differed was the treatment.
A tranger record is a DOCUMENT and is read with the viewer now. Its
dialog showed a <pre> with a Copy button under it -- a jwt payload with
its roles and its allowed origins read by scrolling text -- while
C_YUI_JSON sat one gclass away, already hosted by that same view for the
raw tranger. It gets the whole record, so it never asks for a subtree; one
viewer at a time, destroyed with the dialog (a viewer left alive holds its
service NAME and the next record finds it taken). 7.23.108 gives the
shell’s dialog a wide option for the ones that hold a document: 640px is
a width for a question with two buttons, and at that width the viewer
wrapped every long value and pushed its own view switch behind the
toolbar’s arrow.
And the wheel means one thing now, in all three graphs (7.23.109).
C_G6_NODES_TREE has scrolled on the wheel since 7.23.75 while
C_YUI_JSON_GRAPH and C_YUI_GOBJ_TREE_JS went on zooming with it — the
same gesture with two meanings in graphs a reader has open side by side,
and the JSON viewer is reached from inside the other two. The pair moves
to yui_graph_camera.js, where the rest of the camera vocabulary already
lives, with its two traps beside it: the scroll’s enable and the zoom’s
trigger must name the SAME key (a G6 trigger is a CHORD), and the zoom
carries animation: false, which the treedb graph was the only one
missing. Measured in the demo on all three: a plain wheel leaves the zoom
untouched and moves the drawing, Ctrl + wheel goes 100% → 150%.
The MAP too (7.23.110): maplibre’s cooperativeGestures, which never
blocks a wheel carrying ctrlKey — so the trackpad pinch survives — and
whose own notice teaches the gesture, which a graph has to do without. It
is set where the map is BUILT and not in the attr’s default value, and that
is the part worth remembering: a JSON attr is replaced WHOLESALE by a host
that passes its own, so a default is a suggestion. The demo passes
map_settings and never saw the first version of the change; the same trap
the SDK documents for a crypto override.
And with the gesture, maplibre’s own words (7.23.111): its zoom
tooltips, geolocate button, attribution toggle, popup close and that very
notice were English inside a Spanish app. yui_maplibre_locale(t) gives
the map its locale at construction and yui_maplibre_relocalize(map, t)
does the half a locale cannot — maplibre reads its strings ONCE, when each
control builds its DOM, so a language change has to rewrite what is drawn.
The notice is also HELD for 2s: maplibre shows it 100ms and fades it for a
second more, and this GUI has no transitions, so what was left was a blink
— worse than no notice. That recipe was already solved in yunovatios’
yv_map_base.js, which is where it was found; the library learnt it
instead of inventing a second one, and the app keeps its own.
And then the rest of the tooltips (7.23.112), swept for the two
defects that make one: a title written as a literal, and a title: t(…)
with no data-i18n-title — translated once and frozen for the life of the
view. Four real ones out of 51 hits: the FORM’s toolbar (save, undo,
clear, copy, paste), where the visible label carried its key and the
title beside it was raw English — so what a pointer reads and what a
screen reader announces were the untranslated half; the tab CLOSE of
C_YUI_NAV, the same shape; the map’s own three controls, which read
their keys once and showed maplibre.drag_mark itself wherever a consumer
had not defined them; and the treedb table’s operations column, titled
'Op', which takes a titleFormatter now because a language change
re-runs setColumns() over the SAME definitions. The other 47 are not
defects, and knowing why is the useful half: a title: handed to a modal
or a window is an i18n KEY the helper translates, and the G6 plugin
toolbar is rebuilt whole on a language change.
And the Developer window, which had no i18n at all (7.23.113): its
~30 strings — the trace chips, the view and output selectors, the
direction filters, the search placeholder, copy and clear, the muted row
and the window’s own title — were English literals. A debugging tool is
still a tool somebody reads. Three details make it RE-translate rather
than translate once: TRACE_DEFS carries the i18n KEY where it carried
the label, because refresh_dev_chrome() repaints those chips from
data-label on every toggle; the two COMPOSED titles are gone, since
'Show ' + label + '…' is a string that is no key; and the ⊘ Periodic
chip is a glyph span plus a labelled span, because refresh_language()
replaces the FIRST text node of the element carrying data-i18n and
would have eaten the glyph with the word.
And the schema editor’s forms had no NAMES (7.23.114, 7.23.115):
field() draws Bulma’s shape, where the <label> is a SIBLING of the
control with no for and no wrapping — so it named the box for the eye
and for nothing else, and a reader announced every input of the column and
topic forms as unlabelled. The name is put on the control from the label’s
own key, in field() and not field by field. It took two goes: the first
named the outer tag, and select_input() returns Bulma’s
<div class="select"><select>, so every SELECT stayed anonymous while the
inputs beside them were fixed — measured on a deployed schema. It descends
now. The flag checkboxes are NOT a defect: each sits inside its own
<label>, which IS its name.
And the same defect in disguise, in C_YUI_FORM (7.23.116): there
the label IS written <label for={name}>, which reads like a correct
association — except for matches an id, and no control of that form
sets one; they carry name. So every field of every form the library
builds (the treedb record editor, wattyzer’s, yunovatios’) was unlabelled
for anything that is not an eye. The name goes on the control from the
label’s own key, in the one place a field is finished — not as an id,
because two forms can be open at once and duplicate ids would break both.
The graphs window closed it (7.23.117, 7.23.118), and the last two
came from READING the deployed toolbar rather than the source: the treedb
graph’s two selects had no name — their rótulo is a <span>, and an
is-hidden-mobile one, so on a phone the control says only its current
value (the other three graphs already named theirs); its refresh button is
an icon plus an is-hidden-mobile label and nothing else; and the
toolbar’s own scroll arrows said scroll left in English beside a Spanish
toolbar, because yui_toolbar.js asks for those keys and NO consumer had
defined them. That last one is worth keeping: a key asked for by a module
the app does not import DIRECTLY can slip past validate-locales, which
is how two of them stayed missing in five apps at once.
And the last one is Tabulator’s own DOM (7.23.119, 7.23.120): it
draws one filter box per filterable column and gives it NOTHING — no
label, no aria-label, no placeholder — so the alarms window had five
anonymous text boxes under Equipo / Nombre / Alarma / Estado /
Descripción. The name is composed from the column’s title, which costs one
consumer key with an interpolation instead of one per column, and it runs
on a language change too. It took two goes: the obvious key name,
filter column, ALREADY existed in two apps as that box’s placeholder and
carries no interpolation — so the first version named all five boxes the
same, which the deployed read caught.
Three views of a record, the tree reads down, the wheel scrolls (gobj-ui 7.23.75 - 7.23.83)¶
kernel/js/gobj-ui -> 7.23.83 and yunos/js -> both SPAs on ^7.23.83.
Asked for together, all in the service of one graph with many nodes.
The wheel SCROLLS the graph; Ctrl + wheel zooms, in every operation mode
of every graph on C_G6_NODES_TREE. A wheel that zoomed made a graph taller
than the screen a thing to be looked at from afar or read through a keyhole
-- never scrolled, which is what a wheel does on a map and on every page.
Shift + wheel goes sideways, a trackpad pinch arrives as Ctrl + wheel, the
touch pinch is untouched.
The treedb tree reads DOWN. Read right it was dagre with the siblings
held still, and nobody could tell the two apart; down is where a tree has
room. The algorithm is written once and the top-down tree is the same tree
fed transposed cards and read back transposed. treedb-outline is gone -- a
list that indents is a JSON viewer, and the library has one.
A record has two SHAPES. Open is the card with its ports, the only shape
a link can be edited on; closed is a rounded square of the topic’s colour, no
ports, no text -- the topology alone, as a native G6 rect so focus,
selection and anchor paint on its own stroke. Two persisted toggles in the
view’s toolbar (node_mode, node_labels), and one node against the rule
by double click or its context menu. The ports of an open card are handles
now (radius 10 / 5, 2 px). Consumer i18n keys: closed nodes, node labels,
open node, close node; treedb-outline dropped.
7.23.76 is the one that shipped: 7.23.75 left the outline’s case in the
view’s option_label(), and every consumer’s validate-locales refused the
build for a key asked for and defined in no locale -- the guard doing its job
before a deploy. 7.23.77 shrinks the +N chip beside closed nodes (40×22,
following the shape of the card it continues), and gives the test-app’s
graph a page of one so a chip shows. 7.23.78: expand all and collapse
all leave the ZOOM alone -- both fitted the whole graph, so opening
everything zoomed out to a strip and closing it zoomed in on the roots; the
camera holds the anchor still, else the first root, as a single fold does.
7.23.79: the default ports go to radius 14 / 8, because on a deployed
treedb 10 / 5 still read as dots -- and 7.23.80 is why they did: Save wrote
the SIZE and the port radius of every card into __graphs__, chosen or not,
so every Save froze the library’s size of the day and no later default
reached a saved treedb. A size on the tier’s default is not saved any more,
a closed node saves no size at all, and the node’s context menu (edition)
gets reset sizes / reset topic sizes, which forget every saved size and
put the defaults back on the spot. Consumer i18n keys reset sizes, reset topic sizes.
A port has a SHAPE, and its own properties popover (7.23.81). G6 draws
every port as a circle, so a record’s card is now a node of its own,
treedb-card -- G6’s html node with drawPortShapes overridden -- where a
port’s shape is circle, square, diamond or triangle; r stays the
one size. A selected port shows a gear beside it, as the node does, and the
port context menu has the same entry for a finger: shape, radius and scope
(this port / the same port of every card of the topic / every port), with a
live preview. Remembered as the topic’s default and saved per node in
__graphs__ (port_shapes). The node and edge popovers’ labels had no i18n
key at all and rendered in English in every language; keys added to the
consumers with the port’s.
And a record has THREE views, not two (7.23.82), the same three the
gobj tree view offers: expanded is the card with its pills and ports;
compact is a one-line PILL with the name inside and small ports, so a link
can still be drawn; shape is the FIGURE of the topic’s colour, no ports, no
text -- and the figure is chosen now, square, circle, diamond, triangle,
hexagon or star, in the node properties popover, remembered as the topic’s
default and saved per node as node_shape. The three views are three push
buttons in one group (7.23.83: a view is picked at a glance, and the
pressed one says which is on); the labels toggle is enabled on figures only. Consumer i18n
keys nodes, full, compact, expand node, collapse node, hexagon,
star; closed nodes, open node, close node gone.
The focused legend chip is highlighted, not pressed (gobj-ui 7.23.74)¶
kernel/js/gobj-ui -> 7.23.74 and yunos/js -> both SPAs on ^7.23.74.
Two loose ends of the 7.23.73 review, found reviewing the review.
The body of a focused chip wore a state it does not own. The body is the
show/hide toggle of its topic, and that is what its aria-pressed says; the
focus belongs to the crosshair next to it, which is the button that looks
pressed. 7.23.73 painted the body with pressed_state as well, so the eye
and a screen reader read two different states off one button -- and the
count, has-text-grey by Bulma’s !important, sat grey on the near-black of
the pressed look, at 2.7:1 against the 4.5:1 a text needs. The focused chip
now carries an inset ring (GRAPH_LEGEND_FOCUSED) and repaints nothing in
it. The header of treedb_layout.js also said two layouts and two adapters;
there are three of each since 7.23.70.
A review of the graph and authz rounds (gobj-ui 7.23.73)¶
kernel/js/gobj-ui -> 7.23.73 and yunos/js -> both SPAs on ^7.23.73.
Three findings of a review of the 7.23.66-7.23.72 round; the backend
half of the same review lives in the yunovatios repo.
The treedb layouts died on a deep tree. layout_tree, layout_radial
and layout_outline walked the spanning tree with recursion -- the natural
way to write it, and the one the data breaks: a self-referent hook (a place
inside a place inside a place) is as deep as the store says, and the stack
is not. A chain of 20000 nodes answered RangeError: Maximum call stack size exceeded, and a graph that cannot be drawn is not a layout choice, it is an
exception in the console. The five walks share one walk_order() now --
parent before child, so it reads forwards for a pre-order pass and backwards
for a post-order one -- and the new test was first made to FAIL on the old
code, which is the only way a regression test is worth having.
A state was painted with a colour of the palette. The focused chip, the
focused crosshair, the shown loose records and the chosen main topic each
paired a STATE with is-primary/is-warning -- exactly what
set_pressed_state() was written against in 7.23.12, where every colour of
that palette names a KIND of action, so a state wearing one reads as a
category. Same lesson, the other strip.
And the legend said its topic names at 0.75rem. 7.23.69 raised the
GLYPHS of the chips out of is-small and left the name and the count behind,
the name being the thing a layer control is read for. The button stays small
so the strip keeps its height; the label comes up inside it, the way the
glyphs did.
A treedb form could not SAVE, and a user could not be given a role (gobj-ui 7.23.71 + 7.23.72)¶
kernel/js/gobj-ui -> 7.23.72 and yunos/js -> gui_treedb 0.17.28 /
gui_agent 0.22.57, both on ^7.23.72. Two defects of the same form, found
one behind the other on the deployed console.
An fkey is edited by LINKING, and writable does not govern it (7.23.71).
7.23.55 fixed a real bug -- a <select> ignores readonly, so a column
without writable rendered an EDITABLE select -- by disabling what the
attribute cannot reach, and it caught the fkey with it. An fkey is normally
declared with no writable flag at all (treedb_authzs’s users.roles is
['fkey']), so the Role of a user opened as a dead grey box with its four
options inside, while the topic view was still sending fkeys back for exactly
that reason. The rule is now one pure module, form_field_readonly.js, and a
section of the gobj-ui README: a form opened to LOOK still has no editable
field, fkey included.
And with the control enabled, no record could be SAVED at all (7.23.72).
Since 7.23.64 -- the guard that makes the form busy while it reads the
picked files -- ac_form_save_record() read priv.reading_files without
declaring priv, so the FIRST line of every save threw ReferenceError and
the dialog just stayed open. Three releases with a treedb form that could not
write. priv_declared.test.js now guards the whole class: every function in
src/ that reads a bare priv. must declare it, take it as an argument, or
sit inside one that did -- a missing declaration throws only when its LINE
runs, and that line is usually the first of a path nobody walks every day.
Three layouts made for a treedb (gobj-ui 7.23.70)¶
kernel/js/gobj-ui -> 7.23.70. treedb-tree (a tidy tree read left to
right, the new default), treedb-outline (one node per row, indented by
depth) and radial (a sector per subtree, each ring as far out as its cards
need), all on the same deterministic spanning tree of what is on screen: the
first parent that reaches a node keeps it, children by hook then record
order. O(n), no crossing heuristic: opening a hook moves nothing that is not
under or beside it. The study of G6’s own layouts against a treedb is in the
gobj-ui README.
The treedb graph’s legend is its layer control (gobj-ui 7.23.69)¶
kernel/js/gobj-ui -> 7.23.69. The legend strip is always there, one chip
per topic: the body shows or hides the topic, ☆ makes it the main topic,
+N shows its loose records, ⌖ highlights it. A main topic governs the
tree -- deduced when none is chosen, its parentless records are the roots,
and a parentless record of a topic the schema hangs from it is loose and
drawn only on request. The three settings are preferences per treedb
(hidden_topics, main_topic, loose_topics). A click on the legend
reveals the whole topic; the per-topic route one page (7.23.67).
The treedb graph opens FOLDED, like a JSON viewer (gobj-ui 7.23.66)¶
kernel/js/gobj-ui -> 7.23.66 and yunos/js -> gui_treedb on ^7.23.66. A
treedb drawn whole was a pile: a nave with a hundred and forty devices was a
row of a hundred and forty cards, and a 6400-record treedb built a DOM card
for every record before the first pixel. The graph now reads the treedb as a
tree by its hooks: every record is fetched, only the visible ones become G6
nodes -- the roots, expand_depth levels under them (default 2), one page
(fold_page_size, default 24) of children per hook, a +N chip after each
page, and a pill per hook on the card (▸ devices 142) that opens or folds
it. dagre reads left to right, with the ports turned to the sides.
The arithmetic is treedb_fold_model.js, pure and tested. New consumer i18n
key show more. The JS API docs are repinned to the new tag.
v7.18.2 (2026-09-05)¶
A second arrival under the same name appends nothing¶
store_file_bytes() wrote the manifest’s original_name into the asset node
whether or not it was the name already stored, and every update appends an
instance: a repeated import-assets over the same directory added one
instance per asset that said exactly what the previous one said (12,134 rows
of nothing on the yunovatios census). The name is compared with the stored
one first, and only a NEW name is written. Test 17 of
tests/c/tr_treedb_files.
A create of an existing id stored the file before refusing¶
treedb_create_node() ran treedb_store_files() before looking the id up,
so a create-node refused as “Node already exists” had already written
the blob and the __assets__ row, and left both for the gc. The bytes are
taken only once the create is known to go ahead. Test 18.
The node of treedb_delete_node() is borrowed, not owned¶
The signature said owned. Every caller — C_NODE, gc-assets, the
rollback of a create — hands it the index’s own reference, which the delete
releases on success and leaves alone on a refusal. The comment, in the
header and in delete_node(), says so now.
And with the convention written down, one path contradicted it: the
“Not a pure node” refusal of delete_node() and of
treedb_delete_instance() released the node it had just refused to delete —
a reference neither of them owns, so the node could be freed while still
indexed. A refusal now leaves the node exactly as it was, still indexed,
like the immutable and snapshot guards beside it.
Documentation¶
Two standalone pages on the treedb, carded on the landing and installed by
deploy.sh: /treedb-files, one write
of a file column stepped through, and
/treedb-system-topics, what
__snaps__, __graphs__ and __assets__ hold and the door the system opens
each of them through. docs/doc.yuneta.io/build_artifact.py turns any such
page into the artifact version of itself.
v7.18.1 (2026-09-05)¶
A file column names its OWN asset hook, or the write is refused¶
treedb_store_files() expanded a bare id into the full reference treedb’s
links speak and took a reference that arrived FULL as it came — and
link_file_columns() links by the hook the VALUE names. So a record whose
foto said __assets__^<sha>^as_devices_qr linked the file into qr, left
foto empty and answered “Node created!”: the client asked for one column
and the treedb wrote another. Since the write path started linking file
columns itself, that is reachable through a plain update-node, with no
autolink anywhere.
A full reference must now be exactly __assets__^<the id>^as_<T>_<C> of its
own column, and a malformed one is refused at the door naming the column,
instead of much later by filtra_fkeys() naming only the string. Test 14 of
tests/c/tr_treedb_files.
A create whose file column could not be linked answered success¶
The record existed and was indexed, so it was returned — with the column empty and the caller told the node was created, which is the failure the previous release closed for the write path wearing a success. The create is undone now: one message, and either it happened or it failed. If the rollback itself cannot run the node is returned, because a failure that left a record behind is worse than the success it replaces — and it is only undone when this create made the KEY, since an instance created for a secondary key shares its key with the node that was already there and a delete takes the key whole. With the check above, the path is only reachable on a broken invariant of our own.
A second arrival of the same bytes with no name WIPED the name¶
original_name is the only writable column of an asset and a second arrival
is an update of its node, so a manifest without original_name wrote ""
over the name the file first arrived under — and appended an instance saying
the file had arrived again, nameless. The GUI always sends a name; a C node
forwarding a record with its bytes need not. A manifest that carries no name
says nothing about the file and leaves the stored one alone. Test 15.
gc-assets did not take the bytes with no row, and the docs said it did¶
The collector walked the rows of __assets__ and never read the blob
directory, so the orphan blob that the write order deliberately allows — the
blob goes down BEFORE the index node, on purpose, because a node pointing at
nothing repairs itself never — was garbage nobody could ever reach again.
So was the .tmp of a write that never got to its rename. Both the design
note and the reference page promised that gc-assets took them.
sweep_orphan_blobs() runs at the end of the gc, over the union of every
treedb’s __assets__ index (.blobs is the TRANGER’s, and every open treedb
loads the whole topic, so that union is every row on disk); the dry run says
what it would take. A leftover of an id that still HAS a row is removed and
NOT counted as an asset collected. Test 16.
import-assets could be started AT a symlink out of its root¶
The walk follows no link it finds inside the tree — it reads every entry
with lstat — but the directory it STARTS at is opened, so
source_dir=taller/escape with import_root/taller/escape -> /etc walked
/etc as if it were the root’s own. The confinement is the whole security of
this command, so the path is now resolved with realpath() and must still be
inside the resolved import_root. tests/c/c_assets case 6 baits it with
the symlink it already created.
And two smaller ones¶
A node created for a SECONDARY key inherits the primary’s links, so asking
link_file_columns() to treat it as new linked every file column a second
time — a “Child already in parent hook” per column and an instance for
nothing. And build_blob_rel() in C_ASSETS checked the LENGTH of an id
where treedb_blob_path() checks its alphabet: one of the two builds a
served path, and they must not disagree about what an asset id is.
The write path of a file column links it itself, autolink or not¶
treedb_update_node() skips every fkey in its field loop — right for a link
a person edits by linking — and treedb_create_node() links nothing at all,
so the link of a file column existed only where an autolink followed. A
create-node, or an update-node without autolink, stored the bytes and
the index node, answered success and left the column as it was: an orphan
asset that gc-assets would take, and a device with no photo. The GUI always
sends autolink, so it never saw it; ycommand and the
EV_TREEDB_UPDATE_NODE event do not.
link_file_columns() runs inside both writes now: it links what the kw’s
file columns name, unlinks what they stop naming (""), and leaves a
column the kw does not carry alone. A create with a file column costs one
more instance, as the autolink a caller used to have to remember did. Test 11
of tests/c/tr_treedb_files covers create, move, clear and a reopen.
DESIGN-treedb-files.md §16.8 lists what the same review left open.
The snapshot guard of an asset cannot be bypassed from the wire¶
treedb_gc_files() walks the snapshots once for the whole run and told
treedb_delete_node() so with an internal key, __snaps_walked__, in the
delete’s options — and delete-node forwards its options from the wire as
they are, so a client with delete could spell the bypass and take an asset
a snapshot still needs. The bypass is a parameter of a private delete_node()
now; the public entry always walks. The old key is inert.
A file column that is not fkey is refused everywhere, not only at open¶
derive_file_hooks() refused it at open, fatally, but a run-time
create-topic got nothing but a log line: treedb_create_topic() ignored the
return, and the write path keyed on the word file alone, so the topic stored
its bytes into a column nothing links and gc-assets took them. The check is
one function now, check_file_column(), asked by treedb_create_topic()
BEFORE the topic exists (refused, cause in last_message), by the open — which
now logs with on_critical_error when a topic of the schema could not be
created, instead of leaving it for the first write to meet as “Topic name not
found in treedbs” — and by treedb_store_files(), which refuses the write.
Two treedbs on one tranger: gc-assets and delete-node read every treedb’s links¶
__assets__ is a topic of the TRANGER, so a tranger holding two treedbs
shares it — but each treedb loads its own copy of every node, and a link made
from a treedb lands in that copy. Read from one copy alone, “no live node
links it” was true of one treedb and said nothing of the other, and the row
and the bytes are the tranger’s: gc-assets from the first treedb took what
the second linked, and delete-node did too. No yuno opens two treedbs on
one tranger today; C_TREEDB can.
Both ask every treedb’s copy now. The delete refuses, force or not — force
unlinks the children of THIS copy and the other treedb’s would dangle — and a
deleted row leaves every treedb’s index, or the other treedb keeps answering
a row that is not on disk. The derived hooks are seeded into, and taken from,
every treedb’s copies rather than the deriving treedb’s only: the desc is
shared and get_node_down_refs() expects each of its fields in the node, so
the first delete of a copy that lacked the other treedb’s hook failed with
“field not found in the node”. Test 12 of tr_treedb_files opens a second
treedb on the test’s tranger and walks the whole of it.
gc-assets holds exactly what an activation would load¶
It held every tagged instance as “a snapshot needs this” — but
treedb_save_node() inherits the tag, so after a snap every later instance of
a node carries it too, and an activation loads only the NEWEST instance per
key under the snap’s tag. So the gc kept for ever whatever a node ever named
after its first snap, and nothing could release it. It holds, for every snap
that still exists, the newest instance per key under its tag, and only that:
a node that moves on releases, and deleting the __snaps__ row of a snap
(delete-node; there is no delete-snap) frees what only it held. A treedb
with no snapshot does not walk at all. Tests 6 and 13 of tr_treedb_files.
The gbuffer door works through EV_TREEDB_UPDATE_NODE too¶
The event handler handed the write path kw["record"] while the kw’s one
binary field rode at its top level, so a manifest of slices through the event
met “carries no bytes”, and uploaded_by stayed empty. It hands both over
now — without the extra reference the command path takes, because nothing
copies an event’s kw (take_files_gbuffer(kw_is_a_copy)).
A plain update cannot drop a seed link of a file column¶
The seed guard of mt_update_node() ran only with autolink; since the write
path links a file column itself, a plain update moves the ones it carries,
so those are asked the same question, and only those.
Three small ones of the file columns¶
files_content_types handed as TEXT through open-treedb (what a CLI does)
was read as an empty list and the default stood in, in silence — parsed now,
or refused with a log. import-assets refused any source_dir with .. as a
substring (v1..v2 included) — a path segment equal to .. now. A second
arrival of the same bytes no longer rewrites uploaded_by: that is who put
the BYTES there, and the bytes did not change; original_name still moves,
which is what the history of names is made of. And the files_max_size
description names the websocket’s ceiling: a frame has none of its own, its
gbuffer is capped by MEM_MAX_BLOCK.
gobj-ui 7.23.64: the form is busy while it reads the picked files¶
A second Save during the read sent a second write, and a Cancel threw the save
away without a word. The toolbar is disabled and the save button spins until
the read lands; a second Save is refused and says so; a read whose form was
closed meanwhile is dropped with a warning. Both SPAs take the range
(gui_agent 0.22.53, gui_treedb 0.17.25).
gobj-ui 7.23.63: saving a record unlinked its read-only file column¶
The topic view sends back only the writable cols, the fkeys and the pkey,
and the write goes out with autolink, which rebuilds the links from what
the record carries. A file column IS an fkey but answers type: "file"
since gobj-js 7.16.5, so a file column without writable — the one only a
load fills, which the open declares legal — fell out of the record, and every
save of any other field of its record cut its link, in silence. The exemption
asks is_file now. Both SPAs take the range (gui_agent 0.22.52,
gui_treedb 0.17.24).
v7.18.0 (2026-09-05)¶
open-treedb could not set the attributes a file column needs¶
import_root, files_max_size and files_content_types are SDF_RD on
C_NODE, so they can only be set at creation — and cmd_open_treedb()
builds the C_NODE’s kw itself and forwarded only initial_load. Since
open-treedb is how every real yuno opens a treedb, they were attributes
nobody could set: import-assets answered “import of files is disabled,
‘import_root’ is empty” on a node whose config named a root.
Forwarded by NAME and not by sweeping the kw, because the kw of a command
carries the caller’s own keys too (__username__, the routing metadata)
and none of them is an attribute of a treedb.
Found by migrating a real host (yunovatios) to file columns, which is
what a migration is for.
A file column, and __assets__ as a treedb system topic¶
The storage half of C_ASSETS moves into tr_treedb, where it belonged:
you mark a column ['fkey','file'], and treedb gives you a pseudo-filesystem.
The index lives in memory as the system topic __assets__ (created at open
next to __snaps__ and __graphs__, shown in system mode only), the content
on disk under <treedb dir>/.blobs/ab/cd/<sha256>.<ext>, and the column holds
an fkey into __assets__ — so an asset is linked, graphed, scope-checked and
cascade-deleted like any other node. Design and its review:
kernel/c/timeranger2/DESIGN-treedb-files.md.
The hooks of
__assets__are DERIVED, in memory, never persisted. For every columnCof topicTflaggedfile,__assets__gainsas_<T>_<C> -> {T: C}between the creation of the user topics andparse_hooks(). The host declares nothing but the column; nothing about__assets__is ever versioned.filegoes WITHfkey: every link behaviour of treedb and of the GUI keys on that word, and afilecolumn without it is refused at open.The hooks follow the schema at RUN TIME, both ways.
create-topicanddelete-topicare live commands ofC_TREEDB, so the derivation runs after each one and not only at open: it adds what the schema now asks for — into the desc and into the__assets__nodes already loaded, or_link_nodes()answers “hook field not found” — and removes what the schema stopped asking for, children included. A hook left behind holds children that are gone, and the gc reads a hook that is not empty as “some node links this asset”, so those bytes would never be collected again.__assets__is a topic of the TRANGER, so the removal takes only what maps to a topic of this treedb or to one no longer open at all.The bytes ride BESIDE the record (
__files__, a manifest keyed by column:content64from a browser, oroffset/sizeslices of the kw’s onegbufferfrom a C caller) and are consumed at the door bytreedb_store_files(), called fromtreedb_create_node(),treedb_update_node()andtreedb_autolink(). Treedb re-hashes what arrives — the client’s id is an optimisation, never an authority — checks the size on the base64 BEFORE decoding and the type ON THE BYTES (an svg calledimage/pngis refused), writes the blob, creates or refreshes the index node, and rewrites the column into the full fkey reference. A bare id of an existing asset links it; three writes, blob first.⚠️ A command that receives a
gbuffermust TAKE it out of its kw before its first exit (take_files_gbuffer()inc_node.c).expand_command()builds the command kw by copying the caller’s keys by value (json_object_update_missing,command_parser.c), so a top-level binary field is named by two kws and both areKW_DECREF’d — the command’s answer releases one,command_parserthe other, and one reference is dropped twice. And there is no second owner downstream: every exit ofmt_create_node()/mt_update_node()ends inKW_DECREF(kw), which drops the binary field it finds, sotreedb_store_files()consuming it and thatKW_DECREFare the only two releases and never both.create-node/update-nodeare the first commands in the tree to take agbuffer; every other consumer is on the event path, which does not copy.Two levels of limit: the treedb’s ceiling (
C_NODEattrsfiles_max_size,files_content_types;treedb_set_files_limits()) and the column’s policy in itsproperties(max_size,content_types), which narrows the ceiling and never raises it.gc-assetsreads the SNAPSHOTS, and that is its only guard:treedb_shoot_snap()skips every__topic, so an asset node never carries a tag. The collector walks the tagged instances of every topic with afilecolumn and keeps what any snapshotted version of a node still points at.delete-nodeon an__assets__row runs the same walk, andforcedoes not override it: the tag guard is inert for an asset (shoot_snapskips the__topics), so this is that guard in its place, andforcemeans “unlink the children”, never “ignore what a snapshot needs”.import-assets(confined toimport_root, answers the mappath -> id) andgc-assetsare now commands ofC_NODE;delete-nodeon an__assets__row removes its bytes.The FIRST arrival names the file, for ever. The extension is part of the path a web server serves and the url is cached for ever, so it cannot change under it — and one container can be declared as more than one type (the same bytes as
video/mp4and asaudio/mp4are compatible with each other and with what the bytes say). A later arrival keeps the storedcontent_type, with a warning: otherwise one asset got two blobs and the row named only one, so the other could never be served, never be seen by the gc and never be removed by the delete.C_ASSETSkeeps onlyget-asset— the way OUT: a signed url or the bytes inline, from the treedb’s.blobs/.put-asset,put-assets,list-assets,delete-asset,import-assets,gc-assetsand the store attributes are gone from it (nodes/delete-nodeon__assets__do the listing and the deleting).sha256_digest()/sha256_hex()in gobj-c helpers, standalone: the persistence layer links no TLS backend. Checked againstsha256sum.System schema 16 → 17,
colstopic 10 → 11:filejoins the enforcedflagenum.treedb_field_typesof gobj-js gets the word (7.16.4).Migration: none. The
assetstopic of a host and itsas_<col>hooks are replaced by the derived__assets__; the fkey literal in every record changes, so a store built onC_ASSETSis wiped and rebuilt throughimport-assets. The hosts (yunovatios) change['fkey']to['fkey','file']onfoto/qr/plano, drop theirassetstopic and rebuild.New suite
tests/c/tr_treedb_files(the eight claims of the design that would break quietly);tests/c/c_assetsrewritten for the round trip. Every treedb open now logs one more “Creating topic” (__assets__).
icon becomes a treedb field type, and the C side is the one that ENFORCES it¶
A col flagged icon holds the NAME of an icon of the app’s set (yi-bolt),
not a file. The nearest thing the vocabulary carried was image, which is a
different case: a frontend built an <img src="yi-bolt"> and drew the
browser’s broken-image glyph.
The word is added in three places, and the third is the one that matters:
tr_treedb.h— the documented list of field types. A comment.kernel/js/gobj-js7.16.3 —treedb_field_types, which is what makestreedb_get_field_desc()answertype: "icon". Drawn bygobj-ui7.23.53.treedb_system_schema.c— theenumof thecolstopic’sflagcolumn, which is the listcheck_desc_field()VALIDATES against.
That third one is the lesson. The vocabulary in tr_treedb.h reads like the
source of truth and is a comment; the enforced copy is a JSON literal in
treedb_system_schema.c, reached through
_treedb_create_topic_cols_desc(). A schema using a flag that is only in the
comment is REJECTED at treedb_open_db — “Wrong enum type” — and the yuno
exits at mt_play with “Parse schema fails”, which is a yuno that will not
start rather than a column that does not draw. Verified the hard way on a
live node.
The system schema therefore moves too: schema_version 15 → 16 and the
cols topic 9 → 10, so the persisted __system__ of every store learns the
new flag on the next start instead of offering the old list to the schema
editor.
An asset gets one field a person may edit¶
assets.original_name becomes writable in the canonical topic that
c_assets.h publishes for hosts to copy. It is the only column that can be:
every other one describes the BYTES — id is their sha256, content_type,
size and t measure them, source_path and uploaded_by are facts of the
load — and editing one would lie about the content. A name is a LABEL, which
is a different kind of thing.
Without it the topic has no writable column at all, so its record form opened holding nothing but the read-only id.
The JS layer¶
Carries its own versions and reaches the apps through npm.
gobj-js 7.16.4 → 7.16.5 and gobj-ui 7.23.60 → 7.23.62: the file column control¶
A col flagged ['fkey','file'] can be filled by a person now. Until this it
was written by a C caller or by import-assets and by nobody else.
gobj-js 7.16.4:
filejoinstreedb_field_types, which is what makestreedb_get_field_desc()see it.gobj-js 7.16.5: the ORDER of the flags decided whether a column was one. That function walks the flag array and every type word OVERWRITES the answer, so
['fkey','file']answeredfileand['file','fkey']answeredfkey— the same column drawn as a picker or as a select depending on how its author wrote the list. The C asks for both words withkw_has_word()and does not care about the order; neither does this now.gobj-ui 7.23.60: the control (
yui_file_field.js). Three things it does that are not obvious, each one a bug the shape avoids: the file is read at SAVE and not at pick (aFileis a reference, so cancelling reads nothing); theFilenever enters a kw (plain json, and the trace serialises it), so the host asks the form for it withget_picked_files; and reading is a promise, so it enters the machine asEV_FILES_READ/EV_FILES_FAILED.gobj-ui 7.23.61: the form logged “type unknown: file” twice per record opened — its two value converters did not know the word.
gobj-ui 7.23.62: a read-only form reported the column as WRITABLE. The control hid its button, which hides the way in for a person and leaves the
<input type=file>enabled.
The last two were found by driving a real form against a real census, not by a fixture: the control drew correctly and the log said it did not.
gobj-js 7.16.3 and gobj-ui 7.23.50 → 7.23.59¶
The treedb topic table, one round, in the order the defects were found:
iconas a field type (7.16.3 / 7.23.53), the C side of which is above. A name the app’s set does not carry falls back to the TEXT of it, becauseyui_icons.csspaints acurrentColorbox for ANYyi-name and drawing an undefined one renders a solid black square.A record is READ, not just written (7.23.54–7.23.56). A click on a row opens it outside edition mode — it was the biggest target on screen and the only thing that answered nothing — and the form shows EVERY field, the non-writable ones read-only. What is shown is not what is sent: only writable cols, fkeys and the pkey travel back, because
treedb_update_node()does not checkwritableandtis an integer rendered as a date. Two controls had to learn it:readonlyis an attribute of a text control and of nothing else, so a<select>, a checkbox and a colour swatch accept it and ignore it.The paginator (7.23.50–7.23.52). Picking “All” no longer takes the size selector away with it; the size is remembered per topic; and remote paging is out of use, because a partial topic BREAKS LINKING — a record’s link picker is built from the rows the parent topic’s table holds, which under paging is one page. Local pagination is untouched:
getData()answers the whole dataset whatever page is on screen, so paginating the DISPLAY is safe and only paginating the FETCH is not.A colour shows its value beside the swatch (7.23.58), and
savemoves to the RIGHT end of the form toolbar (7.23.57, 7.23.59) — the form opens as a modal dialog now, and a dialog’s primary action goes bottom-right. The first attempt moved nothing: thetoolbarattr carried the five names written out again, soplan_toolbar()never saw theundefinedthat would have usedDEFAULT_TOOLBAR.
gobj-ui 7.23.49: a search that could not see inside an fkey¶
A treedb row is not flat. With list_dict an fkey arrives as a LIST OF
OBJECTS — [{id, topic_name, hook_name}] — and the topic table’s search box
stringified every value with String(val), which for that is
"[object Object]". So searching a topic of devices for the place they
sit in — the value an operator actually has in mind — never matched anything,
while the cell plainly rendered that id: what you see and what is searched
were not the same thing.
yui_row_search.js walks into lists and objects and reads only the id of
an fkey: topic_name and hook_name are the same two words on every row, so
matching them would turn any such term into a wildcard over the whole topic.
Two more in the same table. The page-size selector offers All on a
remotely paged topic — filterMode: "local" means the filters only see the
page that is loaded, and nodes has taken limit: 0 for “every one” since
paging landed, answering the plain list. And each header filter carries a
✕: a column filter was undone by deleting what you typed, and with several
of them set, getting back to the whole table was an exercise in remembering
which ones you had touched.
gobj-ui test-app: the maplibre worker asset carries its version¶
The worker and its shared chunk were the only assets of the bundle emitted
under a FIXED name, so a static host serving /assets/ with a long max-age
hands a returning browser the OLD worker against the new bundle — and worker
and main thread speak a private protocol that changes between maplibre
versions. Paid for in yunovatios, where a cached 6.4.1 worker made the 6.7.0
GlyphManager answer t.codePointAt is not a function once per tile. The
emitted names now carry the installed version.
yunos-js 0.22.46 / 0.17.18¶
The agent console stops its JSON viewer before destroying it (gui_agent
0.22.45), and both SPAs then raise their ranges to the libraries above —
@yuneta/gobj-ui ^7.23.49 and @yuneta/gobj-js ^7.16.2 — which is what
carries the fkey search fix into gui_treedb’s table. Ranges only: no code
of either app changes. Detail in that repo’s own CHANGELOG.
7.17.3¶
An absent DTP_JSON is json_null(), and four predicates disagreed about it¶
set_default() materialises a DTP_JSON with an empty default as
json_null() — a valid pointer — so if(!jn) over one of them is a
branch that can never be taken. It reads exactly like the guard everybody
writes, it compiles, and it survives every test that exercises the
present-value path. It reaches further than attributes: a subscription is
built with gobj_sdata_create() too, and its optional __config__,
__global__, __local__ and __filter__ are declared that way.
The tree already had the answer written four times, and the four disagreed:
| C NULL | json_null | {} [] | "" | scalar | |
|---|---|---|---|---|---|
empty_json() — helpers.h, public inline | FALSE | TRUE | TRUE | FALSE | FALSE |
json_size()==0 — helpers.c, public | TRUE | TRUE | TRUE | TRUE | TRUE |
json_empty() — tr_treedb.c, private | TRUE | TRUE | TRUE | TRUE | TRUE |
is_unset_value() — c_treedb.c, private | TRUE | TRUE | TRUE | TRUE | per type |
The one in the public header was the wrong one, and wrong on precisely
the case being chased: jansson’s json_is_array()/json_is_object()/
json_is_null() are each NULL-safe on their own, so all three fell through
and empty_json(NULL) answered “not empty”.
Consolidated, not extended. json_size() moves to static inline in
helpers.h (and takes const json_t *); empty_json() is now
json_size(jn)==0, so there is one switch and the two cannot drift again;
the private json_empty() of tr_treedb.c is gone. Alongside it,
json_absent(jn) — true for C NULL and json_null and nothing else — for
the seven places that already wrote (!x || json_is_null(x)) by hand.
The rule, in CLAUDE.md and GOBJ.md §8.15:
never if(!jn) over a json. empty_json() almost always — an empty
filter is as much “no filter” as an absent one, which a null test does not
cover — json_absent() where null and empty differ, and the shape itself
when the shape is the contract.
The guards that were dead, and the double free one of them was hiding¶
🔴 ycli shortkeys. mt_create read the shortkeys attr and created the
dict if(!priv->jn_shortkeys) — never — so on a config that had never saved
one the dict stayed a json null and add-shortkey wrote into it silently.
Fixing only that guard would have broken ycli at exit: mt_destroy
carried a JSON_DECREF of a reference gobj_read_json_attr() hands out
BORROWED, harmless only while the value was the json_null() singleton
(refcount (size_t)-1). With the dict actually created it is a double free.
Both go in the same change; the three cmd_*_shortkey guards now ask for the
shape.
c_tcp_s / c_udp_s cert reload. crypto is DTP_JSON with a null
default, so reload-certs’s “‘crypto’ attribute is empty” could not be
reached and the reload went on to hand a json null to ytls.
C_IEVENT_CLI resent subscriptions without their filter. Two bugs in the
same four lines. The three optional keys were read with KW_REQUIRED, which
logged an error each per subscription — the flag was simply wrong on keys
that are optional by contract. And __filter__ is any json — the callers
pass a list of alternatives — while it was read with kw_get_dict(), the
dict-only reader, which answers NULL for a list: on reconnect the
subscription was resent without its filter and the remote published
everything.
crypto is a dict, so it says DTP_DICT now¶
The guard was necessary because the type lied. DTP_JSON means any
json, null included, and json2item() proves the difference: it refuses a
non-object for a DTP_DICT with a log, while for a DTP_JSON it accepts
whatever arrives. crypto was declared DTP_JSON in seven gclasses
(c_tcp, c_tcp_s, c_udp_s, c_prot_http_cl, c_auth_bff,
c_task_authenticate, c_smtp_session) and is a dict in all seven — and
they hand it down to one another, so a non-dict accepted at the top reached
ytls at the bottom. It fails at the first hop now.
Safe on a node with a persisted value, checked rather than assumed: neither
c_tcp_s nor c_udp_s ever calls gobj_load/save_persistent_attrs (and
they are children, not services, so they could not), nothing in the tree
writes the attr with gobj_write_json_attr(), and there is no
"crypto": null in any config or store. A stored one would fall back to
{} with a log — which is what it effectively was.
With the type honest, the two json_object_set_new(jn_crypto, "trace_tls", …) at listen time can no longer fail — and they say so if they ever do,
along with the ssl_server_name one in c_tcp. All three used to return
-1 into nothing.
gobj-js 7.16.2: the same question, spelled the same¶
empty_json() lands in gobj-js/src/helpers.js with an identical truth
table (verified side by side, both runtimes, eight values). Javascript has no
second null, so if(!x) happens to work there — and that is the divergence a
port crosses in both directions. json_absent()'s twin is the is_null()
that was already there.
7.17.2¶
C_ASSETS: an audit for silent errors, and gc-assets could empty the store¶
Asked to sweep my own code for failure paths that write nothing to the log. Eight, and two of them mattered.
🔴 gc-assets deleted everything it could not judge. node_is_linked()
answered FALSE for a node it had no hooks to look at — and
asset_hook_names() answers an empty list whenever the topic descriptor
cannot be read. So a treedb that would not answer meant every asset looked
like an orphan, and the garbage collector removed the whole store, rows and
blobs, reporting success. The comment right above that function already said
“not being able to prove an asset is an orphan is a reason to keep it” — the
empty-list case walked straight past it. Cannot tell now means linked, and
gc-assets refuses outright when it cannot read the hooks, because a
collector that cannot tell must not guess.
An asset node with no id was skipped silently by both gc-assets and the
census index — one leaves an unreachable blob behind, the other makes every
image that node holds unreachable. Both say so now.
Three refusals of store_asset — empty, over max_size, content_type not
allowed — put the reason in the answer and nothing in the log. The answer
reaches whoever called; the log is what somebody reads when twelve thousand
images went by and a few did not arrive.
And yui_asset_element()'s onerror drew the marker and told nobody. It
reports through log_error now: a failure only a user can see is a failure
nobody measures.
C_ASSETS: source_path is a LIST, and 147 photographs say why¶
An asset is its content, so N files with identical bytes are ONE asset —
and each of those files came from its own path, which is what a loader links
by. source_path held one, so it held the last one written and lost the
rest.
It is not a corner case. In the yunovatios census 148 measurement points share one byte-identical photograph: content-addressed that is a single asset, so 147 of them named a path no asset carried, were never even looked up, and came out with no image. The load reported success, and it was found counting rows afterwards.
source_path is an array now and store_asset accumulates every path
that ever led to that content. A store written before this still reads: a
bare string is taken as a list of one.
And a re-upload of a path already known now writes nothing at all — before, every re-run of an idempotent load appended a row per asset to an append-only store. Proven on the real case: the shared asset came back carrying 115 paths with 115 devices on its hook, where it used to carry one of each.
The reason this release exists at all¶
Both fixes above were in main and reachable by nobody. A node that carries
a runtime-only SDK — outputs/, outputs_ext/, tools/, no kernel/ and no
.git — cannot rebuild the framework, so a fix reaches it as a package or
it does not reach it. The census on the central node stored its twelve
thousand assets and linked none of them, because its store_asset was the
one that writes source_path as a string.
Three releases in a row have been cut for this reason. It is worth saying plainly: for those nodes a commit is not a fix; a tag is.
kernel/js/gobj-ui moves to 7.23.48, which carries yui_asset.js — the
browser half of C_ASSETS: it resolves a treedb reference to an asset id
through all three shapes a schema can hand it (bare id, topic^id^hook, and
a list of either), asks the service for the mode it prefers, and draws a
marker when an image does not arrive and logs it.
7.17.1¶
C_ASSETS takes video and audio too¶
A node of a treedb owns more than photographs. allowed_content_types now
carries video (mp4, webm, quicktime, ogg, x-matroska) and audio
(mpeg, mp4, ogg, wav, webm, flac) alongside the images and the
pdf, and both mime tables know them.
The pairs that share a container are split by extension, deliberately:
the extension is the only thing a web server reads to pick a
Content-Type, so .webm is video and .weba audio, .mp4 video and
.m4a audio, .ogv video and .ogg audio. Getting that wrong does not
fail — it serves a sound file labelled as a film.
max_size moves from 32M to 128M, and it is a memory limit as much as a
policy one. There is no streaming path: an asset is hashed and written
whole. put-asset costs the worst, because the base64 arrives inside the kw
and is then decoded — one call peaks at roughly 2.3x the file — while
import-assets only pays the file itself. Raising it further for big media
means checking the yuno’s own MEM_MAX_BLOCK first: a single base64 string
above that is refused by the allocator, not by this gclass. Said out loud in
the attribute description, in c_assets.h and on the doc page, because a
limit that silently governs RAM is the kind that is found the hard way.
image/svg+xml stays out of the default, for the reason it always was: an
svg served from the app’s own origin runs script.
C_ASSETS: put-assets, because a batch line carries ONE file¶
import-assets reads a directory that is already on the node, which is
right for a migration and wrong for a norm: the initial data of a yuno
belongs in yunos/batches/, on the authoring side, and the deploy sends it
from there. A batch line carries ONE file (content64=$$(...)), so a census
of twelve thousand images is either twelve thousand commands or a few dozen
bundles.
put-assets takes a bundle: a JSON array of
{original_name, source_path, content_type, content64} — the same form the
rest of the batches use, so it needs no parser of its own and a person can
read it. The bundle is TRANSPORT, not storage: the images stay files where
they are authored, and the deploy packs them the same way $$() base64s
them. Nobody commits base64.
One bad entry does not stop the bundle. A load of that size that aborted
halfway would be neither retryable nor comparable, so every failure is
logged with its name and the answer reports stored and failed.
C_ASSETS: two defects a live run found and no unit test could¶
Both were found driving the gclass through a real agent, and neither fails loudly — which is the point.
get-asset id= answered “Yuno not found”. command-yuno hands its
WHOLE kw to gobj_list_nodes() as the filter that picks the yuno, so a
parameter named like a field of the yuno record becomes a filter on that
field: an id of a sha256 matches no yuno. The error names the YUNO and
never the parameter, so it reads as the service being missing. The
parameter is asset_id now, with the bare id kept as a fallback for a
caller that never crosses the agent. CLAUDE.md warns about exactly this
and calls id “the best-known case”; the warning was there and the
parameter was named id anyway.
orphan=1 listed everything. command-yuno does not coerce, so a
boolean arrives as the STRING it was typed as, and kw_get_bool() answers
the default for a string unless it is given KW_WILD_NUMBER. orphan was
read without it, so the filter did nothing and list-assets orphan=1
returned the whole store — which reads as “nothing is an orphan”, the
opposite of the truth, and would have sent somebody hunting for a bug in the
hooks. orphan, force and both dry_runs take KW_WILD_NUMBER now, and
the test drives them as strings the way the agent does.
7.17.0¶
C_ASSETS: the bytes a treedb node owns but cannot hold¶
A treedb node often owns something that is not json — a photo, a plan, a
signed pdf — and today those bytes live wherever whoever loaded them put
them. In yunovatios that is 356 MB of census images under
/yuneta/store/resources/censo_memorias/, referenced from devices.foto
as a free string, pushed to the node by an rsync from a developer’s
laptop. Nothing owns them: no command writes one, so a user cannot add or
replace a photo from the SPA or from ycommand; nothing checks the path
resolves, so a renamed topic breaks every image in silence; and the web
server serves the whole prefix to anyone who reaches the vhost, which
walks straight past the scope the treedb enforces on the nodes themselves.
They cannot go IN the treedb: it is held in memory and timeranger2
rewrites the whole record on every update, so a 40 KB photo would be
rewritten every time its node changed state, and would ride along in every
page of nodes. (The treedb blob column type is not binary — it is
free-form json.)
C_ASSETS does what the agent already does with yuno binaries: the bytes
in a directory the service owns, one node per asset in the treedb, and
commands as the only way in. The asset id is the sha256 of the content,
so the same bytes stored twice are one asset, a whole census reload creates
nothing, a served url can be cached for ever, and replacing an asset is a
new id plus a relinked node — which is what makes the node’s own history
say which photo it carried, and when.
The consumer’s column stops being a path and becomes an fkey to the
asset topic, so an asset is linked, listed, graphed, scope-checked and
cascade-deleted like any other node, and an asset no node links any more is
visibly garbage (list-assets orphan=1, gc-assets).
Commands: put-asset, get-asset, list-assets, delete-asset,
import-assets, gc-assets. Writes are refused on a replica and gated by
the write / read authz of the service.
get-asset answers in one of two shapes, and the SERVICE decides which:
a signed url when public_url and sign_secret are configured, the bytes
inline when they are not. So the caller has one code path, and a node with
no web server in front of it still shows its images instead of showing
nothing. The signed form reproduces, byte for byte, what nginx’s
secure_link_md5 "$secure_link_expires$uri <secret>" hashes — verified
against openssl on four cases. The client address is deliberately not in
the signature: it would tie the url to one ip and break every phone that
changes network mid-session; the 15 minute default lifetime is what limits
a leaked url.
import-assets is the bulk path, and it moves no bytes at all. One
command walks a directory that is already on the node and turns it into N
assets — for the yunovatios census, 12 281 files that are already sitting
in the store. Sending hundreds of megabytes through the control plane, one
base64 message per file, is the thing it exists to avoid. It reads an
arbitrary path, so it is confined to a configured import_root and refused
outright when there is none.
Three separate things do the confining, and a security review of the
pushed commit was right to look even though it holds: only ONE of them is
visible at the call site. The explicit .. guard refuses rather than
silently resolving somewhere else; build_path() strips the leading / of
every segment after the first and clamps .. against it, so an ABSOLUTE
source_dir lands INSIDE the root (/etc → <import_root>/etc), not at
/etc; and walk_dir_tree() lstat()s, so a symlink is neither a regular
file nor a directory and cannot lead the walk out. The test drives all three
with hostile input rather than asserting it from reading the code, and the
comment at the call site names the other two so neither gets “simplified”
away on the grounds that the other covers it.
The topic belongs to the HOST, not to this gclass: an asset’s fkeys point
at the host’s own topics, so only the host can write those hooks.
C_ASSETS never creates it and refuses to work when the topic it was
pointed at cannot hold what it is about to write — a blob on disk whose row
failed to be written is a file nothing can ever find again. The canonical
topic and the nginx block are in c_assets.h.
image/svg+xml is not in the default allowed_content_types on purpose: an
svg served from the app’s own origin runs script.
tests/c/c_assets covers the lot, and two of its checks exist because the
bug was written first and both were silent. gobj_topic_desc() answers
{topic_name, pkey, ..., cols} while topic_desc_hook_names() walks a LIST
OF COLS — handing it the dict does not fail, json_array_foreach() over an
object iterates nothing, so it answers “no hooks” and every asset looks like
an orphan. And with the hook_size option a hook is not a number: it
renders as [{"size": N}], so an EMPTY hook is a NON-EMPTY list and “the
list has elements” is the wrong test — it made every asset look linked and
gc-assets delete nothing. node_is_linked() now reads the shape it asked
for, and counts anything else as linked: that answer decides what a garbage
collector removes, and not being able to prove an asset is an orphan is a
reason to keep it.
public_url and sign_secret are SDF_WR and repeated in mt_writing.
They were SDF_RD with a cached const char *, which is a pointer INTO the
attribute — writing one at runtime left the copy dangling. It also means a
node can be switched between the two serving modes without a restart.
nginx and openresty are built with secure_link¶
--with-http_secure_link_module added to both configure-libs.sh
configure blocks. It is what lets a web server check C_ASSETS’s signed
urls by itself, with no round trip to the yuno. Core module, no new
dependency, no path/name/symbol change — but it does mean the nginx that
ships in the .deb/.rpm has to be the rebuilt one before a node can
serve assets that way. Until then those nodes answer inline, which is the
fallback get-asset was designed around.
C_AUTHZ says WHICH authz db it did not find¶
“No authz db, authz only to local access” carried the path it looked for as
"path", "%d", path — a const char * formatted as an int, so the one field
that names the missing directory printed the pointer ("path": 676502376).
It is the field the message exists for: a follower (master=false) that
loses the start-up race against the master that creates the store, and a
deployment that genuinely has no authz db, produce the identical line. The
format is %s now, and the line also carries master, which is what tells
the two apart.
Found on a from-scratch yunovatios install: db_tracks_ce and db_tracks_co
checked 24 ms and 22 ms before their store existed, came up with authz
disabled for the life of the process, and answered list-users with nothing
while the master listed three accounts.
initial_load: a link a seed declares is as immutable as the seed¶
The immutable mark of 7.16.4 protected the record: delete-node refused a
seed, force included. Its links were not covered — the mark is one md2 bit,
and tr_treedb does not know which links matter — so unlink-nodes, an
update-node with autolink that did not repeat the fkeys (what kw omits,
treedb_clean_node() drops), a force delete of the parent, or a
link-nodes into a single-valued fkey could still leave the seed hanging off
nothing until the next start re-linked it. A scope in yunovatios is exactly
such a link. C_NODE, the owner of initial_load, now refuses those four
writes when they would cut a link a seed is declared with (“initial_load:
cannot unlink a seed link”, “... update would drop a seed link”, “...
cannot delete the parent of a seed link”, “... link would overwrite a seed
link”); force overrides none. No column flag was added: a flag would
freeze the column for every record of the topic, and the declaration already
says which links matter. The links a person adds to a seed afterwards stay
ordinary.
The fourth one is the least obvious, and it is why the guard cannot be read
off the other three: a link does not always add. _link_nodes() branches
on the shape of the child’s fkey column — a list takes the new ref beside the
ones already there, an object keys it, but a string column has room for
one and is written over without a comparison. So link-nodes to another
parent through a single-valued fkey cuts the declared link as surely as an
unlink does, and a test whose fkeys are all lists cannot see it.
apply_initial_load() now runs in two passes, the way treedb_open_db()
loads a store: every record first (created without its fkey values), every
link second. A child declared before its parent used to be created with its
link failing (“parent node not found”) and healed on the second start; now
the seed comes up whole on the first, whatever the topic order. New test
tests/c/c_node_initial_load, which carries both fkey shapes on purpose:
a list (users.departments) and a string (machines.department).
YUNO_TREEDB.md §3.10 carries the contract.
7.16.5¶
The memory audit follows its own switch¶
CONFIG_DEBUG_TRACK_MEMORY is the Kconfig knob that “enables track memory to
find leaks”, and every guard in gbmem.c also demanded
CONFIG_BUILD_TYPE_DEBUG. A RelWithDebInfo build with tracking enabled
therefore tracked nothing: get_cur_system_memory() answered 0 for the life
of the process, "cur_system_memory": 0 on every log line, and the
shutdown audit (print_track_mem()) was not even compiled in -- which also
made the get_cur_system_memory()==0 checks of every ctest pass on any
leak. The second half of the guard is gone; tracking now follows the one
switch the menu shows. Proven with a probe that leaks a block on purpose:
silent before, “system memory not free” + the block after.
7.16.4¶
initial_load: a treedb declares what it cannot come up without¶
The records a system cannot start without -- the seed role, the admin account, the root of the tree a scope hangs from -- were written by a batch that somebody had to remember to run. A batch writes them once. The day one is deleted the system is up, answering, and showing nothing, and the deletion raises no error because a missing record is not an error.
C_NODE gains an initial_load attr: one entry per topic, a list of
records, with the links riding inside each record as fkey values
(parent_topic^parent_id^hook, the same form the child stores). It is applied
in mt_start right after the treedb opens, master only, on every start:
a record that is missing is created and autolinked from its own fkeys;
a record that is present is never rewritten -- only its declared links are checked, and a missing one is re-linked;
either way the record is marked immutable, so
delete-noderefuses it andforcedoes not override.
The re-link is the half that is easy to miss. Deletion is not the only way to
blind a seed: a scope is a link, and treedb_unlink_nodes() carries no
immutable guard, so an untouchable record can still be left hanging off
nothing. Re-linking on start repairs that; immutability is the defence between
restarts. And it re-links without ever re-writing, because an autolink over an
existing node runs treedb_clean_node() first, which would drop every link the
seed does not declare -- the ones a person added on purpose.
Reachable two ways: the new initial_load parameter of C_TREEDBS’
open-treedb, or the attr set directly by a gclass that builds its own
C_NODE.
C_AUTHZ is now one caller of it. The same loop lived in its mt_start,
written for one treedb; it hands Authz.initial_load down to its C_NODE
child instead, and the loop is gone from c_authz.c. The attr keeps its name,
its shape and its behaviour -- command_delete_user, which seeds an immutable
user through it and checks that the delete is refused, passes unchanged. The
explicit time stamp the old loop injected went with it: a col flagged time
with no value supplied already gets the current time from treedb itself, so it
was always redundant.
7.16.3¶
Cut so that the webstats change reaches the yunovatios nodes: they carry a
sparse SDK, so an SDK yuno cannot be built there and only arrives in the
.deb / .rpm.
C moves in four places -- the flat json (json2flat / flat2json),
c_tranger’s delete-key, diff-schema answering differences nobody had
made, and webstats saying when a log stopped rotating -- plus the three bench
tests that were double frees in the tests themselves. The rest is the JS layer
(gobj-js 7.13.9 -> 7.16.1, gobj-ui 7.23.16 -> 7.23.45,
yunos-js 0.15.1 -> 0.22.44), whose detail lives in each repo’s own CHANGELOG.
webstats says when a log stopped rotating¶
A rotation is two steps -- logrotate renames the file, the web server reopens
-- and when the second one is lost the server keeps writing down the old
descriptor: <path>.1 grows while <path> stays as logrotate created it.
Nothing reported it. The server serves, logrotate exits 0, and the daily report
was still built, because this yuno reads both files.
And it does not heal. Once <path> is empty, notifempty skips it, so the
rotation is never attempted again: one lost signal costs the rest of the life
of the node. It cost eight days on a real node, and it was found by
somebody listing the directory.
The signature needs no history: after a healthy rotation the live file is the
one being written, so its mtime runs ahead of the .1; the other way round
means the reopen was lost. New attr rotation_stall_minutes (default 120,
0 disables) is the margin that keeps a site with no traffic since the
rotation from reading as broken -- there neither file moves, and neither is
meaningfully newer.
Reported in three places, because they have different readers:
a
WARNINGin the log at the moment it is seen -- the node is watched from there, and the mail is once a day;stalled_rotationsin the stored record, and a line in Needs attention of the mail;list-sources, which now answers error and names both files and the gap, so the command an operator already runs to check the yuno checks the rotation too.
It does not fix it: signalling a web server is not this yuno’s business, and a report generator that restarts services is a different and worse thing.
The three broken tests of the bench, and why they broke on one particular day¶
test_tr_treedb, yev_events/test_yevent_traffic5 and traffic6 were aborting
with heap corruption -- “corrupted double-linked list”, “free(): chunks in
smallbin corrupted”. All three were reference accounting in the tests
themselves, not in the library.
And they have a date of origin: on 2025-07-13 test_json() started FREEING
its argument (e6b480f83). From that day, every place that handed it a
borrowed pointer was broken. The offending calls are from 2024-11-02 and
2024-11-08: they were correct when they were written.
test_tr_treedb.c:json_array_get()returns a borrowed element andtest_json()owns what it receives, so the element lost a reference it never gave -- and thejson_decref(data)on the next line freed it a second time.test_users.c:kw_get_dict()returns borrowed. TheJSON_INCREFpaid for the call toload_treedbs(), which does own; theJSON_DECREFafter it spent the parent’s reference, and freed the child under it.traffic5/traffic6: a doublejson_decref(msg)on the error branches. The test corrupted itself exactly when it detected a mismatch, so it died instead of reporting one -- and the mismatch is deliberate: it closes the socket on purpose and the error messages are in its expected list.
A real defect of the SDK, found along the way: json_check_refcounts()
tested if(!jn), logged it and carried on to json_typeof(jn): it blew up
on the very NULL it had just reported. A checker that dies of what it exists to
detect. It returns now.
How it was found, which is what matters next time. ASan saw nothing: the
test tree links the installed libraries from outputs/lib, not the ones in
the build tree, so instrumenting the build does not instrument what runs. It
took building the SDK and jansson with -fsanitize=address -- jansson with
its own generated config headers, or it produces invalid json -- and
relinking the test by hand against those libraries. With that, ASan pointed
at the three exact lines in a minute.
Two fixes of the JS runtime that reach EVERY SPA¶
Both came out of using the consoles against real nodes, and both were in the library, not in the applications.
gobj-js 7.16.1 -- an event addressed to a service that no longer exists is
dropped, not broadcast. C_IEVENT_CLI looked for the destination service
and, on not finding it, fell through to the default delivery: publish it to
every subscriber of the transport. That path is the right one for a message
that names no destination and the wrong one for a message that does -- an
addressed message belongs to its addressee or to nobody.
What triggers it is the normal end of a view’s life: it is mounted under a
service name, it subscribes to a backend event, the user navigates elsewhere,
the view is destroyed, and the frames already on the wire keep arriving for a
name nobody answers to. They ended up on the application gobj’s null event
subscription -- which is null on purpose: naming EV_ON_OPEN in a subscription
forwards it upstream and the remote rejects it -- and its FSM does not declare
an equipment frame, so it said so once per frame: 38 errors in a 26 ms
burst on a node with 38 devices. Now it is dropped with a warning naming
which service and which event were lost. The path with no destination is not
touched, and tests/ievent_dispatch.test.js pins the three cases together,
because the fix is only correct if the third one still works.
⚠️ The C side carries the same open TODO (c_ievent_cli.c) and was left as
it was: it is consolidated kernel, and how a backend routes is not a decision
of the JS side. Until it moves, a C client and a JS client do different things
with an orphaned addressed event.
gobj-ui 7.23.45 -- the wheel over a graph popover, and a card that can be
read and closed. No graph popover could be scrolled with the wheel: you had
to drag the bar. G6’s zoom behavior binds a wheel to the container and
calls preventDefault() on every notch whatever the target is, and then
declines to zoom because the target is not the canvas -- so the gesture was
cancelled and used by nobody. It is stopped at the popover, which fixes all
five. And a node’s detail card came at caption size with a 13x16 px ✕: the
sizes moved out of the inline styles into the CSS -- a media query cannot
reach an inline style -- and the header is sticky, because with it scrolling
away the ✕ disappeared exactly on a phone.
The JS versions say again which SDK they are built against¶
gobj-js 7.13.9 -> 7.16.1, published. From now on a JS package of the SDK
does not run ahead of the C one except in the third index: the first two
are those of YUNETA_VERSION and the third is the package’s own life between
releases. 7.14 and 7.15 are skipped on purpose -- the number does not count
releases of the package, it says which SDK it belongs to.
gobj-ui breaks the rule and it cannot be fixed by renumbering: it is at
7.23.45 with the C at 7.16.2, and publishing a 7.16.x behind it would be a
version lower than the published one, so npm would keep 7.23.45 as latest.
The rule is forwards; it is written down in CLAUDE.md.
The flat json: json2flat / flat2json, in C and in JS¶
A json seen as a table: one row per LEAF, the id is the path of the item and the value its value. To store, to compare and to diff it is a far better form than the native one -- and it is the only one a person can read when two configurations disagree.
{"a": {"b": 1}, "c": [10, 20]} -> {"a`b": 1, "c`[0]": 10, "c`[1]": 20}Half the piece already existed and it was broken. json_flatten_dict() /
json_unflatten_dict() had been in kwid.c for a long time, with a warning in
the header: “digit-only keys are reserved for array indexes”. Measured before
touching anything:
| input | what it returned |
|---|---|
{"1630": {...}} -- a yuno id | an array of 1631 elements |
{"a": {}} / {"a": []} | the key disappeared |
{"a`b": 1} | it was split into {"a": {"b": 1}} |
{"a": {"": 1}} | it came back {"a": 1}, one level short |
"a`1000" | 1001 elements materialised for a single value |
The rule the format imposed was broken by our own data: a dictionary indexed by yuno id is as ordinary as it gets here.
The new grammar, the same in both languages:
separator
`, which is already the path delimiter ofkw_get_dict();a
`inside a key is doubled -- and with that the format forbids nothing;an array index is
[N], canonical and with no leading zeros. It costs one byte more than a bare number and it gives the rule of keys back;a key starting with
[doubles the bracket ([[0]), so it can never be read as an index;an empty container is a leaf:
{}and[]have no leaves of their own, and"properties": {}is everywhere in our configs and schemas.
flat2json() refuses instead of guessing: an id that is a leaf and a
container at once (the result would depend on the order the keys are read
in), an index above the cap, a path deeper than the cap -- the old code
truncated to 256 segments in silence.
Index and key are two TYPES, not two spellings: flat_key_split() returns
the index as an integer and the key as a string. Its own test exposed it: as
strings, the key "[0]" and the index 0 both came back as "[0]", which is
exactly the ambiguity [N] is there to remove.
Also flat_diff() -- {added, removed, changed} -- and flat_apply(), which
works on the flat form on purpose: there an id addresses one value, so
applying is putting and removing, with nothing to walk and nothing to guess.
The two old names remain, delegating to the new ones and marked deprecated:
there is no stored data in the old format to be compatible with -- the only
production caller, flatten-subscribers of the MQTT broker, prints it on a
console.
New tests: tests/c/kw/test_json_flat.c (33 checks) and
kernel/js/gobj-js/tests/json_flat.test.js (36). The ids they pin are the
same in both: a flat json is written by one side and read by the other.
Mostly the JS layer (yunos-js 0.15.1 -> 0.22.44, gobj-ui 7.23.45,
gobj-js 7.16.1) in eight parts, plus two things the C side was missing: a
tranger topic could be listed key by key and never PRUNED, and diff-schema
answered dozens of differences nobody had made.
Two things learn to act on a SET: the connections table of gui_treedb,
because a pasted deploy centre is two hundred rows, and the treedb GRAPH, which
could be rearranged one node at a time and no other way. Its camera toolbar
also stops promising what it does not do.
And the URL learns to hold a POSITION. Three separate places threw one away: a
reload on a deep tab route answered with another tab’s default, an action route
came back to the route a view is declared at rather than the one you were
looking at, and both consoles were deciding this in their own c_app.js --
which is how one of them could be wrong while the other was right, and nothing
said so. The deciding is one shared piece now.
And the JSON viewer learns to READ the same document three ways. A tree is the
right shape for finding one value in a large document, and the wrong one for
reading it as it is written or for seeing its shape — so C_YUI_JSON now also
shows the raw text of what it holds, and a graph of it.
And last, the three G6 graphs learn to be OPERATED by a finger. They have always drawn correctly on a telephone; what could not be done on one was anything else — the zoom was two buttons because G6 gives it to the wheel, the context menu had no door at all, the controls a finger has to land on were sized for a pixel, and the treedb graph’s multi-selection hung off a key a telephone does not have. And the plainest thing of all could not be done either, which is why it was found last: a finger could not MOVE a node.
And the sixth is what a developer READS. Two rounds fell out of looking at the gobj tree: the trace of the FSM is the framework’s whole debugging story, and in the browser it was hard to read for three separate reasons. It arrived in the legacy shape (three lines per transition where a node writes one, because the JS port’s default never matched the kernel’s, and then because four of its six trace sites had no other shape to write). It lost its INDENTATION on the way to the screen, so it read as a flat column where the console read as a tree. And it could not be quietened: the filter that promises to hide recurring events could not see the machine trace at all, so a yuno with one timer buried everything else under two lines a second — in the window, and then, once the window was fixed, in the console beside it.
And the seventh is the gclass itself. view-gclass answers a complete
description -- attrs, commands, methods, trace levels, FSM -- and the only
thing that drew it was a JSON tree: correct, and unreadable. It is laid out by
ZONES now, with its machine as a matrix or a graph, and reading a real gclass
with it found two defects in the drawing and one in the framework’s own
teardown.
Detail in those repos’ CHANGELOGs.
Added¶
The gobj tree says what a gobj IS and what it is DOING (
gobj-ui7.23.30). The card carried a gclass, a name and a coloured dot, so learning which node was a service, which had been disabled and which was running without playing took opening the popover of every one of them. It carries a status SYMBOL now (▶playing,‖paused,■stopped,⊘disabled -- a dot has one shape, so telling running from stopped meant telling green from red at 10 pixels), badges for the role and fordisabled/bottom/commands, the FSM state on its own line, and a dashed dimmed border for a gobj that is out of the game.runningandplayinghad the SAME dot and are not the same thing: a gobj that runs and does not play is PAUSED. The palette went with it -- five saturated hues, one family per role; the fill had been a 9% tint, at which every card was the same near-white rectangle and the tree read as a wireframe.And the structural half, which is what this view is FOR next: it draws DESCRIPTORS now, not gobjs, whose field names are the ones the kernel’s
gobj2json()writes. Two producers -- one for the browser yuno, one for the answer of theview-gobj-treecommand -- so showing a backend yuno is one line, and nothing below that line changes. The remote fetch is not wired yet.A gclass viewer (
gobj-ui7.23.30, rewritten by ZONES in 7.23.36), which the framework did not have.view-gclassanswers a full description of a gclass -- attrs, commands, methods, trace levels, FSM -- and the only thing that ever read it was a terminal. The gobj tree’s popover carries agclassbutton that opens that same document in aC_YUI_GCLASSwindow: zones for what a gclass HAS, and the raw answer one button away, because the backend’s own answer is the authoritative one.The machine is a matrix, rows = events and columns = states -- the shape the FSM is declared in, and the one that survives 12 events against 3 states as well as 86 commands against one. Two things it says that a JSON dump cannot: an empty cell is INFORMATION (that event, in that state, is refused with “Event NOT DEFINED in state”, so the empty cells are the map of what breaks, and they are hatched), and a state nothing declares a way INTO is MARKED -- an action may jump with
gobj_change_state(), asC_IEVENT_CLIdoes intoST_SESSION, and no description can see inside an action, so the alternative was drawing the working half of a gclass as unreachable. A second view draws the same machine as a G6 graph, one edge per PAIR of states and no self-loops.It reads BOTH dialects of the document: the C kernel and the browser registry answer it in different words (
flagjoined with|against an array,typeas"string"against"DTP_STRING",command/parameteragainstid), and one of those is not a rename --states2json()cannot NAME an action, so a backend cell says that there is one and never which.opts.current_statelights the column the reader’s own gobj is standing in: the description describes the CLASS and cannot carry it, so the tree passes it.A find box in the frontend view (
gobj-ui7.23.39). A tree of a real yuno is a hundred cards and the only way to locate one was to read them all. It matches gclass, name, full name and FSM state; a match wears an amber ring OVER its role colour -- the card still has to say what it IS -- and a chip counts them, because a graph that did not move looks the same whether nothing matched or the match was already on screen. The same drawing as the JSON graph’s, which sits in the same console.A folded branch leaves its space behind, so nothing else moves (
gobj-ui7.23.29). Folding re-packed the tree: measured, five surviving cards all slid 240px sideways for a fold that removed nothing they could see, and the card under the finger slid out from under it. The clicked card is put back where it was, and the survivors stay because a fold leaves a PHANTOM child -- an invisible node of exactly the width its children had, in the place they had in the order. A width reserved as a NUMBER was tried first and does not work: the layout centres a parent over its children, so a lump appended at the end moves every sibling by half of it.The gobj tree remembers how you left it (
gobj-ui7.23.29): layout, zoom, camera, folds and anchor, inlocalStorage. The layout and the folds go back BEFORE the first build, because they decide what is built; the camera after, and once only, or every fold would drag the reader back.The cards of the node viewers can be MOVED (
gobj-ui7.23.24), in the JSON graph and the gobj tree alike. The position is deliberately not kept: both are rebuilt from their source on every refresh, fold and layout change, and neither is a document of its own to save it to -- pulling two cards apart to read the lines between them is worth having even for one session. The finger’s half of it needed the mounts to refuse the browser’s gestures: G6 putstouch-action: noneon its CANVAS and nothing on the html nodes over it, so a drag that began on a card was a page scroll that died after ~20px. Same defect the treedb graph paid for at 7.23.14.A camera ANCHOR in every node viewer (
gobj-ui7.23.19, made to work in 7.23.25 through 7.23.27, which is where the interesting part is). It had no visible state and never centred anything: the class it used was styled only for the G6 plugin toolbar;graph.focusElement()does not move these graphs at all; a camera move issued inside G6’s click dispatch is swallowed;drag-elementturned every click that drifts two pixels into a drag, so the node could not be picked with a real hand; the rules that drew the crosshair and the mark were scoped to a host gclass that is not always there; and the camera translate is not in pixels, so stepping by the measured gap oscillated instead of converging. The last one is a closed form now, read off G6’s owngetTranslateOptions():T = (canvasCentre - nodeWorldPosition) * zoom): pick one element from the toolbar’s crosshairs and every zoom leaves it in the MIDDLE.C_YUI_JSON_GRAPH,C_YUI_GOBJ_TREE_JSandC_G6_NODES_TREE, from the same button. A graph that FITS on screen is unreadable at the zoom that makes it fit -- one topic’s schema fits at 37%, where every card is grey texture -- so the useful view is always a fraction of the document, and which fraction was nobody’s decision:1:1translated to the layout’s origin, a corner with nothing in it. Three states (off,arming,on) because two could not say what a press does, and an armed anchor takes the next node click before selection, ports or popover. A zoom re-centres and a pan does not --aftertransformfires for both, so the zoom LEVEL is what tells them apart. The target is remembered by identity (thepath, thefull_name), never by node id, because ids are generated per build. The viewers also open at ACTUAL SIZE now, centred on the anchor or the root, and a JSON card stops listing the containers it already draws as cards -- what made an array of N dicts an N-row card beside N cards.The JSON viewer remembers which of its three views you read in (
gobj-ui7.23.17). It opened on the tree every time, however many times you switched to the graph. The choice is kept inlocalStorageunder one key for the whole library, because which view someone reads JSON in is a habit of the PERSON and not a property of the document. Precedence is host, then memory, then the tree -- which is why theview_modeattr no longer declares"tree"as its default: as a default and as a host’s explicit choice it was the same string, so nothing could tell “show me the tree” from “I have no opinion”, and a memory that cannot see the difference has to lose to both. Only a view the READER picks is remembered.A tranger key can be deleted —
C_TRANGERgainsdelete-key(topic_name,key,force).tranger2_delete_key()has been in the timeranger2 API all along (master-only, and it propagates the delete to the in-process subscribers and to thert_by_diskfollowers), but the only way to reach it wasdelete-nodeon a topic that belongs to a TREEDB. A plain tranger topic had no path at all:list-keyscould show a key born of a port scan or a typo, and nothing could remove it — it stayed in the topic, and in every view derived from the topic’s keys, for ever.The command refuses a key that is not there (
tranger2_delete_key()answers 0 for one that never existed, so a bare wrapper would report a delete that deleted nothing), and refuses a key that still holds records unlessforce=1— naming the record count in the refusal, which is what tells the operator what forcing would cost. Samedeleteauthz asdelete-topic.gui_treedb’s Keys picker grows the matching button: a third action on each key row, next to Rows and Live. It asks first, with the topic, the key and the record count in the question, and closes that key’s open cards before the delete — a Rows card holds a server iterator on the key and a Live card a realtime feed, and both would otherwise be left pointing at something that no longer exists.The JSON viewer shows the same document three ways (
gobj-ui7.20.0, 7.21.0, 7.22.0). The third is a GRAPH — it hosts theC_YUI_JSON_GRAPHchild that already drew JSON as a hierarchy, so what is new is not the drawing but that you no longer leave the viewer to get it. The switch became one button per view: three views do not fit a toggle, because a cycling button cannot be aimed. Two layout facts came out of it, both only a browser could tell you — a canvas pushes no height, so the graph body needs a DEFINITE height and not a minimum (a percentage height does not resolve against a box sized by a minimum: G6 came up 1061x2); and at 390px the toolbar’s search box was taking 294 of 320 visible pixels, which put every button off the edge.The graph then got the two facilities the tree already had (7.22.0): a find box that highlights matching rows, outlines their cards and says how many matched without moving the camera, and expand/collapse that folds every card but the root and marks each cut with a count. Its highlight is baked into the card’s markup and not set as a G6 node state — the key shape of an
htmlnode is a DOM element, and G6 paints no state style on it. Each non-leaf card then got its own fold handle (7.23.0), because a graph you can only open whole or close whole is not navigable — drawn as the same filled chip the gobj tree uses (7.23.1), after the bare glyph it shipped with turned out to be two pixels of ink at the zoom that fits a document on a phone. The graph also picks its layout now (7.23.2) — vertical tree, dagre top-down, dagre left-right — and the treedb topic’s JSON popups became movable, maximisable WINDOWS on a laptop, a document being something you read while looking at the table it came from. And its camera stopped using a different picture from the treedb graph’s for the same two buttons (7.23.3) — one action, one drawing, zoom readout included. Every graph’s camera is built in one module now (7.23.4), because unifying two files and leaving the third is how they drifted in the first place — and the offline demo finally shows the treedb graph too (7.23.5), which is the one the other two borrow that camera FROM and the one nobody could see there. Same drawings in the same PLACES, last (7.23.6, 7.23.7): the GLOBAL fold leads the toolbar ahead of the find box, and the PER-NODE one sits on the right of each card’s own header — two controls, two jobs, two places. And both graph find boxes gained the clear (✕) the tree viewer’s always had (7.23.8), reported on a phone, where there is no keyboard shortcut to fall back on.The JSON viewer shows the same document as raw text (
gobj-ui7.20.0).C_YUI_JSONhad one way to read a document — the lazy tree — and a tree is the wrong shape for some of what people do with JSON: read a command answer as it is written, take a slab of it into a ticket, find a string with the browser’s own Ctrl+F. A toolbar switch turns it into aJSON.stringify(…, 4)dump of the working document; theview_modeattr ("tree"|"text") lets a host open straight into it, andEV_SET_VIEW_MODE {mode}moves it at runtime, toggling when no mode is given.Nothing there is lazy: it prints what the client currently holds,
__collapsed__sentinels included, because that is honestly what it has. Over 2M characters the dump is cut and the cut is announced. Search and expand/collapse hide with the tree — they act on tree rows and have nothing to act on here — while copy stays. Long lines scroll sideways inside the viewer instead of wrapping: in a raw dump the indentation IS the structure, and a wrapped line restarts at column 0 and lies about the depth of everything under it.Connections acts on many, and each gesture is one write (
yunos-js0.16.0, 0.17.0). The browse column gets its header checkbox — three states, covering what the filter leaves on screen, counted over services so it cannot read “all” while half a connection is unticked. And connecting several gets its own dialog, opened on the connect INTENT of every connection: ticking everything is one click, disconnecting a handful is the same box, and its count says what Apply will CHANGE rather than what is ticked. Neither borrows the row checkbox, which means browse and cannot mean two things. Both apply in a single write (EV_SET_CONNS_BROWSE,EV_SET_CONNS_ENABLED), so the app root reconciles the transports once instead of once per connection.Several graph nodes can be selected, and they move together (
gobj-ui7.16.0). In edition mode, shift+click adds a node to the selection or takes it out, shift+drag on the canvas is a rubber band, and dragging any selected node moves the whole set — as one undo, because G6 batches it. The group move costs nothing because of where the selection is kept: G6’sselectedelement state IS the selection, which is whatdrag-elementasks the graph for. The ring had to be painted into the card’s own html, an html node drawing no state style — the same trap that kept the amber highlight invisible until 7.3.0, so turning the rubber band on and nothing else would have selected correctly and shown nothing.And the keys that selection needed (
gobj-ui7.17.0). Esc clears it, ctrl/cmd+A takes every node, Delete deletes it — behind the same confirmation the per-node icon shows, from the same function, with the children about to be UNLINKED and the parents about to be detached summed over the set: these views delete withforce, so twelve cards can detach eleven children. The keys reach the graph only while it has FOCUS, G6 giving its canvas atabIndex, which is what keeps ctrl+A typed in the find box a selection of the text. They arrive asEV_KEY_DOWNand the action decides, the two full-screen keys included — they used to call the plugin straight from the callback.The two decisions a runtime-opened tab costs its url, in one place (
gobj-ui7.19.0,yunos-js0.18.0). Both consoles have a workspace whose tabs the operator opens, and both answered the same two questions in their ownc_app.js— which is exactly why one could have the cold-load one WRONG while the other had it right, and nothing said so. The deciding isyui_tab_routes.jsnow (yui_tab_split_subpath,yui_tab_position_plan), tests and all; the wiring stays in the hosts, which are each right about when their own tabs become real.Zoom to the selection (
gobj-ui7.18.0).fitgives back the whole graph; there is now the same action for the part being worked on, as a button next to it — edition only, disabled while nothing is selected, and wearing thefitbrackets with a marked object inside so the two read as one action at two scopes.fitView()has no subset form, so the bounds are measured off the elements and the zoom clamped to the graph’s ownzoomRange. New host key:zoom to selection.The G6 graphs, operable on a touch screen (
gobj-ui7.23.9). Measured in a real touch context rather than read off the CSS, and two of the five are facts about G6 that nothing in its documentation says:Pinch to zoom, in all three graphs.
zoom-canvasbinds the WHEEL and nothing else, and a telephone has no wheel. G6 does ship a pinch recogniser, but asking for it (trigger: ['pinch']) REPLACES the wheel — itsbindEventsis anif/else— and itsPinchHandlerkeeps its instance and its callback list in STATICS, so on a page with two graphs the second registers against the first one’s emitter: pinching graph A zooms both and pinching graph B does nothing. Recognised per graph instead, over azoom-canvasthat keeps the wheel — the same shape thedrag-canvasreplacement already had, so every graph gets it with no change to its behaviors list.It reads the NATIVE touch events, not G6’s forwarded pointer stream:
@antv/gre-issues pointer ids in the middle of a two-finger gesture (measured: a pinch that started on ids 2 and 3 finished on id 1), so anything keyed onpointerIdloses a finger halfway and reads the gesture as a fraction of what it was.event.touchesneeds no bookkeeping — it IS the list of fingers down, restated on every event.A long press opens the context menu. It never could, on any platform: G6 does not read the DOM’s
contextmenuevent at all. ItsBehaviorControllersynthesises the event frompointerdownwithbutton === 2, so the menu was a right click and only a right click, whatever the browser does with a long press. The press re-emits G6’s own forwarded event under the name the plugin listens for, sogetItems(e)sees exactly what a right click gives it — the port under the finger included, which is what tells the port menu from the node one.Touch targets, all behind
(pointer: coarse)so a mouse sees no change: resize handles were 8x8 and eight of them (now a 14px mark in a 44px box, corners only — eight fingertip-sized boxes around a 90px node overlap into one blob, and a corner resizes both axes anyway);node propertiesanddelete nodewere two 28px circles 4px apart, one fingertip covering both with the destructive one underneath; a port’s hit area was a flat+4in WORLD units, a different target at every zoom and 5 screen px at the 50% a telephone lands on after fit.The floating toolbars fold. Drawn inside the canvas, one on each edge, the two of them took a third of a 356px telephone canvas and stood on top of the nodes. Under 480px of container they collapse behind one button. Measured on the CONTAINER and not the window: the same graph is a full page in one app and a card in a column in another.
New host keys:
show toolbar,hide toolbar.Multi-selection reaches a finger (
gobj-ui7.23.11, 7.23.12). Both of its gestures hang off Shift — shift+click adds a card, shift+drag draws the band — and a phone has no Shift; nor is there a spare gesture to give them, since G6 binds panning and the band to the same plain drag. So Shift becomes a MODE: the graph’s edit toolbar carries a selection mode toggle (a dashed marquee, next to+), and while it is on a tap picks a card and a drag on the background draws the band, with panning standing aside. A button rather than a heuristic, so the toolbar says which of the two the graph is listening for; and not device-specific, because it spares a desktop reader the key just as well. New consumer key:selection mode.The toggle looks pressed rather than taking a colour of the toolbar’s palette (
7.23.12, newset_pressed_state()): every colour there names a KIND of action — blue creates, orange means pending changes, violet is undo/redo, red destroys — so painting a STATE with one of them put the same violet on two neighbouring buttons for two different reasons, and carried the whole state change in the hairline of an outline glyph.
Fixed¶
diff-schema answered 43 differences nobody had made¶
A treedb of 45 columns reported 43 differences, every one of them
fillspace “only in stored” with the value 10. No Apply could ever settle
them, because there was nothing to apply: fillspace DEFAULTS to 10 in the
system schema, almost no schema writes it, and the two sides of the
comparison were not symmetric about that.
The projection of a column copies the attributes the schema DECLARES and no
more; the stored node went through treedb, which fills every attribute the
descriptor gives a default. So an attribute the schema never mentions is
absent on one side and holds its default on the other, and the comparison
read that as an operator addition.
It is the same failure is_unset_value() was written for — its own comment
says “those 586 defaults buried the 6 that somebody made” — except that
one only knows the EMPTY value of a type, and what bit here is the DECLARED
default, which is a different thing. is_declared_default() knows it.
The projection is deliberately NOT filled with defaults instead: it is what the projector UPSERTS, so a default written there would overwrite the value an operator set by hand on an attribute the schema does not declare. The asymmetry is real and belongs in the comparison, not in the projection.
Measured after the fix: 43 differences -> 0 on a treedb whose schema matches, and on a second node the panel now reports exactly ONE — a column that really is only in the store. Which is what the panel is for.
The gclass window’s ✕ left its viewer alive, and the button never opened another one (
gobj-ui7.23.38).close_window()callson_closeand then destroys the WINDOW; the viewer is not the window’s child -- it hangs from the host -- so nothing took it down. It kept its service NAME, so the next click answered “service ALREADY registered” and no second gclass could be opened for the rest of the session; and it was still RUNNING when the host’s own window closed, which logged “Destroying a RUNNING gobj”. The defect predates the new viewer: it is the arrangementC_YUI_JSONwas opened with, and it needed somebody to close the window and click the button AGAIN to show itself.An event declared with NO action was drawn as if it had one (
gobj-ui7.23.37). The C side writes the literal"action"for every action it cannot name, and the matrix marked anything that was not an empty string -- so{EV_TX_READY, 0, 0}, whichC_WEBSOCKETdeclares in all four of its states, claimed an action the gclass never wrote.A reciprocal pair of states drew one arrow on top of the other (
gobj-ui7.23.37). Two G6 facts came out of it, both measured: atypewritten on an edge DATUM is overridden by the graph-leveledge.type(vary it per element with a FUNCTION), andcurveOffsetis measured from each edge’s OWN direction, so a reciprocal pair needs the SAME sign on both -- opposite signs bow them to the same side of the screen.The popover called two different things “Estado” (
gobj-ui7.23.39). The run status and the FSM state sat next to each other under the same label in every Spanish app, one saying Parado and the other ST_IDLE. The second row has its own key now.“1 matches” (
gobj-ui7.23.40). Four find chips counted with two spans -- a number beside a word -- so nothing could see the count and one match read as many, in both languages. They say it as one counted string now, with the singular inmatches_oneand the plural in the base key.A service the app creates is the app’s to START (
yunos-js).c_yuno’smt_playstarts only the DEFAULT service, andC_YUI_WINDOW_MANAGERhas nomt_start-- its dock is built inmt_createand driven by events -- so it WORKED stopped in two consoles and nothing ever complained. It just read!!C_YUI_WINDOW_MANAGERin every trace line, and stopped in the frontend view.The machine trace was written in the LEGACY shape in the browser (
gobj-js7.13.7, completed in 7.13.8).trace_machine_formatwas0in the JS port and1in the C kernel (gobj.c: “0 legacy, 1 simpler”), so a browser yuno wrote THREE lines for one transition -- the call, the state change and a<- mach(…) ret: Nreturn -- where a node wrote one. The same trace, three times the wall, and the two sides did not look alike read side by side.Moving the default was not enough, twice over. Four of the SIX trace sites had no simpler shape at all -- the state change, the event injection, the publish and the subscriber forward still wrote
mach(…)-- so the trace came out MIXED, which looks exactly like a default that did not take. And a stored preference beats a default: every browser that had ever touched the dev window’sSimple machchip carried an explicit0, so the change reached fresh profiles only until the preference moved to a new key (gobj-ui7.23.34, 7.23.35). Both shapes stay; the chip swaps them.With them, the publish path stopped dumping an empty
{}under every publication -- unguardedtrace_json(), one blank line per tick in a yuno whose timer publishes.The
machinetrace could not be quietened, and then lied about it (gobj-ui7.23.32 through 7.23.35,gobj-js7.13.6). The dev window’sPeriodicchip promises to hide recurring events and hid none of them: it matches on a SIGNATURE and every line mirrored from the log was signed by its LEVEL, so a hundredEV_TIMEOUTtransitions and a hundred unrelated debug lines were one thing to it. A machine line is parsed now and signed by its EVENT, matched by name (never by count -- a busy FSM crosses a recurrence threshold on nearly every event within seconds), and an error or a warning is never hidden whatever event it names. The filter is ON by default.It then had to reach the CONSOLE, where the same flood was arriving one pane over. It could not be done from the GUI side: gobj-js writes the console line BEFORE it calls the log sink, so nothing downstream can un-print it. It hands out a per-line say instead (
set_console_log_filter), and the window installs the same predicate its own filter uses -- one rule, two sinks.Two more in the same window: the trace lost its nesting INDENTATION on the way to the DOM (
createElement2()trims a string content, and those leading spaces ARE the nesting), and the empty state said “Waiting for activity — enable Traffic or Automata for more” while the status line beside it said0/36 shown · 36 hidden.A graph viewer started while its DOM was still DETACHED never resized again (
gobj-ui7.23.31). ItsResizeObserverwas looked up withdocument.getElementById(), which answersnullthere, and the guard around it skipped the attach silently -- so the canvas kept the size it was born with for the life of the window while the window resized around it. Two hosts hit it, the gclass viewer and the treedb graph’s raw-JSON viewer, and only when the reader had last used the GRAPH view, becauseC_YUI_JSONremembers the view mode and builds its graph child insidemt_start. The mount is read from the gclass’s own$containernow, and both hosts start the viewer AFTER its presenter has put it in the document.A JSON card was a single anchor for every line leaving it (
gobj-ui7.23.21, refined in 7.23.22). Fourteen edges came out of one point, and which row a line belonged to was a guess the reader made from where it landed. Every container key now opens a G6 port and its line leaves from there -- on the bottom edge, and exactly ON it, because an html node draws its HTML in a DOM layer above the canvas and a port fully inside the box is painted under the card. A container with no card of its own hands the port its row opened in the parent down to all its children, so the fourteen columns ofcolsleave the singlecolsport.getPointPosition()moves tolib_graph.js, shared with the treedb graph it came from: it is the same decision about where a hook’s port sits. Each port then moved onto the LINE of its own row (7.23.22): spread along the bottom edge they were distinguishable but still not attached to anything, since the reader had to count dots, count rows and trust the two orders matched. And a line gained an ARROWHEAD and a port to arrive at, centred above the target card’s title (7.23.23): with ports only on the source side a line said where it left from and nothing about where it went.A JSON graph drew a pure collection as a node of its own (
gobj-ui7.23.20).colsis one key of the topic dict likepkeyis, and the card the graph gave it held no data, said nothing its parent’s row did not already say, and pushed everything below it one level down. Every key is a row again, containers included -- a container row shows what it IS (cols: [14]) and nothing more, since what it HOLDS is the cards, and that is the part which has to scale. A container with no scalars of its own now gets no card at all and folds from the chip on its row.Highlighting a JSON value made it HARDER to read than not highlighting it (
gobj-ui7.23.18). On the graph view’s amber match chip an orange boolean measured 1.80:1 and a red number 2.98:1, in both themes. No chip colour could have fixed it:#FF8C00is mid-luminance, so it reaches at most 2.33:1 against any background that exists, white included -- the ceiling is a fact of the colour, not of the surface. Orange moves to#8C4D00and blue to#4359C6, and every value type now clears 4.5:1 on all five surfaces the viewer paints (two light cards, the match chip, two dark cards); the worst in the matrix is 4.66:1. The graph also stops deriving its dark colours by mixing the light ones with white -- which is how one document was green in the tree and a paler green in the graph -- and carries the tree’s two palettes instead, so all three views agree.The JSON viewer read better in LIGHT than in dark (
gobj-ui7.23.17), which is normally the other way round, and the measurement said why: the dark graph card mixed 30% of the group’s tint into the SURFACE, so green string values sat on a green card. The light card is near-white and lets saturated dark text carry the colour; the dark one carried it twice. Thelistcard was worse -- a yellow tint at 30% lands mid-luminance, a muddy olive where nothing contrasts with anything, and its purple values measured 1.60:1 against 11.48 for the same values on the light card. The surface now stays out of the hue on both themes and the group’s colour lives in the border and header bar, where the light theme always put it; worst case goes to 4.74:1. The search highlight moved with it: its amber failed at both ends once the values brightened, and a chip dark enough for the text is invisible against the card, so a match is now a LIGHT chip on both themes carrying light-surface text -- the same flip the header bar already does.A container row would not say WHICH record it was (
gobj-ui7.23.17). An array of dicts carrying anidis a list of records -- a topic’s nodes, for instance -- and the row said only2: {15}, so telling record 2 from record 9 meant opening all fifteen fields of both. It now carries the id beside the size, open as well as closed: it is the row’s label, and a label that vanishes on expand makes the row jump and costs you the name of the thing you just opened.A table whose
fillspacewas a string printed a stack trace per cell (ycommand,ybatch,mqtt_tui,msg2db_list,treedb_list).fillspaceis a column WIDTH, a presentation hint, and a command that declared it as"30"instead of30-- whichjson_pack("s:s", ...)does the moment the field is written next to the other three, all of them strings -- madekw_get_int()log “path MUST BE a json integer” withLOG_OPT_TRACE_STACK. Three reads run per column (header, separator, and once per data cell), so a four-column answer with two rows emitted twelve stack traces and buried the answer, and every column fell back to the default width of 10, which is what truncated the ids the operator was reading.yclihad been passingKW_WILD_NUMBERhere for exactly this reason and its four siblings never got it; all fifteen read sites now do, so a width is accepted as written, string or int.The toolbar lost its scroll arrows once it shared a row (
gobj-ui7.23.16). Ayui_toolbar()alone in its row works, so the defect hid until the treedb graph was reached through the navigation, where the pinnedGRAPH_BACK_TOPICSlink sits beside it. A flex item’smin-width: autoresolves to its CONTENT minimum, and everyyui-horizontal-toolbar-sectionisflex-shrink: 0; white-space: nowrap— so the toolbar refused to shrink, keptwidth: 100%of the row, was pushed out of it, and its right arrow (absolutely positioned at its right edge) landed outside the ancestor that clips the view. The arrow was drawn all along, off-screen, while the items it would have scrolled to were plainly cut off.A group move that moved the whole graph (
gobj-ui7.23.10). Three defects behind one report: selecting two cards in a treedb graph and dragging them moved everything, erratically, and the saved result was right anyway. (1) Edition’sdrag-canvasis given anenableso it stands aside for Shift, the marquee’s key — and anenableREPLACES G6’s default, which is thetargetType === 'canvas'test that keeps panning off a node drag: both behaviours ran on the same gesture, so a 150px drag moved the card 255px and every other card 105px, and a pan writes nothing, which is why saving and refreshing showed the right thing. (2) The history plugin was installed at one MOMENT — the arrival of the last topic of the load — and only if the graph was in edition right then, so reaching edition through the mode selector, which is the ordinary way in, left dead Undo/Redo buttons and ahistory_pause()nothing answered, while Save still lit on its own. It follows the mode now. (3) G6’sbrush-selectrewrites theselectedstate of every element on every canvas click, behind the gclass that owns the selection and outside its history pause — a recorded command whose before and after are identical, so a click on the background lit Save on a graph nobody had touched.The treedb graph’s two selects speak the app’s language (
gobj-ui7.23.13). The layout one and the operation-mode one rendered their raw names —reading,edition,dagre,manual— in every language and in every app: not a missing key, a missing call, since neither went throught()at all. The Spanish console of a production node said “Modo de operación: reading”, label translated and value not. Two things came with it: the option’svalueis set EXPLICITLY (an<option>with none answers with its own TEXT, so a translated label would have sent"Edición"to the FSM as the mode to enter), and the labels are literals rather thant(name), because a consumer’svalidate-locales.mjsreads literals and a variable key is invisible to it. New consumer keys:reading,operation,writing,edition,manual,dagre,antv-dagre,d3-force,force-atlas2.The demo’s treedb chapter opens its edition mode (
gobj-uitest-app). It was mounted read-only because there was nothing behind it to write to — still true of the RECORDS, where the refusal is the honest answer twice over (the graph draws a created or deleted node from the treedb’s own node events, which a backend living in one page does not send, so a write that answered yes would leave the graph still). It is not true of the ARRANGEMENT: moving cards, the selection mode, undo/redo and Save all end in one write,update-nodeon__graphs__, and that topic is the view’s own bookkeeping.C_DEMO_BACKENDdeclares it and takes that write, so the public demo now shows the whole of edition — including the selection mode added in7.23.11.A mode that could not move the camera (
gobj-ui7.23.9). The treedb graph’soperationmode left itsbehaviorslist empty wherereadingandwritingnext door both fill it. On a desktop the toolbar still zoomed and nothing panned; on a telephone, where the gestures ARE the camera, the graph was a picture.A view that talks to a backend has to be a SERVICE, and for two reasons (
gobj-uitest-app). Only one of them is visible, and that is the trap:gobj_save_persistent_attrs()refuses a gobj that is not a service and says so, butC_IEVENT_CLIroutes an answer back withgobj_find_service(gobj_name(src))and simply finds nobody. The offline demo mounted the treedb graph as a pure child and WORKED — its in-page backend answered thesrcpointer it had been handed, which is what a backend in the same page has and a real one never does. So the chapter modelled a contract that does not exist. Both halves fixed together: the graph goes throughyui_mount_service_view(), the way both real consumers mount it, and the demo backend resolves its destination by NAME and refuses, loudly, to answer a caller that is not registered.Three framework contracts, each found the same way: by a loud error nobody had seen because nobody had built that shape yet (
gobj-ui7.20.1, 7.23.0, 7.23.2). They are gobj rules, not component details, which is why they are here and not only in that repo’s changelog:gobj_destroy()destroys the children BEFORE callingmt_destroy(). A hosted child torn down inmt_destroyis torn down after the framework has already destroyed it — while it was still running. Retire a hosted child inmt_stop, where everything is whole.C_YUI_WINDOW.close_window()callson_closeand THEN stops and destroys itself. A host that also destroys the window destroys it twice. On the ✕ path the host drops its reference and keeps its hands off; it destroys the window only when IT initiates the close.Adding an output event to a gclass hosted as a CHILD is a BREAKING change. The child subscribes its host to everything it publishes, so a new output event is a new mandatory declaration in every host’s FSM. Six gclasses answered “Event NOT DEFINED in state” to a graph click because one viewer started republishing its child’s event “so the host has one contract” — which is the framework backwards.
F5 on a deep Schemas route landed on somebody else’s default (
yunos-js0.17.6). Reloading on#/schemas/node/<node>%1F<yuno>/treedb_authzs/__graphs__answered with.../treedb_system_schema/edit. A node tab’s route is registered when the node is OPENED, so on a cold load it does not exist yet: the shell resolves as far as the workspace home and hands the whole rest over as the subpath. The restore read all of it as the node id, matched nothing, and fell back to the first tab — which stamped its own default treedb, and that treedb its own default view. Every segment past the id was thrown away, which is why a bare node tab survived a reload and nothing deeper did.An action route came back to the MOUNT, not to where you were (
gobj-ui7.19.4). Withredirect: "back"or"none"the shell restores the previous resting route, and it read that offstages.main.active_route— the route the view is DECLARED at. Under aC_YUI_NODEtree everything below is subpath the node owns, so switching the theme from a graph five levels down landed on the workspace root, position gone; the same held for/preferences,/sitemapand the dev-tools routes. It restorescurrent_routenow, the same route WITH its subpath. An app whose views are all declared routes never saw this, which is why it took a node tree to surface it.Shift+clicking a card smeared a text selection across the graph (
gobj-ui7.19.2 — 7.19.3). Shift+click is the browser’s own extend-the-text-selection gesture, so the moment7.16.0gave it a meaning here, marking three nodes also painted their labels blue. The canvas is a canvas — except the popovers, which carry record data an operator copies out, and7.19.2took the readable one (g6-node-detail) down with the rest before a measurement caught it.1:1was a dead button, and the minimap drifted in full screen (gobj-ui7.19.1). G6’s toolbar callsonClickonly when the PRESSED element carriesg6-toolbar-item— which is why its own CSS makes the icon<svg>click-through. The text glyph7.15.0added had no such rule, so a real press landed on the<span>and nothing happened; it survived a live check because that check usedelement.click(), which dispatches where it is told rather than where a pointer is. The minimap is placed in PIXELS computed once from the canvas size, so growing the container left it halfway up the left edge, over the graph it explains: it is anchored in CSS now, inset, and follows the theme instead of being a white box on a near-black canvas.After clicking a node, the graph’s keys did nothing (
gobj-ui7.18.1 — 7.18.3). The keyboard reaches the graph through G6’s canvas, the element carrying atabIndex, and a card is a DOM element inside the container — so clicking one sent the focus to<body>and Escape, ctrl+A and Delete stopped working immediately after the click that had just selected something. Three releases, each closed by MEASURING rather than reasoning (a probe reportingdocument.activeElementand every keydown): the first focused the wrong canvas of the four a graph stacks, the second ran before the browser’s own mousedown focus handling. The focus is restored onfocusoutand only when it goes nowhere — moving to a real element is the user leaving.The graph toolbar’s house never took you home (
gobj-ui7.15.0). The button people press to get the treedb graph back runszoomTo(1): it sets the SCALE and leaves the camera where it was, so from a corner of a large graph it answered with the same corner at 100%. A house means the initial extent in a map and the starting view in an editor, never a scale, and it sat directly underfit, which is the control that really does give the graph back. The action is right, so only its name moves:1:1, written and not drawn, the way every editor that offers actual size labels it. With it, the zoom level is shown as a readout, groups are separated by gaps rather than one more hairline, and both floating toolbars follow the theme — they were pinned light in both, two bright islands over a dark canvas. New host keys:actual size,zoom level.A header checkbox that could not be clicked (
yunos-js0.16.1). Two defects on one click, both found by driving the deployed app. It was born disabled: Tabulator draws the header before the rows exist, and the first load was the one data path that never repainted it afterwards. And the click cancelled itself:preventDefault()on a checkbox makes the browser revert the tick when the dispatch ends, while the repaint the event triggers runs in a microtask BEFORE that — state written, table repainted, count updated, and then the revert landed last.A proposed url built on the realm id (
yunos-js0.17.1). The agent console’s For TreeDB copy read the FQDN from one place only, the__ssl_certificate__config variable, and fell through to the realm id when it was absent — proposingdemo.hidrauliaconnect.esfor a backend whose certificate sayshidrauliaconnect.es. The evidence was in the same config, written where it is USED (crypto.ssl_certificateof the gate), and is now read there too: the gate wearing the top port wins, a wildcard certificate names no host, and the realm id is the last resort.The
yunos-jsCHANGELOG was in Spanish from 0.13.15 to 0.15.1. Translated in place: this is a public repo, and those are English.A finger could not move a node (
gobj-ui7.23.14, 7.23.15). The plainest thing there is to do in the treedb graph’s edition mode, and two separate defects were in front of it — the first one not in this code at all.The browser was taking the gesture. G6 puts
touch-action: noneon its canvas and nothing on its HTML nodes, which are ordinary DIVs layered over it. So a drag that began on a CARD was a page scroll as far as the browser was concerned: it let twopointermoves through, decided, and killed the pointer stream with apointercancel. The node followed the finger for about 20px and stopped dead while the page slid underneath — which reads as a delta bug and is not one. The pointer log is what says so: it ends atpointercancel/lostpointercapturewhile the touch stream runs on to the end of the gesture..graph-containerrefuses those gestures whole now, withtouch-action: autokept for the panels and the context menu, which a finger must still be able to scroll.And the press meant two things at once. The long press fired on a TIMER, 500ms in, while
drag-elementwas already carrying the node: the menu opened over a card that then ran away underneath it. A timer cannot arbitrate a gesture, because at the moment it fires the gesture is not over. Nothing decides now until the finger moves or lets go — moved → drag, still and let go quickly → the element’s own action, still and held past 500ms → the context menu. The rule isclassify_press()in the new purepress_arbiter.js, with tests; its 10px slop is G6’s owndragstartDistanceThreshold, so “still” means the same to the arbiter and to the drag.Three more things the same press was doing, all fixed with it: the
clickthat@antv/gSYNTHESISES in itsonPointerUp(it does not takeclickfrom the DOM either, so swallowing the DOM one never helped) went on to click the node the menu had just opened on; the browser’s own menu opened on top, while the finger was still down; and each release of a PINCH looked like the end of a press.Deciding at the release costs exactly one thing — while the finger is down, nothing says that letting go would now give the menu rather than the node’s own action — so
7.23.15adds a 15ms haptic tick at the 500ms mark. It is a NOTICE and not the decision: a finger that buzzes and then carries the node away still gets its drag.
7.16.2¶
JS layer only (gobj-js 7.13.2 -> 7.13.5, gobj-ui 7.10.5 -> 7.14.3,
yunos-js 0.13.3 -> 0.15.1). Each repo carries the detail in its own
CHANGELOG.
Fixed¶
A subscription filter that never filtered (
gobj-js7.13.3).gobj_publish_event()read the answer ofkw_match_simple()— a JS boolean — with the C runtime’s=== 0guard, so a subscription whose filter did not match was published to anyway. A treedb view subscribes toEV_TREEDB_NODE_DELETEDonce per topic with a{treedb_name, topic_name}filter, so deleting one row reached the table five times and the four strays logged “record not found”. Measured on the wire first — one frame out, one answer in — which is what said the copies were made in the browser.The audit of the same trap across the runtime (7.13.4, 7.13.5) fixed four more sites:
mt_publication_pre_filter,mt_play,mt_subscription_added, and therc_walk_by_tree/rc_walk_by_listreturn, now normalized through awalk_ret()helper. The other half of the trap is the caller: an action or framework method that answersundefined(a barereturn;) reads as stop to a< 0guard and as nothing to a=== 0one. 31 off-contract returns were corrected across the five JS repos, with regression tests for the publish path.The schema editor’s buttons and icons are the size of buttons (
gobj-ui7.13.4 - 7.13.6).is-smallhad been applied to every control ofSCHEMA_CARDandTREEDB_CARD, form fields included, and the card grid was 12rem where a treedb name needs 14. In the same round, the header-filter hairline that 7.13.1 and 7.13.2 both failed to land: the cause was neither the token nor the specificity but the CSS import order — our Tabulator fixes were emitted before the theme they fix.
Added¶
Rows can be selected and removed in bulk, in any table (
gobj-ui7.11.0 - 7.13.0). A shared facility,yui_table_select.js(yui_selection_column,yui_selection_settings,yui_selection_bar,yui_wire_selection), instead of a per-table checkbox: the treedb topic table takes it behind an opt-in flag (with_selection_bar), and so do the tables of the SPAs. The bar reports what is selected and clears itself; a tri-state parent checkbox writes its children in ONE gesture (theindeterminatestate is a property, so those formatters must return DOM nodes, not markup).A column drag can be undone (
gobj-ui7.14.0, 7.14.1). Reordering a schema’s columns raised Version not raised with no way back short of reloading the store. The remembered order is now reset only when the store is loaded, not on every write — which is what made the first Undo button invisible.Connections and the Schemas picker read the same (
yunos-js0.15.0, 0.15.1). The two tables show a backend and the treedbs it exposes, and one is pasted literally into the other, but one was a tree and the other a flat table with a nested sub-table; the checkbox meant open in one and select for deletion in the other. Connections adopts the picker’s shape: services are child rows, the checkbox means BROWSE in both with the three states on the connection row, and search, count and fold sit in the same place — the fold and the search are one wrapping unit, so a phone takes them to the next line together instead of leaving the fold stuck to the title. Removing several connections moved to its own dialog, with its own checkbox list, so marking rows to browse can never delete them.
7.16.1¶
Fixed¶
The identity ack freed a record it does not own (
c_ievent_cli).ac_identity_card_ack()read the ievent stack record withmsg_iev_get_stack()and then released it. That function returnsjson_array_get(jn_stack, 0), a pointer thekwstill owns — its own signature says so: “Return is NOT YOURS!”. The decref freed__md_iev__.ievent_gate_stack[0]while the array kept pointing at it, and theKW_DECREF()that follows inac_on_message()walked the tree and freed it a second time.The second free reads a refcount out of released memory, so what happens next belongs to the allocator and not to the code: the block can still read
1and be freed again, or it can have been reused and corrupted in silence. That is how one stray decref per connection survived since the gclass was imported. It is no longer silent on glibc 2.43 — the agent aborts with “corrupted double-linked list” while it processes the controlcenter’s ack, seconds after start, and a run that gets past the ack dies later in_int_mallocinstead.The path runs on every successful identity ack, which is every agent-to-controlcenter link and every citizen-yuno-to-agent link. This was the only site in the tree that treated the borrowed record as owned:
c_controlcenteralready pairs itsmsg_iev_get_stack()with an explicitJSON_INCREF().The watcher changed directory without looking (
ydaemon).relauncher()calledchdir(work_dir)in the child and dropped the result, which the compiler had been reporting as-Wunused-resulton every build. A failedchdirleaves the yuno running from whatever directory the watcher was in, and it said nothing about it. It now logs the path and theerrno.
7.16.0¶
Added¶
The daily report says which machine sent it (
webstats). The mail left every node with the same sender, the one persisted inemailsender’sfromattr (no-reply@artgins.comhere), so five nodes reporting into one mailbox were indistinguishable by sender — the hostname was in the subject and nowhere else, which no mail client sorts or filters by.email_fromnames it, and the config composes it:json_configresolves(^^__hostname__^^)inside a string, so"C_WEBSTATS.email_from": "(^^__hostname__^^)@artgins.com"reaches the yuno already resolved towattyzer@artgins.com. The yuno stays out of it — no domain and no address shape built in — and any other variable works the same way. Empty leaves the sender to the email service, as before.
Fixed¶
A probe that spells itself
/%2eenvis still a probe (webstats). The probe patterns were matched against the raw path, so a scanner that percent-encodes the interesting characters —/%2eenv,/%2egit/%63onfig,/%2f%2eaws%2fcredentials— shared no substring with.envor/.gitand was counted as ordinary traffic. Eight such requests hid in one day of one node, and the packagedfail2banfilter is blind to them for the same reason: it matches the literals against the raw log line, and it has no way to decode first.The path is now percent-decoded before the patterns are applied, which is what it already means at the HTTP layer. Chasing the encodings from the pattern list instead is a game with no last move (
%2E,%2f, and the combinations of both), so the patterns stay in their plain form.A malformed escape is copied verbatim, and so is
%00: decoded it would end the string and hide the rest of the path from the match./cgi-binjoins the default patterns, inwebstatsand in the packagedfail2banfilter together — they are one list in two files and drift between them is the bug. No node serves CGI, and what asks for it is not after a script: the requests are path traversal reaching for/bin/sh.The agent’s init script started a second web server, by hand (
yuno_agent).yunos/c/yuno_agent/service/yuneta_agenthad drifted from the script the packagers generate, and it still carried a bare/yuneta/bin/nginx/sbin/nginxin itsstartcase — no check for one already running, and plain nginx hardcoded while the openresty line sat commented out beside it.It is not a dead file:
CMakeLists.txtinstallsservice/into/yuneta/agent, which is where the.deb/.rpmpostinst readsyuneta_agentfrom to seed/etc/init.d/. So a build on a source node overwrote the packaged script with the stale one — the same way a build overwrites the agent binaries.On a node serving with openresty, the result was a second web server of the wrong kind racing
yuneta-webserver.servicefor ports 80 and 443 every boot. It lost by eight seconds and left 21[emerg] bind() ... Address already in uselines behind; had it won, it would have served the stockserver_name localhostconfig for every vhost on the node.The copy here is now byte-identical to the one the
.debgenerates, which starts the web server withsystemctl start yuneta-webserver— idempotent, and it picks nginx or openresty as the node chose.ENTRY_POINT.md§8.3.1 now says who owns the web server at boot, and §8.3 adds theulimit -lthe script has always set and the doc never showed — the memlock ceiling an io_uring ring is counted against.The GeoJSON owes maplibre a real boolean (
gobj-ui7.10.5).devices2geojsoncopieddevice.connectedinto the feature properties verbatim. Four style expressions test it with['case', ...], and maplibre asserts a strict boolean there — sonull, absent, or the1/0a backend may send all fail the assertion. Reproduced against maplibre’s own expression engine.The damage was not the console. The failed assertion leaves the cluster accumulator NULL, and the cluster colour compares it against
point_count— never equal, so the cluster paints red as if a device were down. The unclustered point and its label lose their colour expression too and fall back to the property default, black, instead of green or red. The map reported a state that was not true.!!device.connectedalso makes the1/0case work rather than merely stop erroring.In the same pass,
get_coordinatesreaddevice.settings.coordinateswhile its own comment sayssettingsmay be null — a TypeError that would unwind out ofdevices2geojson.The map asks for its source instead of guessing from the style (
gobj-ui7.10.4).C_YUI_MAP’s refresh guarded itself withmap.isStyleLoaded()and then calledmap.getSource('devices').setData(). The style is the wrong milestone: the source is added on the maploadevent, andloadfires one render frame afterisStyleLoaded()turns true. A refresh landing in that frame threw “can’t access property setData, map.getSource(...) is undefined” — and the throw unwound throughgobj_publish_event, aborting the publisher’s own loop, so the caller stopped processing the rest of its batch.Measured against maplibre-gl 5.24.0 in Firefox, the window is exactly one frame, 10-42 ms, on every run — cold cache, warm cache and a
display:nonecontainer alike.Testing the style was wrong in the other direction too:
style.loaded()requires every tile manager to be loaded, so it returns to false while new tiles come in. Once the map was up, every refresh during a pan or a zoom was silently dropped and the devices stopped moving until the tiles settled.No in-repo consumer registers
C_YUI_MAP, so no yuno changes here. The same fix ships on the frozen v1 line as 1.0.2 (npm dist-taglegacy), for estadodelaire and hidraulia, which is where it was found.
7.15.0¶
Added¶
A delete says what it takes with it (
gobj-ui7.10.3,gui_treedb0.13.3). The question wasare you sure— the same words for one loose record and for one with six children hanging off it. And these views delete withforce, which on a treedb node does not only remove it: its children are UNLINKED (they survive, loose) and it is cleaned off its parents. An operator could detach six records believing they had removed one.It names what is going and adds a line per thing at stake, each only when there IS something at stake — a loose record must not be dressed up as a dangerous one. Counted off the record the table already has, so asking costs no round trip; the counting is pure and tested, because a hook or fkey value arrives as a list of refs, a dict keyed by id or a single ref string, and a column can be BOTH — which counts on both sides, because the delete does both things.
In the graph the node-delete popover carries the same lines, and the unlink popover carries the reassurance that is its whole point: neither record is deleted. Next to a delete button painted the same red, that is not obvious.
Three traps, all found by running it, and all worth knowing before composing any message from keys:
yui_shell_confirm_*renders its message as an i18n key (so a composed sentence can never be one — pass DOM);createElement2trims text nodes (so a" "separator vanishes and it reads “BorrarDeveloper” — space with CSS); and a counted word must carry noi18nattribute, because the dialog re-translates from the key alone, without the count, and puts the plural back over the singular.nodescan answer a PAGE (c_node.c). A treedb lives in memory, so walking it is not what costs: serializing every node, pushing it through a websocket and parsing it in a browser is.nodestakesfrom(1-based) andlimit, and cuts the answer on the way out.The contract is deliberately the one
list-keysofC_TRANGERalready uses: with nolimitthe answer is the plain list it has always been, so every client written before this keeps working; asking for a page gets the envelopeget-pageuses,{total_rows, pages, data}, so a client pages nodes exactly as it pages records. Filtering happens BEFORE the cut, sototal_rowsis the size of the match and not of the topic, and a page past the end is empty while still reporting the true total.KW_WILD_NUMBERon both, because they arrive as strings through the agent’scommand-yunoforwarding, which does not coerce — and that half is pinned by the test too. New test:tests/c/c_node_paged_nodes(124/124 green). Documented inYUNO_TREEDB.md§5.3.The SPA asks for pages (
gobj-ui7.9.2,gui_treedb0.12.2): each topic table pulls the page it is showing instead of the host pushing the whole topic down. The page size is generous on purpose (200), so a treedb that fits in one page behaves exactly as it did — paginator hidden, every filter seeing every row — and only a topic that does not fit pays for paging. Safe against a backend that cannot page: it answers the whole list, which the table reads as one page.The header filters and the search box work on the page that is loaded (
filterMode: "local"), which is what the tranger browser’s Rows card already does and for the same reason: the alternative is pushing every filter to the backend and changing what “search” means.Verified end to end against staging:
treedb_system_schema’scolsarrives whole in one page of 200 (114 rows, paginator hidden), and at 50 per page it really pages — 50, 50, 14, with different rows on each.Two traps worth keeping, both found by running it: the correlation id must be read flat off the command stack (
C_IEVENT_CLIEXTRACTS__md_command__and pushes it AS the stack’skw), and whether there is more page must come fromgetPageMax(), not from the row count against the page size — 50 rows of a 50-row page reads as “it all fits” and it does not.Edit a cell of a topic table in place (
gobj-ui7.8.0,gui_treedb0.11.0). Changing one field meant opening the record form, changing it, saving and closing — four clicks for a word. A writable scalar is editable in the table now, in edition mode.Which cells: the schema decides (
writableonly, never the pkey — renaming what a record is KEYED by is not a field edit) and the type decides the rest. A hook holds children and an fkey IS a link, so both are edited by linking; a dict or a list is a document the form has an editor for; a date cell shows a formatted string over an epoch, so typing into it would write the string. Those stay with the form, one click away on the same row.The write is a partial update with no
autolink, and that is the whole safety of it.treedb_update_node()merges (json_object_update), so the fields it does not carry are left alone;autolinkwipes a node’s links to rebuild them from the fkeys the record carries, and on a partial record it reads that as “no parents”, detaches the node and answers success. So a cell edit travels as its own event and does not reuse the form’s, which does send autolink and may — the form hands it the whole record with its fkeys in it. Verified against a live treedb: a role edited in place kept its parent, and the parent kept it.A refused write puts the topic back to what the treedb has: leaving the typed value on screen is tolerable for a form, which stays open on the values it failed with, while a cell edited in place would just look saved.
The agent console hands its backends to the TreeDB GUI (
yunos-js:gui_agent0.9.0,gui_treedb0.10.0). Copying a dozenwss://host:portrows between two tabs of one browser was the alternative. The Schemas picker copies the yunos it shows as gui_treedb connections; that app’s Connections page pastes them (a clipboard read is not always allowed — Firefox refuses outright — so a refusal opens a box to paste into rather than failing).The picker already asked every yuno what services it runs, which is how it knows which ones hold a treedb, and threw the answer away one line after it arrived. What it lacked is the ENDPOINT, and one
view-configper yuno gives it:the port, from
global["<gate>.__json_config_variables__"].__top_url__. A yuno without it exposes no top gate and can never be a treedb backend, so its absence is the filter — a node with 17 yunos contributes the 2 that are reachable.the host, which is not in that url (the yuno binds
0.0.0.0, its realm binds127.0.0.1): the filename of__ssl_certificate__— the FQDN the certificate is issued for — else the realm id. Neither is guaranteed, so the url is a proposal and the connections arrive disabled.the service, which is the one NAMED like the yuno’s role. Taking “the first top-service” instead made all nine backends of a real scan read
authz, which is a connection the backend refuses.
A topic table you can read, and a graph you can search (JS submodules:
gobj-ui7.3.0,yunos-js—gui_treedb0.8.0). The treedb topic table had a single global search box while the read-only tranger browser sitting next to it had per-column filters, a column chooser and a CSV export: the richer table was the one that cannot write. It has all three now (with_header_filters,with_columns_button,with_export_button, all default on).The filter box is not put on every column. A hook holds children, a dict holds a subtree, and a date cell shows a formatted string over an epoch number — a text match against the raw value there is a box that lies, so those columns get none. A
booleangets a tristate tick, anenumthe list of its own values, and anfkeya box that stringifies the value first, because which rows point at X is the question fkey columns exist to answer. Search and header filters are separate layers: clearing the search does not drop the column filters. The CSV carries what the table HOLDS — loaded rows, visible columns, both filters applied — and not the topic, which is a server-side dump this view cannot stream.The graph gets a find box. It matches the term against a node’s label as well as its id — a topic keyed by
rowid/uuid/qualifiedis keyed by a counter or a path while the name a human knows sits in a secondary key — and it says how many it found, because a graph that did not move looks the same whether nothing matched or the match was already on screen.Searching in the table now crosses the FSM (
EV_SEARCH) instead of calling Tabulator straight from the DOM handler, where themachinetrace could not see the one action a reader performs most often.A legend for the topic colours, a minimap, and one click to take every service (
gobj-ui7.4.5,gui_treedb0.9.0). A node’s port colour encodes the topic it links to — a deliberate, functional cue, and nothing on screen said what any colour meant. The graph gets a legend strip, and it is a strip rather than an overlay because it is read AGAINST the graph. Its entries are buttons: one focuses its topic and travels up asEV_TOPIC_SELECTED, so the host turns it into the URL and what you are looking at stays linkable.A minimap appears from 30 nodes on (
minimap_min_nodes). One of a graph that already fits on screen is decoration, so it is not a preference to find and set. Its shapes are drawn by hand — a block in the topic’s colour — because G6’s minimap clones each element’s KEY SHAPE and these arehtmlnodes; at that scale a card is a rectangle anyway.In Connections, the browse column header selects every service of a connection at once. A yuno routinely exposes a dozen, and one click each was the only way there was; the header carries the state of the whole column (all / none / some) and flips it.
A treedb nobody has arranged opens laid out, not piled up (
gobj-ui7.5.7,gui_treedb0.9.2).manualmeans “leave every node where it was put”, and where none was put it means a cascade —get_default_ne_xy()walks x and y together, so a first open was a diagonal pile of cards, 126 of them ontreedb_system_schema, and the way out was knowing to pick a layout by hand. It opens in dagre unless most of its nodes carry a saved position; the pick is not persisted, so one dragged node makesmanualright again by itself.And it opens FITTED — but never zoomed out past the point where a card carries no readable text. dagre spreads a schema treedb (one topic, a hundred columns) far enough that fitting lands near 0.2 zoom: technically the whole graph, and a hairline. It stops at a legible zoom and centres, and the minimap such a graph always has says where that part is.
Two traps came out of this that are worth knowing beyond it, and both cost a wrong guess before the real value was pulled out:
gobj_write_attr()writes the private field of the same name. Apriv.xtherefore stops answering “did the HOST set attr x?” the moment anything resolves a default into it.After a graph is built,
record._geometryis not evidence of anything a human did.get_node_graph_props()returns the record’s own_geometrywhen it has one — and a treedb hands back{}for a record nobody moved — and thekw_get_int(..., KW_CREATE)that follow write the invented cascade coordinates into it. Every node of every treedb read as “placed”.
The record form says what is wrong, and says it in the app’s language (
gobj-ui7.7.1,gui_treedb0.10.2). Pressing Save on a form with a bad field did nothing at all: the branch that refuses the save wroteabort_closeandwarningon the click’s own kw, which nobody reads — that pair belongs to the close path. It callsreportValidity()now, which marks every bad field and puts the caret and the viewport on the first.The message a field did show was
input.validationMessage: the BROWSER’s sentence in the BROWSER’s locale, so a Spanish form on an English Firefox read “Please fill out this field.” The empty required field has its own key now; the rest still falls back to the browser, whose wording for a bad pattern or an out-of-range number is better than anything generic.And two dialogs of the editing flow were untranslatable:
yui_shell_confirm_*renders its message as an i18n KEY, so the English sentences passed in were keys nobody had defined — they render as themselves, in every language, and no locale validator can see a key that travels as data. New keys for consumers:this field is required,all changes will be lost.
Fixed¶
The amber highlight in the treedb graph had never appeared (
gobj-ui7.3.0).EV_FOCUS_TOPIChas been setting G6’sactiveelement state since the topic-cards landing, and that state is defined as an amberstrokeplus ahalo— both properties of a node’s KEY SHAPE. Every node in this graph is anhtmlnode, whose key shape is a DOM element, so neither property ever had anything to paint on: the topic focus centred the viewport and marked nothing. The highlight is drawn into the card’s own html now, repainting only the cards whose state changes, and a theme switch carries it across — that path rebuilds every card and would otherwise clear what is on screen.The row-count footer lied under a filter (
gobj-ui7.2.4 + 7.2.5). It read “5 Filas” over four visible rows.dataProcessedfires onsetData, not on a filter; and subscribing todataFilteredwas not enough either, because Tabulator dispatches that event from inside its ownfilter(), which only returns the surviving rows to the pipeline afterwards — sogetDataCount("active")answers the pre-filter set there. The footer takes the rows the event hands over. It went unnoticed while the only filter was the global search box.A filtered column had no room for its filter (
gobj-ui7.2.2 + 7.2.3). The table lays outfitDataFill, which sizes a column to its DATA, so aRolecolumn holding “root” came out narrower than the box it had just been given and its placeholder was cut to “filtrar c”. Text and list filters carry aminWidthof 150 — what the placeholder needs in the LONGEST locale, which is the measure that decides, since a column width cannot follow the language.The schema diagram keeps the size it is drawn at, and its arrows say who references whom (JS submodules:
gobj-ui7.0.1,yunos-js). Selecting Diagram inC_YUI_TREEDB_SCHEMAdrew the graph at its own scale and then, a frame later, zoomed it to the container: the reader saw one size and got another, and which one depended on how many topics the treedb had. The fit is gone — the canvas still follows the container when the view is shown, since G6 draws at 0x0 while hidden, but the camera no longer moves, so the scale and any pan/zoom survive.The arrowheads were also backwards. This view exists to draw a treedb the way its
.cliteral draws it in ASCII, and the.clands the arrowhead on the parent’s HOOK row: the reference belongs to the child’s fkey and points up at its parent, which is what the↖in the(↖)/[↖]/{↖}marks the same view prints has always said. The edge stays declared parent -> child, because that is what ranks the parent first under left-to-right dagre; only the marker moved.
Changed¶
maplibre-glmoves to6.4.1, and gobj-ui’s peer floor moves with it (gobj-ui7.0.0, a dependency-only major — no component API moved). 6.4.1 fixesDOM.sanitizeleaving dangerous attributes behind when several of them sit next to each other: it iterated a liveNamedNodeMapwhile removing from it, so every removal skipped the attribute right after it and anontogglecould survive the scrub. A floor is the only thing that stops a consumer from resolving to a version without that fix, so the floor is what moves.yunos/js/gui_treedbdeclares maplibre itself and moves too.Not affected: the v1 line.
estadodelaireandhidrauliaare the two SPAs that actually mount a map in production and they are onmaplibre-gl^5.24.0, which is the last 5.x there is — the fix is not backported, so the only way out of it is the migration to v2.
7.14.0¶
A schema stops being read by its rows. The order of the columns is part of the schema and no longer the order the filesystem happens to hand the projection back in; and the console that edits schemas stops showing the three topics they are STORED in and shows the schema they ARE — and stops losing your place in it.
The consoles also start saying WHICH machine you are looking at, in the two places that never said it: a treedb tab named after a treedb that lives on four backends, and a collapsed node hiding the yunos it has open.
Added¶
The Schemas workspace of the agent console edits a schema AS A SCHEMA (JS submodules:
gobj-ui6.3.0,yunos-js). Every schema a yuno holds lives in itstreedb_system_schema, stored as data in three flat topics linked by fkeys —treedbs->topics->cols. That is the right storage and it was the whole screen: adding one column to one topic meant finding it in a table holding every column of every topic of every treedb the yuno has, composing the parent fkey by hand, and remembering to raise atopic_versionthat nothing asks about.The new
C_YUI_SCHEMA_EDITORputs the schema back together — treedb -> topics -> columns in declared order, reordered by dragging, flags as checkboxes that say what they do, the schema DRAWN from the records being edited, a check of what the treedb would refuse, an export as the C literal and an import shown as a plan before it runs.Both versions travel with every write, and the operator is asked to remember neither:
topic_version, without which the persistedtopic_cols.jsonmasks the edit — the restart succeeds and nothing moved — andschema_version, which is what publishes the schema as a whole. Raising it is safe becausereconcile_treedb_schema()comparesc_schema_version, the version of the LITERAL, precisely so that an edit made there survives every start until a newer literal arrives.No C changed for this. It is listed here because the two JS submodules move with the release, and because the behaviour it depends on is the kernel’s:
autolinkon a partialupdate-noderewrites a node’s links from the fkey fields the record carries, finds none, and reads that as “no parents” — the editor learned it by detaching a topic from its treedb on a dev node, with the write answering success.Both consoles say WHICH machine you are looking at (JS submodules:
gobj-ui6.3.0,yunos-js). The same blind spot in two shapes, and the same cost: an edit made on the node you did not mean.A gui_treedb tab is labelled with the TREEDB name, and a treedb name is not unique across backends — two tabs reading
treedb_authzsare two different machines.C_YUI_TREEDB_TOPICStakes asource_urland prints it in its toolbar, which is where it fits: a tab wide enough forwss://host:1996is a tab bar with room for one tab.In the agent console what is open in Statistics and Schemas is a YUNO, and the checkbox that says so is on a child row — invisible while the node is collapsed, which is how a node is read most of the time. The node row now carries the labels of its open yunos, read from the SELECTION and not from the loaded children, because it is asked exactly when the children are not on screen.
Fixed¶
system-schemaanswered with a validator, not with the schema. The command ofC_NODEandC_YUNOreturned_treedb_create_topic_cols_desc()— the descriptor a user column is validated against, which is DERIVED from thecolstopic oftreedb_system_schema: one topic of three, with the storage-only fields dropped andvaluerenamed back toid. What it claims to answer is the meta-schema, andC_NODEalready publishes the schema of its own treedb throughdesc/descs, so there was nothing else it could mean. It now returns the whole thing —treedbs->topics->colsplus theschema_versionthat says how a schema is stored — and it reports a parse failure instead of answering an empty success.The literal is handed over by the new
treedb_create_system_schema()(tr_treedb.h), which is where the parse lived three times: the column descriptor andc_treedb’s materialization of the__system__treedb now go through it too.The consoles stop losing your place (JS submodules:
gobj-ui6.3.0,yunos-js). Two reports from the agent console’s Schemas workspace, and the same shape twice:A strip of treedbs dropped what each one had open. Open a topic, move to a sibling treedb, come back — its cards, with browser Back the only way to the table that was there. A
C_YUI_NODEnav item pointed at the canonical route of its child; with the newremember_positionit points at the tail last active under it. It stays a REAL position, which is why this belongs to the tree and not to a viewer restoring itself: clicking is a navigation like any other and nothing argues with the url.Shell tabs had it too, and there the same trick is impossible: an item’s route is where the tab’s view is MOUNTED and what a deep link resolves to, so moving it would move the mount. The position is replayed when the tab is ENTERED AGAIN — never when you arrive at the root of the tab you were already in, which is the view’s own way OUT of a topic. gui_agent already did this for its node tabs; gui_treedb does it now for its treedb tabs, with the decision in a tested function.
The picker was ERASING the position it should have replayed. A tab replays where it was left when it is entered again, and “again” was decided by whether the previous route was another tab — but the picker left that decision by its own early return, so the tab it had left behind stayed on record as the one we were in. Going to Select and back read exactly like walking up out of a topic: no replay, and the root recorded over the position. The tab we were in is now consumed at the one point every route passes through and written again only by the tab branch, so a route that forgets to clear it cannot exist.
Also: a yuno tab names its node (
yuneta_agent · wattyzer). Every node runs ayuneta_agent, so two tabs of two nodes read the same and there was no way to know which one you were typing into.A table paints its columns in the order the schema declares them again. Since 7.13.0 a treedb opens from its
__system__projection, and the projection is a treedb like any other: its nodes load in the orderreaddir()returns the key directories, which on ext4 is hash order. That order became the order ofcolsin the rebuilt schema, then got FROZEN into each topic’stopic_cols.jsonthe next timetopic_versionrose, and from there into every table:list-yunoscame back with_channel_gobjfirst andidthird, in a different scramble on each node.Four things had to change for an order to survive:
The order is stored.
topicsandcolsof__system__carry anorder, stamped by the projector from the position the node occupies in the schema compiled in C. It is storage, not something a column may declare:_treedb_create_topic_cols_desc()keeps it out of the descriptor a user column answers to, andget_treedb_schema()takes it away again on the way out — in a schema the order IS the sequence of the dict, and a schema carrying both would hand every topic a column attribute nobody wrote.The rebuild sorts by it, and falls back to where the schema compiled in C declares the node when the projection cannot place it — one projected before the index existed — so a store that has not been re-projected yet already comes back in order.
orderdefaults to 9999: a column created in__system__by hand says nothing about where it goes, and what says nothing goes last.The keys load sorted (
find_keys_in_disk()).readdir()order was never a contract: two replicas of the same store read it back differently, and so did the same store twice.A topic_cols.json that differs from the schema ONLY in the order is rewritten, instead of waiting for a
topic_versionbump. The file is deliberately frozen until the version rises — that is what stops a change to WHAT a column declares from arriving unannounced — and a change to the order announces nothing new.
A phone stopped cutting the node list in half (
yunos-js). Tabulator ends a cell that does not fit with an ellipsis, andfitColumnsleaves the status column about half of what “Ejecutando Solo lectura” needs. Name and status wrap now. Two things were needed:variableHeighton the columns, so the table grows the rows, andheight: autoon those cells, because Tabulator measures once and writes that height into every cell — a column that gets narrower afterwards (a rotated phone, a resized window) wraps into a height nobody re-measured and loses the second line.gui_treedb opens on the work. Its rail is Topics / Graphs / Connections, the backends last: they are where you go when one has to be added or fixed, which is not what a session is spent doing.
/connectionsis the same route — linkable, in the site map, and the landing is unchanged.Meta-schema at
schema_version15 (topicstopic_version 7,cols9), which re-projects every store on the next start and rewrites everytopic_cols.jsonwith the order the schema declares.
7.13.2¶
A key that says where it belongs: the schema stored as data stops being addressed by a counter, and starts being addressed by its own name.
Changed¶
BREAKING (stored data):
topicsandcolsof__system__are keyed by the QUALIFIED name, not by a rowid. The id of a node is now the id of its parent, a dot, and its own name —treedb_yunovatioscodb.yunosfor a topic,treedb_yunovatioscodb.yunos.yuno_rolefor a column.The collision it has to avoid is real: a topic name is unique only inside its treedb, so keyed by the bare name the second treedb of a store declaring
userscollided with the first and lost its schema. A rowid handed out fromtranger2_topic_size() + 1avoided it and paid for it — it does not reproduce, so two projections of the same schema do not align;find_col_id()andfind_topic_node()had to be linear scans overvaluebecause the key said nothing; and a rowid pkey has no update, so an editor saving a column appended a second one under the same name instead of changing it.The separator is a dot and cannot be
^. That is the character an fkey reference is split on —decode_parent_ref()requires exactlyparent_topic^parent_id^hook— so a column keyedtreedb^topic^colmakes every reference to itself undecodable.A new
idflag,qualified, besideuuidandrowid. A create that sends noidgets one composed from the parent named in its fkey and the value of the topic’s first secondary key (build_qualified_id()intr_treedb.c). So an editor creates a column the way it creates any other record, and the rule lives in the store instead of in each client. Declarable by any schema;tr_treedb.handlib_treedb.jscarry the vocabulary.Meta-schema at
schema_version14 (topicstopic_version 6,cols8), which re-projects every store on the next start.migrate_schema_ids_to_qualified()runs first and moves a projection made with rowid keys, node by node: it rewrites each under its qualified id with its content intact — an operator’s addition is in there too and is not the projector’s to drop — and deletes the old one, columns before their topic because deleting a parent only UNLINKS its children. It is the one delete the projector does, and it happens once per store.An id that does not fit is refused, never trimmed. Both composers built the qualified id with a bare
snprintfinto a buffer of one record key, and both halves can be a full key on their own. The id is the ADDRESS of the node, so a truncated one silently addresses something else — two long names sharing a prefix would land on the same record. They read the return value now, log, and answer nothing; every call site skips that topic or column. GCC saw only one of the two:-Wformat-truncationcould bound the projector’s buffers and not the ones fed by a json string.Published: gobj-js
7.13.2and gobj-ui6.0.0. gobj-js ships ahead of the SDK release it belongs to, at the number the SDK will catch up to (the same way7.6.7did):qualifiedjoinstreedb_field_types, the list that turns a column FLAG into thetypeevery form and table switches on. gobj-ui 6.0.0 is a dependency-only major — no API moved — that raises its peer floor to gobj-js^7.13.2, because on an older runtime the word is missing and the qualified paths fail quietly. The in-repo SPAs (gui_agent,gui_treedb) take both.gobj-ui labels a
qualifiedkey by the secondary key too (treedb_node_label.js), the same way it already did forrowidanduuid: the id names the record but names every ancestor with it, and the card wants the leaf. The pkey stays as the tooltip. The form treats a qualified pkey as a key the store hands out — never typed on create — and an existing row opens for update, which is what stops an edit from appending a twin.
7.13.1¶
A schema you can read: the descriptor learns to say what NAMES a record, and the two treedb views stop showing storage where a schema was asked for.
Added¶
A topic descriptor now carries
pkey2s.tranger2_topic_desc()clonedtopic_name,pkey,tkey,system_flag,topic_versionandcols— so a viewer of a treedb learned which column is the primary key and never which one is the SECONDARY key.That is the difference between identifying a record and naming it. A topic whose id column is flagged
rowidoruuidkeys its records by a value nobody reads, and the name a human knows them by lives in itspkey2scolumn:treedb_system_schemakeystopicsandcolsby rowid and holds the topic/column name invalue. Withoutpkey2sin the descriptor the only thing a GUI could print was the rowid, which is how the agent console’s graph came to draw cards reading181,225,193.Additive and safe on both sides:
kw_clone_by_path()skips a key a topic does not have, and the reader (gobj-ui 5.17.0) falls back to the id when the descriptor does not carry the field, which is what an older node answers.diff-schema: what the stored schema says that the schema in C does not. A treedb opens from its projection in__system__, the projector never deletes, and a re-projection publishes undermax(stored, literal) + 1. So the three numbers of thetreedbsnode say that SOMETHING was published and never what:schema_version24 overc_schema_version23 is the shape of an operator edit and the shape of a plain re-projection alike. The new command ofC_TREEDBtells the two apart.ycommand -c 'command-agent service=treedbs command=diff-schema \ treedb_name=treedb_yuneta_agent'It answers one row per difference (
treedb,kind,topic,col,attr,stored,from_c) plus a comment with the stored versions, the ones the running yuno carries and the count.kindischanged,only_in_stored(an operator addition, or a topic a later schema dropped and the upsert kept),only_in_c(declared and never projected) orversion.Two rules keep the answer readable, and both are about the store rather than about schemas. A record carries EVERY column of its topic filled with the empty value of its type, so an attribute nobody wrote is stored as
"",{},[]or0and is not a difference — read as one, those defaults buried 6 real differences under 592. And the version stamps are not compared as content, since the projector raises them itself: only the anomaly is reported, a projection made from another release or a topic the re-projection never reached.The comparison is projection against projection: the schema from C is projected in memory through the same two builders the projector uses (
build_topic_projection()/build_col_projection(), extracted fromupsert_treedb_schema()for this), so no default the projector fills in can read as a difference nobody made, and the two paths cannot drift apart.C_TREEDBnow keeps the schema each treedb was opened with, which is the other half of the comparison; a treedb opened from its projection alone has nothing to compare with and the command says so.yunos/js: gui_agent asks it from the Schemas workspace. ADifferencesbutton beside Apply asks eachC_TREEDBservice of the yuno and shows the answer as a table, with what the command says about the versions above it. A service that refuses, or that is too old to know the command, says so in the same dialog beside the ones that answered. Detail in that submodule’s ownCHANGELOG.md.Note it takes the
readpermission of theC_TREEDBservice, which is not theC_NODEservice of the treedb: an identity that reads a schema can still be refused here.gobj-ui submodule 5.16.0 → 5.17.0: the treedb views become readable on a schema. Two views, two different failures, one release.
C_YUI_TREEDB_SCHEMA— the/schemalanding of a treedb — drew one 40px circle per topic with its name underneath, which answered neither of the questions a schema is opened for: what a topic holds, and what links to what. It now draws what the.cliterals draw in ASCII: one CARD per topic listing its fields in schema order, and one edge per hook, leaving the row that declares the hook and landing on the fkey row of the child it names. The marks are the notation of those literals ({}[]()(↖)[↖]{↖}*#), so the drawing and the source read the same. Both ends of an edge come from the declaration —'hook': {'yunos': 'realm_id'}names them — and a self-referent hook draws as the loop it is.C_G6_NODES_TREE— the RECORD graph — labelled every card withrecord.id, which ontreedb_system_schemameant a fan of cards reading181,225,193. It now reads the pkey column’s flags and, when the key is synthetic, labels by the secondary key, keeping the pkey as the tooltip. That is what thepkey2sabove is for.The two answer different questions and the confusion between them is the whole story: the node graph draws RECORDS, so on a treedb whose records ARE schemas it draws a box per column — a correct picture of the storage and an unreadable picture of the schema.
yunos/jsfollows:gui_agentandgui_treedbat^5.17.0.
Changed¶
Every JSON file the framework writes is indented with FOUR spaces, not two.
save_json_to_file()is the one helper behind all of them — persistent attrs, configs, whatever a gclass saves — so the change reaches every file at once, and it puts the C side on the rule the rest of the project already follows: one level of structure is four characters.Only whitespace moves. Nothing reads these files by column, and jansson parses either shape, so an old file stays readable and is rewritten with the new indent the first time something saves it.
7.13.0¶
A schema stops being something only the C literal knows.
The __system__ treedb holds it as DATA again — projected on every open,
reconciled by version, and able to rebuild the schema a treedb opens with — and
gui_agent grows the Schemas workspace that edits it: the treedbs of any
yuno of any node, over the one control-center session the console already has,
as tables or as a G6 graph. The node’s own agent is one of the entries.
The other half is honesty about who may write. Only the MASTER of a treedb’s
tranger can, and until now a replica answered success and lost the row at
the next reload; it refuses now, treedb-info says which one a yuno is, and
the editor opens read-only instead of turning every click into a toast.
And a family of small lies found by pulling one thread: every authzs command
answered with its permission list in the COMMENT field — which the client
refuses — four of them answered nothing at all, and the one helper behind them
leaked a whole kw per call.
Fixed¶
Eleven full-path buffers were sized with
NAME_MAX.NAME_MAXis 255 and documents a FILENAME; these were filled with a whole path bybuild_path()oryuneta_realm_file(), so a long realm, role or config name truncates. Nothing was silent —build_path()logs the overflow — but the write failed.The one that started it:
dbsimple.c’ssave_json()usedNAME_MAXfor the same pathload_json()reads withPATH_MAX, so a yuno could load its persistent attrs and fail to save them. Alsoentry_point.c(the file log handler),c_resource2.c,c_agent.c(audit file, yuno data dir, two temp files) and five callers each inc_ycommand.candc_cli.c.get_persist_filename()keeps itsNAME_MAXbuffer: that one really is a filename.gobj_build_authzs_doc()never releasedkw, and everyauthzsleaked one. Its siblinggobj_build_cmds_doc()decrefs on all three of its return paths; this one decrefs on none of its four — while all fifteen callersKW_INCREF(kw)before calling, precisely because they believe it consumes one. Neither header said// owned, which is how the two helpers drifted apart.It now consumes
kw, with a single exit:authzandserviceare BORROWED out of the kw and used byauthzs_list()and by the error messages, so the decref cannot happen before the last of them.Measured, not argued: 20
authzsagainst a treedb service and an orderlykill-yuno, reading thegbmemaudit at shutdown — 753 leaked blocks before, 0 after. This is what the “system memory not free” line at the end of a yuno’s log has been reporting.authzsanswered with the permission list in the COMMENT field, andycommandrefused it.msg_iev_build_response()takes(result, jn_comment, jn_schema, jn_data, kw)and every one of them is ajson_t *, so the compiler cannot tell them apart — butjn_commentis a STRING (ycommandreads it withkw_get_strand prints it with%s), while the authzs doc is a list. Everyauthzscommand therefore loggedERROR: ... "function": "kw_get_str", "msg": "path MUST BE a json str", "path": "comment"with a full stack trace, and printed no permissions at all.
cmd_helpis NOT affected and was left alone:gobj_build_cmds_doc()returns a json string, which is exactly what a comment is.Eleven copies of the same three lines —
c_agent,c_authz,c_node,c_task,c_tranger,c_treedb,c_postgres,c_pepon,c_testonand the two timer tests.C_PROT_MQTT’slist-topics,list-clients,list-usersandcreate-userhad it too, with the resource list in the same wrong slot.The SKELETONS shipped the broken shape.
gclass_serviceandyuno_citizenboth scaffolded the barereturn gobj_build_authzs_doc(...)— which is where the four gclasses above got it. Fixed at the source, with the reason in the template so the next reader does not have to find this entry.Four
cmd_authzsanswered nothing at all.C_IDP_KEYCLOAK,C_MQTT_BROKER,C_CONTROLCENTERandC_WEBSTATSreturnedgobj_build_authzs_doc()raw, andcommand_parser()hands a handler’s return value straight back with no wrapping. That is not a display problem:msg_iev_build_response()is also what sets__md_iev__back on the answer, so a bare doc has nothing to route it to the requester and the command hung until the caller gave up — verified live, 25 s with no answer against an unpatched yuno, correct output against the patched one.
TreeDB¶
C_NODEanswerstreedb-info, and refuses writes on a replica. Two halves of the same hole.The master flag was not reachable from the control plane: it is an
SDF_RDattr of the tranger, absent fromservices, fromtreedbsand from the stats, so a client wanting to know whether a treedb can be edited had to fetch the entireprint-trangerdump — megabytes to read one boolean.treedb-infoanswers{treedb_name, master, schema_version, topics}— the topics BY NAME, and the__schema_version__the treedb was built with, which is what tells a client whether the schema it is looking at is the one it knows. It is also per TREEDB, not per yuno: the same yuno is routinely master of itstreedb_system_schemaand a replica of a data treedb. The role comes from the config — only a yuno configured as master opens the store in exclusive mode, and one configuredmaster: falsenever competes for the lock, so the start order does not change it.And writing to a replica used to answer success.
create-node,update-node,delete-node,link-nodes,unlink-nodesandimport-dbbuilt the node in the in-memory treedb and returned it; the append was never attempted, sotranger2_append_record’s own “NO master” guard never fired and nothing was logged; the row was gone at the next reload. They now answerERROR -1: <yuno>: treedb '<name>' is READ-ONLY, this yuno is not the master of its trangerchecked BEFORE the authz check, because on a replica nobody can write whoever they are, and a
-403sends the operator after a permission that would not help. Reads are untouched — a replica exists to be read.Verified against two live yunos sharing one store, one master and one replica: the four write commands refused on the replica, reads kept working, writes on the master still persisted to disk. 123/123 ctest.
Not covered, and deliberately: the snap commands (
shoot-snap,activate-snap,deactivate-snap) also write, and their master semantics were not verified here.
Documentation¶
The JavaScript side had no API reference — it had a listing. Looking for
gobj_post_event()in the docs found nothing, and the function had been in the JS runtime for releases. It was on the site, as one line inside a code block onapi/js/events.md, with no signature and no anchor: the JS pages carried zero(name)=labels and the appendix index was C-only (2540 C functions, 0 JS symbols), so no JS name was reachable by search. The whole JS reference was 790 lines against 31 608 for C, written once in April and touched since only by the STE sweep and a link fix.The one line documenting
gobj_post_event()was also wrong: it said “next microtask”, and the queue has drained withsetTimeout(…, 0)— a macrotask, so the browser keeps its turns — since the JS side aligned with the C contract in gobj-js 7.10.0.gobj_deliver_posted_events(),gobj_posted_events_size(), the 10 000-message ceiling and the purge-on-destroy were all undocumented.Both packages now have a real reference, every symbol anchored and linked to its source: gobj-js 255/255 and the gobj-ui public surface 92/92. New pages for what was missing entirely — the whole trace API (
gobj_set_gclass_trace()and the silencing side, which is the tool the project calls an axiom for debugging), theSDATACM/SDATAPM/SDATAAUTHZcommand descriptors, thekw/kwid/msg_ievfamilies, the DOM and i18n helpers, thejdbdatabase — and a first gobj-ui section: the shell API, the dialogs, the component gclasses, the period algebra, the theme and the dev panel.scripts/verify_js_api_coverage.py— the gate that stops this from happening again, the JS half ofverify_api_coverage.py. It reads every export of both submodule packages, compares them against the(js_<name>)=anchors, and reports MISSING and STALE.--writegeneratesapi/appendix_js_api_index.md, which lists all 439 symbols with their signature, module and a source link — so a search finds a symbol whether or not it has a reference entry yet. The links are pinned to each submodule’s own tag, read from itspackage.json: gobj-js is at 7.10.0 and gobj-ui at 5.11.0, and pinning either to the SDK tag points at a tag that repository does not have. The script warns when a submodule HEAD is not its tag, because then every#Lanchor is a guess.
Added¶
The
__system__treedb holds a schema again. EveryC_TREEDBservice builds a__system__treedb whose topics (treedbs→topics→cols) exist to hold a treedb schema as data — the V6 design that makes a schema listable and editable at runtime through the ordinary node commands. It has been created empty on every start and never filled: each of the 13 stores on the dev machine hadtreedbs/keys,topics/keysandcols/keyswith zero records, and nothing could have filled them, for two reasons in the meta-schema itself.Its
colstopic had lost thevaluecolumn while keeping'pkey2s': 'value', a secondary index on a column that no longer existed; andcols.idhad lost itsrowidflag, sotreedb_create_node()refused every column with “Field ‘id’ required”. Both were collateral of merging two descriptors that are not the same object: the storage schema of a column (keyed by rowid, name invalue, because a column name is unique only inside its topic) and the validator for a user column (keyed by name). The diagram at the top of the file still documented the original shape —id (rowid),* value (2)— which is what it was checked against.Both are restored,
_treedb_create_topic_cols_desc()now derives the validator from the storage schema (renamesvalue→id, drops the storage-only fields) instead of copying it, andtopicsgainedsystem_topicso the flag survives the round trip. System schema 7 → 8. Since the three topics were empty everywhere, the regeneration the bump forces has nothing to lose.open-treedbprojects a treedb’s schema into__system__the first time it sees it, under the defaultuse_internal_schema=1too — the C literal stays the source of truth, and the schema becomes visible. Withuse_internal_schema=0the projection is what the treedb opens with; the deadreturnthat shadowed that path inget_client_treedb_schema()(it was in V6 as well) is gone.The
treedbsnode carries two version numbers, because one counter cannot serve two writers:schema_versionis what the schema is worth totreedb_open_dband whoever edits the schema raises it, whilec_schema_versionrecords which literal the projection came from and only the projection writes it. Reconciliation compares against the second. With a single counter, the first edit made here — which has to raise the version to reach the treedb at all — would silently outrank every later release of the literal. A re-projection publishes undermax(stored, literal) + 1, or the persisted schema file, sitting at the edited number, would keep masking it. System schema 8 → 9.Afterwards the two homes reconcile by
schema_version, strictly newer wins — the same ruletreedb_open_dbalready applies between that schema and the persisted schema file, so an edit made in__system__survives every start until a higher version arrives from C. Reconciling is an upsert; nothing is ever deleted. A column’sidis a rowid handed out from the topic size, so re-creating columns renumbers all of them and can hand a retired number to a different column — and a delete is the one destructive primitive of the store: it drops the schema’s own history, which is the reason to keep a schema in a treedb at all, and it refuses a snapshot-tagged node, so re-projection would fail outright on any store that has ever been snapped. An update appends a new version instead, and what a column used to declare stays readable withinstances.And a write to those topics is a schema change, so it answers to the rules of a schema: a column is checked against the descriptor a user column answers to (the one
parse_schema_cols()applies at open, so the failure lands on the writer instead of on the next open);pkeymust beidandsystem_flagsf_string_key, whichtreedb_open_db()otherwise answers by silently dropping the topic;pkey/tkey/system_flagcannot change once the topic exists, becausetopic_desc.jsonis never rewritten and the change would be stored, shown by every reader, and ignored by the topic for good; and two columns with the same name in one topic are refused when the column is linked, which is when the clash becomes real.New test
tests/c/c_treedb_system_schemacovers the three steps: project; delete the schema file and re-open, and the topics and columns that come back are the ones that went in; then move the schema forward and check the projection updates while the existing columns keep their rowid. Docs:YUNO_TREEDB.md§3.11.gobj-ui submodule 5.11.1 → 5.12.0:
C_YUI_TREEDB_TOPIC_WITH_FORMshows the topic’s schema. A toolbar button (with_schema_button, on by default) opens that topic’sdesc— pkey, cols, types, flags and fkey targets — in the standardized adaptive dialog, on the lazy JSON viewer. The table shows the data; nothing showed the contract the data answers to, which is what you need in front of you when a value is refused or a link does not appear, and reaching it meant reading the backend’s schema by hand.register_c_yui_json()became idempotent along the way, the courtesyregister_c_yui_form()already had: the topic view auto-registers the JSON viewer for its dialog, and every consumer registers it explicitly after the topic view — an order that would otherwise trip “GClass ALREADY created” on boot.yunos/jstakes it ingui_treedb, built and deployed. The JS API reference is repinned to the new tag withscripts/verify_js_api_coverage.py --repin(105 links, 2 line anchors) and its appendix index regenerated.gobj-ui submodule 5.12.0 → 5.13.0: the JSON of a treedb cell is one click away. A col that holds a JSON document —
dict,list,object,array,blob,template,coordinates,gbuffer— has only ever shown the first 20 characters ofJSON.stringify()in its cell, which for the fields carrying the actual configuration of a node is a preview of the opening brace. Reading the value meant opening the edit form: edition mode, a raw text editor, and a dialog whose purpose is to change the record.Clicking the cell now opens the whole value in the same adaptive dialog the schema button uses, on a hosted
C_YUI_JSON— read-only, collapsed and searchable. The record is already in the table, so the dialog issues no command and touches no backend. The click crosses the machine (EV_SHOW_CELL_JSON {row_id, col_id}) like every other action of this view, and the kw carries the identity of the cell and never its value: the trace dumps the kw. A cell whose document is empty gets no link, and that absence is what makes the click a no-op.The preview is built as a DOM node (
JSON_CELL/JSON_CELL_ICON/JSON_CELL_PREVIEW) instead of the bare string the formatter used to return, so record data is no longer parsed as markup on its way into the cell.Every v2 consumer took it and is deployed:
gui_treedb(artgins.ytreedb.com), wattyzer (app.wattyzer.com, which was two releases behind and also gains the schema button), both yunovatios GUIs, andgui_agenton both control-center planes — that one mounts no treedb table, so what it takes is 5.12.0’s idempotentregister_c_yui_json(). The two v1 apps, estadodelaire and hidraulia, stay on the frozen1.0.1line by design.Three of those apps had to define the keys the library now asks of their own i18next, which their prebuild validators caught:
show jsoneverywhere, plus the JSON viewer’stoo many rows; collapse some branchesin wattyzer, reachable there since the viewer entered the treedb stack. An undefined key renders as the key itself and never changes language — the failure that is invisible by construction.
Fixed¶
use_internal_schemais gone: a treedb opens from its projection, always. The flag chose between the C literal and the__system__projection, defaulted to 1, and was read-only in most yunos — so an edit made in__system__reached nothing anywhere until every yuno’s config was changed one by one, which is the difference between a feature and a demo. It distinguishes nothing now: the projection is seeded from the literal and re-made whenever the literal or the projector moves ahead, so opening from it is opening from the literal until somebody edits it. The literal stays as the fallback for a projection that cannot be rebuilt into a valid schema. Removed fromC_TREEDBand from the yunos of this tree; the project yunos that still pass it keep working untouched, becausecommand_parsermerges an unknown key and nobody reads it.Opening every treedb from its projection surfaced one more thing the flag had been hiding, caught by the mqtt ACL test: a column with no
defaultis stored with an empty one, because the attribute is a blob and a blob with no value is{}. Left in the rebuilt schema that empty dict is a real default, and creating a record without that field handed{}to a string column — “Value must be string”, from a schema nobody wrote that way. It is dropped now, but only for a column that cannot hold a container:[]is a legitimate default for an array column, and mqtt_broker’spublish_acl/subscribe_acldeclare exactly that.Two treedbs of one store could not both hold a schema.
__system__keyed itstopicstopic by the bare topic name, and a topic name is unique only inside its treedb:usersis a topic ofauthzs,mqtt_brokerandcontrolcenteralike. The second one to be projected got “Node already exists”, its topic node was never created, its hooks then failed with “link topic not found”, its rebuilt schema failedparse_schema, andget_client_treedb_schemafell back to the C literal — in silence. The local controlador’s store holds five treedbs and its projection held one.topicsis keyed by rowid now, with the name in the pkey2value, exactly ascolshas been since V6 — the same flaw, solved for columns and never for topics. The sweep that found it now holds all four schemas in one store, which is what used to break. Addressing a topic costs what addressing a column costs: its rowid, not its name. System schema 12 → 13.A change to a schema now publishes itself. Raising
topic_versionandschema_versionis what makes a change visible — forget either and the change does nothing and says nothing, becausetreedb_open_dbkeeps the persisted schema file on a tie and tranger2 keepstopic_cols.jsonunless the incoming version is higher. Leaving that to whoever writes means every editor, script and console carries the rule; the author of this code got it wrong three times in a row while debugging, knowing it. So writing acolsortopicsnode of__system__raises the versions that publish it, walking up the fkeys to the column’s topic and its treedb. A new column publishes when it is linked to its topic, which is when it becomes part of the schema — at create time it has no topic yet. The projector sets the versions itself and marks the tranger while it works, which is also what stops the rule from answering its own writes.The projection of a schema was losing three column attributes, and a scalar
default. Measured across the 19 schemas in the tree and the project repos:enum(18 uses) andtemplate(6) had no column in the meta-schema at all, andpkey2shad one that the projector never filled. Theenumloss is the worst of the three because it is silent and disarming: the column keeps itsenumflag while its enumeration disappears, so it declares an enumeration it no longer has and every value passes. On top of that, a column’sdefaultis declaredblob, and a blob replaced anything that was not an array or an object with{}— so'default': 'es'was stored as an empty dict. A blob that discards a scalar is not a blob; it keeps what it is given now, on both the write and the read back.The cause of the missing attributes was a list written by hand in the projector: adding an attribute to the descriptor did not add it there. It now copies whatever the descriptor declares, and the test asserts fidelity the same way — reading the attribute list from the descriptor instead of repeating it — so the next attribute that gains no storage fails a test rather than changing every schema quietly.
Fixing the projector was not enough to fix the stores it had already written: reconciliation compared only the C literal, so a projection made by an older SDK stayed frozen — and an older SDK is exactly the one whose projection is missing what it did not know how to store. The
treedbsnode now also recordssystem_schema_version, the meta-schema version that produced it, and a projection made by an older one is re-made on the next start. Raising the meta-schema’s own version is the migration lever, and it moves when the projector changes even if no field does — 11 → 12 was exactly that, because 11 could write atopic_versionlower than the one already published, leaving a re-projection that fixed the columns in__system__and never reached the topic. A topic’s version cannot go backwards now, for the same reasonschema_versioncannot.Verified end to end on a real schema: the local yunovatios controlador, degraded by the lossy projector, recovered
language’s["es","en"]and its"es"default, andcluster’sfalse— a scalar boolean default that a blob used to replace with{}. System schema 9 → 12.close-treedbon a playing yuno was a use-after-free anyone could trigger. It destroys the treedb’sC_NODEand itsC_TRANGER, and an owner keeps raw handles no framework cleanup can reach: the service pointer, thetrangerjson_t read from it, copies of both on a hot path, and whatever else was opened on that same tranger —db_history_coopens itsmsg2db_alarmsthere. Freed underneath, the next record processed writes into released memory. Every in-tree consumer closes only frommt_pause, so the command now refuses while the yuno plays and points at the pair that does this correctly,pause-yuno+play-yuno(which reopens through the owner’s own lifecycle, without restarting the process).force=1remains for a caller that opened the treedb and holds nothing of it.A treedb
update-nodevalidated nothing, and answered success when it failed.treedb_update_node()set each incoming field straight onto the node — no type, nonotnull, noenum— so a column could end up holding what its own schema forbids, andmt_update_node()then dropped the return value and answered the collapsed view of the unchanged node, so a refusal read as a success all the way to the client. Updates now run the same normalization as creates, validate every field before touching the node (a refusal leaves nothing half-applied), and the failure reachescmd_update_node, which already knew how to report it.enumwas decorative on writes. A column’senumlist was checked only when a schema was parsed (check_desc_field), never when a node was written:normalize_node_field_value()'senumcase only looked at whether the container was a string or an array. So the promise did not survive the first write, on any treedb. Checked now on both paths, for the supplied value only — an absent optional field is stored as""or[], which no enum has to list. Audited before shipping: 122 enum columns across the local stores, 0 values outside their list.C_NODE’smt_node_treereleased its options before reading them.JSON_DECREF(jn_options)also nulls the variable, and thewith_metadatalookup came two lines later — so the option never took effect, and every call logged “kw must be list or dict”. Reading a node tree with metadata was impossible, silently.get_treedb_schema()askedgobj_node_tree()whether a treedb had a schema in__system__yet, and a first open — the ordinary answer — logged “Node not found” as an error. It asks with a list now, which is silent when empty.Landing page: carded yunomusica.com as the second live example — a shipped SPA on
C_YUI_SHELL+C_YUI_NAV, installable and offline, next to the shell demo that exists to show the shell.gobj-ui submodule 5.11.0 → 5.11.1, which exports
yui_shell_confirm_danger()from the package root.confirm_ok,confirm_yesnoandconfirm_yesnocancelwere all in the barrel and the destructive one was not, so the dialog the library’s own comment tells you to use for “this deletes an account” — red button, safe answer last — was the single one thatimport { … } from "@yuneta/gobj-ui"resolved toundefined. Found while writing the gobj-ui reference above. A missing export fails at the call and not at the import, so it read as a bug in the caller, and the only consumer had reached for the deep import and moved on. Deep imports keep working, so nothing breaks.Both JS packages track their
package-lock.jsonnow. Each one ignored it, so a build of either was reproducible only by accident: every machine resolved its own tree from the ranges inpackage.json, and the tree behind a publisheddist/was whatever that disk held that day. Nothing recorded it. The lockfile is never published — npm keeps it out of the tarball, and it is in neither package’sfiles— so this changes what a developer and a CI run get, not what a consumer gets. Both verified with a realnpm cifrom the committed lock in a clean directory: 76 packages for gobj-js, 251 for gobj-ui.The staleness warning of
verify_js_api_coverage.pyhad to change with it. It compared the submodule HEAD against its tag, so a lockfile commit made it fire for ever — and a warning that everybody learns to ignore is worse than no warning. It now diffs the documented files (src/,index.js) between the tag and the working tree, which is both quieter and stricter: a commit that cannot move a line number says nothing, and an uncommitted edit that can move one — which the old check missed entirely, because it compared two commits — now reports.Every JS consumer refreshed, by a rule instead of by hand. Eleven packages across nine repos moved their floors to
^5.11.1for gobj-ui and^7.10.0for gobj-js, plus maplibre-gl^6.3.0, vite^8.2.1and vitest^4.1.10where they applied. The rule was to raise each floor to what npm calls wanted — the newest version the existing range already accepted — and never to latest, so a major cannot enter through a sweep.That rule is what protects the two v1 apps: estadodelaire and hidraulia keep gobj-ui at
1.0.1(the frozen line), maplibre-gl at5.24.0(v6 is ESM-only and needs the worker wired) and vanilla-jsoneditor at0.23.8. All three are majors, and a major is a migration to plan, not a number to raise.The peer floors of gobj-ui are untouched for the mirror-image reason: a peer floor is a contract with every consumer, and
^6.1.0already accepts maplibre 6.3.0. Raising it would force every consumer up and cost a republish, for nothing.Verified by build, not by the number in the file: all eleven install, the nine with a build script build, and the two libraries pass 39 and 256 tests. Worth knowing — estadodelaire and hidraulia declare vitest and ship no test file, so the build is the only gate their bump passed, and they are the two that cost the most to repair.
--repinfor the JS docs. A submodule bump moves a tag, and every hand-writtenblob/<tag>/link into that repository goes stale with it.check_doc_line_refs.py --repincannot help — it matchesgithub.com/artgins/yunetas/blob/only — so the 5.11.1 bump would have left 104 links pointing at 5.11.0 with the pages looking correct. The JS script now retags them and recomputes each#Lanchor from the symbol that the entry’s own(js_<name>)=label names, which is more reliable than the link text: a heading is free to read## C_TIMERwhile the symbol isregister_c_timer().The agent dropped every counter-driven answer of a client behind a controlcenter.
kill-yuno,run-yuno,play-yunoandpause-yunodo not answer when the command is parsed: they answer when it is DONE — the agent raises aC_COUNTERthat waits for the killed yuno’s channel to close, or for the launched one to connect back, andac_final_count()sends the answer to the requester. It looked for that requester withgobj_child_by_name(__input_side__, …)only, and a command cascaded from a controlcenter arrives through the OUTBOUNDcontrolcenterC_IEVENT_CLI — a top-level service, not an input-side child. So the lookup missed, the answer was dropped with “requester channel child not found” in the agent’s log, and the caller waited forever for work that had actually been done.This is the same defect
ac_command_yuno_answer/ac_stats_yuno_answerwere fixed for ind2d833109(7.11.0); the deferred path was missed because it is reached only by the four lifecycle commands, and only from a SPA — a localycommandrequester (“input-N”) IS an input-side child, so every CLI test passed. Same fallback here:gobj_find_service(requester, FALSE).Found while giving gui_agent’s Schemas workspace its Apply button, whose whole job is
kill-yuno→run-yuno play=0→play-yunochained on those answers. Nodes need the new agent binary for it to work there.gobj-ui submodule 5.15.0 → 5.16.0, and the JS yunos with it: the treedb GRAPH learns read-only and reports its writes. The whole write surface of
C_YUI_TREEDB_GRAPHhangs off ONE operation mode —editionis the only one that draws the create / delete / link affordances — soreadonlydrops it from the mode select, and because the mode is a PERSISTED preference, a graph left in edition on a master comes back inreadingon a replica. The five write events are refused as well: the G6 child raises them from its undo/redo history and from saving the node geometry, neither of which goes through the toolbar.It also publishes
EV_RECORD_WRITTEN, the event the topics editor has had since 5.14.0, because a host that edits a SCHEMA is not finished when the record is written — the yuno still has to be restarted to re-read it.__graphs__is excluded on purpose: the graph writes that topic itself on every layout save, and reporting it would say the schema changed because somebody dragged a node.And a third thing, which only a SECOND view on the same treedb could surface: an updated node the table never loaded threw an unhandled rejection. Tabulator’s
updateData()rejects on a row it cannot find and nobody awaits it, so it arrived as a bare “Update Error - Unable to find row” naming neither gclass nor topic.yunos/js: gui_treedb’s Connections moved from the account menu to the rail. The backends are where a session starts and both work entries depend on them, so they lead the rail — Connections / Topics / Graphs. The item declares a route and no target:/connectionsis already inshell.routesandbuild_item_index()fills it from there, so one place says which gclass the backends page is. The workspace picker (tab 0 of Topics/Graphs) is renamed Select in the same move, path included (/<ws>/connections→/<ws>/select): two different pages had carried the name Connections, and it only showed once both were on screen at once. Detail in that submodule’s ownCHANGELOG.md.yunos/js: gui_agent reaches the AGENT’s own treedbs. The agent never appears inlist-yunos— it is the daemon that answers it — so its treedbs (treedb_yuneta_agent,treedb_system_schema,treedb_authzs) were unreachable from the console. The Schemas picker offers it now under the sentinel yuno id__agent__, addressed with the agent’s owncommand-agent service=<treedb>; on the.ovhplane the same row is agent22. Apply is disabled there — the agent is not a managed yuno and is restarted on the node. Detail in that submodule’s ownCHANGELOG.md.yunos/js: a gui_agent tab remembers where you were inside it. Open a topic in a Schemas tab, look at another tab, come back — and you landed on the treedb cards with the topic gone: a tab’s nav item carries a FIXED route (its base), so clicking it always navigated to the root.The route cannot simply be made deeper —
yui_shell_set_submenu()registersitem.routein the shell’s item index, so it is where the tab’s view is MOUNTED and what a deep link resolves to.C_APPremembers the last position per tab and replays it when the tab is entered again, which it tells apart from walking UP inside the tab (the view’s own Topics button) by looking at whether the previous route belonged to another tab. Detail in that submodule’s ownCHANGELOG.md.yunos/js: gui_agent’s Schemas workspace draws the treedb as a graph. Under a treedb the subpath was<topic>[/info]orschema; it is now alsograph[/<topic>], and the third icon of every topic card goes there — the treedb’s viewer hosts the graph beside the topic editor, on the SAME routing adapter, and swaps the two bodies by url. Mounted lazily (G6 is the heaviest thing the workspace draws, and it measures its canvas when first SHOWN), and opened with thedagrelayout rather than the library’smanual, which places nodes where the records say they are — a schema treedb carries no geometry, so it opened as one diagonal pile of 265cols.The adapter had to learn the treedb LINK events: the graph subscribes to
EV_TREEDB_NODE_LINKED/UNLINKEDon its transport as soon as it loads a topic, andgobj_subscribe_eventrefuses an event that is not in the publisher’s output list — so leaving them undeclared did not cost an edge that fails to redraw, it cost the SUBSCRIPTION.gui_treedbdeclaresEV_RECORD_WRITTEN, which had been an FSM error on every write since gobj-ui 5.14.0. Detail in that submodule’s ownCHANGELOG.md.gobj-ui submodule 5.13.0 → 5.14.1, and the JS yunos with it.
5.14.0publishesEV_RECORD_WRITTENfromC_YUI_TREEDB_TOPICS: the view refreshes itself from the treedb’s own node events, which arrive for every writer and therefore say nothing a host can act on — a schema editor, where changing a column also has to raise the versions that publish it, cannot use them without answering its own writes in a loop.5.14.1teaches the treedb table therowidtype: a topic keyed by rowid — which is how the__system__treedb has stored a schema since this release — logged “unhandled type ‘rowid’” once per cell over every render oftopicsandcols, the two topics a schema editor exists to edit.yunos/js: gui_agent grows a Schemas workspace — gobj-ui’s treedb editor, unchanged, over a routing adapter that re-wraps each command ascommand-agent+cmd2agent="command-yuno …", so one control-center session reaches thetreedb_system_schemaof any yuno of any node. It discovers which treedbs a yuno exposes, marks in the picker the yunos with none, carries<treedb>/<topic>in the URL, and applies a schema by restarting the owning yuno. Detail in that submodule’s ownCHANGELOG.md.yunos/js: the two SPAs give their primary rail back to the work. Ingui_agentthe rail is the four workspaces (Commands, Statistics, Terminal, Schemas) and the settings page moved to/preferencesunder the toolbar avatar; ingui_treedbthe rail is Topics and Graphs, and what/settingsheld became two pages with their own names —/connections(the backends, where each workspace picker sends you) and/preferences(the live buffer). All of them stay ROUTES, so they are linkable, survive an F5 and appear in the site map.On the way, gui_agent’s Schemas tab stopped routing its own url by hand: the treedbs of a yuno are now a tree of nodes (
C_YUI_NODEat the tab’s route, onelinkchild per treedb), which is what makes the shape of the navigation a user choice — Preferences → Navigation: stacked strips, back to parent, or breadcrumb, applied live to the open tabs. Detail in that submodule’s ownCHANGELOG.md.
7.12.0-2¶
A packaging revision, not a new version: the tree under kernel/, modules/,
utils/ and yunos/ is the same one 7.12.0 was cut from. The packages are
rebuilt as yuneta-agent-7.12.0-2 and attached to the existing 7.12.0 tag.
Fixed¶
Nothing ever restarted the SECOND agent, on either distro. An upgrade replaces both agent binaries, and only the main one is ever bounced: the init script’s
stop_yunos()stops just that one (start brings up both, stop takes down one), and no package scriptlet touchesyuneta_agent22at all. “Do not bounce both at once” had quietly become “never bounce the second one” — it kept running its old inode for as long as the node stayed up. Found on all five nodes at once, four of them five days deep, and the code the spare was running was the version-comparison bug this release exists to fix. The escape hatch was the oldest thing on the node, which is the exact opposite of what a second agent is for.install.shnow refreshes it — but only after confirming the main agent is up and running the binary that was just installed. (That script is fetched frommain, not shipped in the package, so it takes effect without a package revision; the%prechange below is what the-2packages carry.) If the main agent is unhealthy, the spare is the only way into the node and is left strictly alone, with the manual command printed. A spare that does not come back is reported without failing the install, because the node still has its main agent.The staleness test is the process, not the file:
readlink /proc/<pid>/exeending in" (deleted)".rpm -q,--versionand the installer’s own sign-off all read the file, and all three said everything was fine.An
.rpmupgrade left the OLD agent running, and every check said it had worked. dpkg stops the agent inprermand starts it again inpostinst, so a.debupgrade lands on the new binary. rpm has no equivalent hook:%preunstops only on uninstall ($1 == 0), and%post’sservice yuneta_agent startis a no-op against a process that is already running. So the new binary went to disk and the old process kept running on the unlinked inode.Nothing reported it.
rpm -qsaid7.12.0-1,yuneta_agent --versionsaid7.12.0, and the installer signed off with “yuneta_agent and yuneta_agent22 are running” — all three read the file. Onlyreadlink /proc/<pid>/exe, ending in" (deleted)", read the process. This is how 7.12.0 reached a node installed and not running: the release whose whole point was a fix in the agent.%prenow stops the main agent on an upgrade ($1 == 2), mirroringprerm. Only the main one —yuneta_agent22stays up, the same asymmetry the init script’sstop_yunos()already has (start brings up both, stop takes down one): the second agent exists so each can recover the other, and an upgrade that bounced both at once would give that up for the seconds it matters most. The yunos are untouched — they outlive their agent and the new one adopts them back over the control channel, measured across five nodes as the same yunos with the same PIDs.
7.12.0¶
A minor for what was found, not for what was added.
The agent had been demoting yunos: a version comparison packed its segments
into an int, 1.9.0.0-2 overflowed into a negative number, and every
deactivate-snap re-appended the OLDER release as the primary — for eleven
days on a client node, with nothing in any log to say so. The comparison was
both the decision and the only guard, so when it lied there was nothing left to
notice. It is version_cmp() in the SDK now, comparing segment by segment,
tested against that node’s real version chain; and the direction is checked on
its own, logged either way, and a release that does not move forward needs
force=1.
Two more of the same shape: a msg2db accepted records it could never load back
(4447 of them on that node, and 4447 log lines at every start), and the package
would unpack another machine’s build over a node that compiles its own —
foreign glibc archives into outputs/, which a static link takes in silence and
the heap pays for at run time. Both refuse now, at the point where the mistake
is made.
Fixed¶
The agent promoted the OLD binary, and kept doing it.
get_n_v()weighed each segment of a version by 1000 and accumulated into anint. A four-segment release with its revision —1.9.0.0-2— needs 10¹², so it overflowed and came out negative:1.7.1.0-2 -> 1,978,652,738 1.9.0.0-2 -> -317,314,558promote_highest_release_yunos()compares with that, so it read the older release as the newer one and re-appended it as the primary. Not a failure to promote: an active demotion, repeated at every restart. The store on a client node showed it plainly —1.9.0.0written at 13:46:25,1.7.1.0written back two seconds later, and again on 01/08, 04/08, 05/08, 09/08 and 10/08. Eleven days ofdeactivate-snapbringing back the binary it was supposed to replace, with nothing in any log to say so, because a comparison that comes out negative says nothing.The function moved to the SDK as
version_cmp()— it is a string utility, not agent business, and aPRIVATEin a gclass cannot be tested. It compares segment by segment and packs nothing into a number, which is what the JS (version_tuple) and Python sides of this project already did; the C side was the only one that accumulated, and the only one that broke. A wider accumulator would have moved the ceiling without removing it: weighing segments by 1000 also assumes every segment stays under 1000, so1.2000.0outranked2.0.0, and it right-aligned the segments, so7.11and7.11.0were different versions. Neither survives the rewrite. All 8 comparisons inc_agent.cgo through it.Covered by
tests/c/helpers, with that node’s real version chain. The test was checked against the old implementation first — it fails six times there: the exact pair that cost the eleven days, the1.2000.0/2.0.0inversion, and the7.11=7.11.0equality.A release that does not move forward is refused, and says so. The comparison that picked a candidate was the only thing standing between a deploy and a downgrade, so when it was wrong there was nothing left to notice — eleven days of demotions and not one line about it.
cmd_find_new_yunos()now checks the direction on its own and logs it either way: an info line namingfrom→towhen a release moves forward, a warning when it does not, and it refuses to take it unless the caller passesforce=1.promote_highest_release_yunos()gained the matching guard: a promotion that would go backwards cannot happen by construction, so if it ever does it is an error, not a silentgobj_update_node().
Changed¶
The package refuses a node that builds from source. It carries libraries, headers and binaries built on another machine, and it lays them into the very tree
yunetas buildowns —outputs/andoutputs_ext/. From then on fresh objects link against archives built for a different glibc: a dynamic link fails loudly, a static one resolves them silently and the binary corrupts its heap at run time. This is not hypothetical — it cost a client node a week, and it is whylibc_guard.cmakeexists.preinst/%prenow abort when/yuneta/development/yunetasholdskernel/and.git, andinstall.shchecks first so the refusal is not buried under apt’s own error plus two failed fallbacks. Installing is still supported — it just has to be a decision:sudo YUNETAS_FORCE_OVER_SOURCE=1 apt install ./<package>.deb sudo touch /etc/yuneta/allow-package-over-source # or make it stickThe variable goes after
sudo, which resets the environment; the marker file is there because that footgun is not obvious. Forcing prints what has to happen next: rebuild everything, external libraries FIRST. Rebuilding only the SDK leaves the packaged externals in place, which is the same mismatch by a shorter road.A msg2db says once what it used to say thousands of times. Every record whose
pkey2value is empty is dropped at load, and each one wrote its own error line: a client node was printing 4447 identical lines at every start, which buries everything else that log exists for.Now the first dropped record is logged whole, so it can still be diagnosed, and
msg2db_open_db()reports the count when the topic finishes loading. Two lines instead of thousands, and neither of them silent — the number of records that did not load is stated, which is the part that matters.
Fixed¶
BREAKING: a msg2db refuses a message whose
pkey2value is empty. The write side only asked whether the key was present, while the load side asks for a non-empty value — so a record carrying"alarm": ""was accepted, written to disk, and then dropped at every load for the life of the store. A client node was carrying 4447 of them, written across a fortnight and unreadable ever since.A store must not take what it cannot give back. The refusal belongs where the caller still holds the record and can be told, not in somebody else’s log a restart later.
⚠️ An app that writes those records will now see them refused, with an error naming the record. That is the intended consequence: the mistake surfaces where it is made. Records already on disk are untouched, and the load side keeps its guard and its reporting for them.
tr_msg2db.cno longer keeps a file-staticjson_t *cache.topic_cols_descwas aPRIVATE json_t *kept alive across open/close with an incref dance — the pattern CLAUDE.md forbids for library helpers, and the same variable thattr_treedb.cis cited as the fixed example of. Nothing outsidemsg2db_open_db()ever read it, so it is a local now, built per call and released before the function returns.⚠️ This was not the leak it looked like: msg2db still leaks 8 tracked blocks per open/close cycle, measured and written down in
TODO.md. The static was removed because it is forbidden, not because it was guilty.
7.11.0-3¶
A packaging revision, not a new version: the tree under kernel/, modules/,
utils/ and yunos/ is the same one 7.11.0 was cut from. The packages are
rebuilt as yuneta-agent-7.11.0-3 and attached to the existing 7.11.0 tag.
One change, and it closes the difference the previous revision opened.
Changed¶
/etc/yuneta/webserveris handled the same way in the.deband the.rpm, and neither ships it. 7.11.0-2 fixed the.debby dropping the file from the package; the.rpmstill shipped it as%config(noreplace). Both worked, differently, which is the kind of asymmetry that is fine until the day it is not.They now follow the pattern this packaging already proved for
nginx.conf:preinst/%presaves the node’s value to/etc/yuneta/webserver.pkgsave, andpostinst/%posttransputs it back if the transition removed it. A node that already chose keeps its choice; only a node that never had one gets the build default seeded.The save matters more than it looks: a file an upgrade no longer provides is a file the package manager removes, which is exactly how 7.9.1 deleted two nodes’
nginx.confwhile fixing the overwrite that preceded it.
7.11.0-2¶
A packaging revision, not a new version: the tree under kernel/, modules/,
utils/ and yunos/ is the same one 7.11.0 was cut from. The packages are
rebuilt as yuneta-agent-7.11.0-2 and attached to the existing 7.11.0 tag.
Everything here comes from one install that failed on three nodes at once, and from what that failure left behind.
Fixed¶
A conffile question killed the install, and the install path cannot answer one.
install.shis meant to be run ascurl ... | sudo sh, which leaves dpkg with no stdin. It called apt with noDEBIAN_FRONTENDand no conffile policy, so the first question dpkg asked read EOF and took the whole install down — “end of file on stdin at conffile prompt” — with the package left half configured and re-running it hitting the same wall.It fired on three nodes because
/etc/logrotate.d/yunetahad reached them by hand before it shipped in the package, so dpkg does not own it and asks what to do with it.install.shnow runs noninteractive withconfdef+confold, and says afterwards when a file was kept. Thepreinstalso adopts an unowned copy, so the question cannot come back.⚠️ The failure is worse than a failed install, and that is the part to remember: dpkg had already unpacked and
prermhad already stopped the agent and the web server. Three nodes were left with their web server down, theirnginx.confdeleted — dpkg removes what the package no longer ships — and nothing to restart them. It was only recoverable becausepreinsthad saved the configuration tonginx.conf.pkgsave.The package no longer decides which web server a node runs. It shipped
/etc/yuneta/webserverwith whatever the build machine chose, and the build machine has no opinion, so it shippednginx. On 7.11.0 both openresty nodes came back serving from the wrong tree with a default configuration — one of them the company server, the other a client’s.The file is node state, exactly like
nginx.conf, which stopped being shipped in 7.9.1 for this same reason. It is not in the package any more; the scriptlet creates it only when it is absent.The handover to the web server unit stops the old server for real. It asked with
-s quit, which is the graceful shutdown: it waits for the requests in flight, and on a node with websockets that wait does not end. Twenty seconds were not enough, the old master kept:80and:443, and the unit went into a restart loop that failed to bind 17 times. It now escalates to-s stop, andyuneta-webserverlearned that verb.packages/deb/README.mdsaid a non-interactive install reboots the node. It has not been true since the auto-reboot was removed. Reading that paragraph before installing on a client’s node is a bad minute to have.
7.11.0¶
A minor for one reason: a public call changed its name and its signature, one day after it got one.
7.10.0 added gobj_post_message() to C. The name was wrong and the shape was
wrong, and it took the compiler to say so — gobj_post_event(dst, event, kw, src) already existed in gobj-js and in the ESP32 port, and the ESP32
component includes the Linux gobj.h, so the three-argument version did not
even build. Two implementations already agreed; C is the one that moved.
gobj-js then moved too, off a setTimeout(…, 10) and onto the same
contract.
The rest is the node’s web server: it has its own systemd unit now instead of being a side job of the init script, and on Rocky it needs a label to be allowed to start at all.
Fixed¶
The web server unit did not start on Rocky, and the node served nothing.
/yunetais outside the SELinux policy, so everything under it is labelleddefault_t. The SysV script could exec that, becauseinitrc_tmay; systemd cannot. The unit shipped in 7.10.0-2 died with203/EXEC— “Failed to locate executable /yuneta/bin/yuneta-webserver: Permission denied” — retried four times, gave up, andyunovatios-centralstayed two and a half hours with no web server at all.The
.rpmnow labels the wrapperbin_t(withsemanage fcontext+restorecon, falling back tochconwhere the policy tools are missing), and the scriptlet that already warned when the unit does not start now says where to look. Only the wrapper needs the label: it is the one file systemd execs, and what it execs in turn is reached from its own domain.Debian is untouched by this, which is exactly why it was not seen before shipping: the same revision on
yunovatios-controladorcame up fine.
Changed¶
BREAKING:
gobj_post_message()is nowgobj_post_event(), withgobj_posted_events_size()andgobj_deliver_posted_events()renamed to match.The call it adds to C already existed in JavaScript, and had for years:
gobj_post_event()ingobj-js, under a comment that reads “post_event, by now only in js”. 7.10.0 shipped it into C under a third name for the same idea, which is the one thing a framework with two implementations cannot afford.The signature changed with the name, and the compiler is what said so: the ESP32 port declares the same call too,
gobj_post_event(dst, event, kw, src)over anesp_eventloop, and its component includes the Linuxgobj.h, so the three-argument version did not build. Two implementations already agreed on four arguments, the same four asgobj_send_event()— because posting is sending, later. C is the one that had to move.So the destination is no longer always the caller. Lifetime is handled both ways instead: destroying the DESTINATION drops what it had pending, and destroying the SOURCE clears
srcand keeps the entry, because the destination still wants its event and gets it withsrc == NULL.The only caller is
webstats, so nothing outside this tree breaks.gobj-jsaligned to the same contract (package7.10.0). It deferred withsetTimeout(…, 10)— one timer per event, and ten milliseconds standing in for “later”, which is the very thing this call exists to stop writing. It now keeps a queue drained once per turn, a snapshot at a time, with the same lifetime rules, the same check at post time, the same ceiling and a trace line. Nothing called it, so the change breaks nobody. ESP32 still has its own contract, and what differs is inTODO.md.
7.10.0-2¶
A packaging revision, not a new version: the tree under kernel/, modules/,
utils/ and yunos/ is the same one 7.10.0 was cut from. The packages are
rebuilt as yuneta-agent-7.10.0-2 and attached to the existing 7.10.0 tag.
Both changes are about the node’s web server, and both came out of one daily report that arrived empty.
Fixed¶
The web server logs rotate to
.1on every node, never to a date./etc/logrotate.confon RHEL and Rocky carries a globaldateext, so the same drop-in was givingaccess.log.1on Debian andaccess.log-20260808on Rocky.That is not cosmetic. The
webstatsyuno reads<path>and<path>.1: the day it reports is the day each LINE carries, and the previous day lives in the.1thatdelaycompressleaves uncompressed. Withdateextthere is no.1, so the yuno reads today’s file alone and reports the previous day as empty, with nothing in any log to say why.Found by accident, which is the only reason it was found at all:
logrotate -f /etc/logrotate.d/yunetadoes not readlogrotate.conf, so forcing a rotation by hand produced a.1and the difference showed itself.
Changed¶
The node’s web server is a systemd unit now, not a side job of the init script.
/etc/init.d/yuneta_agentran nginx and let it daemonize, so nothing owned the process afterwards: a laterstartfound nothing to look at and tried again, and the second master died with “Address already in use” while the first one kept serving.yuneta-webserver.serviceruns it withdaemon off, so$MAINPIDis the real master and stop and reload reach it. The unit calls/yuneta/bin/yuneta-webserver, which is the one place that reads/etc/yuneta/webserverto choose between nginx and openresty — that choice used to be repeated in three functions of the init script, and a unit file cannot branch on the content of a file.systemctl reloadsends HUP, which reloads the configuration; reopening the log files is USR1 and stays in thepostrotateof/etc/logrotate.d/yuneta, where logrotate does it itself.⚠️ The upgrade interrupts the node’s web server briefly. The unit cannot take
:80and:443while the old daemonized master holds them, so the scriptlet asks that one to finish first and waits for it, then starts the unit and says so in the log if it did not come up.
7.10.0¶
A minor, not a patch: the framework gets a call it did not have.
gobj_post_message() names something every gclass already did and had no way
to say — do this, but not on this stack. Until now that was written as a
C_TIMER0 of one millisecond, which is a time for something that is not a
time, and which cost the name of the event: every deferred continuation
arrived as EV_TIMEOUT, so the machine trace, which is the execution log of
a yuno, said “timeout” instead of what happened.
The call is not new to Yuneta. It existed in the first versions and was cut when io_uring came in; what was missing was the wiring, not the design.
Nothing is removed and nothing changes shape: a gclass that does not call it
behaves exactly as before. webstats is the first consumer and the only yuno
that changed.
Added¶
gobj_post_message(): an event a gobj sends to itself, delivered on the next cycle of the event loop. It is the way for an action to leave the stack it is standing on — the usual case being a subscriber that must destroy or stop the publisher whose synchronousgobj_publish_event()is still on the stack below it.The call replaces the idiom of a
C_TIMER0child armed with 1 millisecond, which was never a time: it was “later”, written as a duration. Saying it that way costs an io_uring timeout for something with nothing to wait for, needs a child gobj with its own start/stop, and throws away the name of the event — every deferred continuation arrived asEV_TIMEOUT, so themachinetrace, which is the execution log of a yuno, said “timeout” instead of what happened, and a gclass with two deferrals had to tell them apart with a flag.The contract, in full in
gobj.h: self-send only, posted from the thread of the event loop,kwowned, the event checked against the gclass at post time so the error names the caller, a ceiling of 10000 pending (it is not a work queue), and delivery as a snapshot per cycle — the messages queued when a cycle begins are the ones delivered in it, so a chain of posted events advances one step per turn of the loop and never starves the io_uring completions.gobj_destroy()drops what its gobj left posted, andgobj_end()says how many were never delivered.yev_loop_run()delivers them at the top of each cycle — before the completions, so an event posted inmt_play(), with the loop not yet running, does not wait for a completion that may never come — and does not block on the ring while any are pending. A gclass never calls the delivery itself. The ESP32 port carries its own copy of gobj and does not have this yet.Covered by
tests/c/gobj_post_message, which checks each clause and, for the snapshot, that a chain posting itself for 200 ms still hears a 1 ms periodic timer: draining the queue until empty instead gives 2.2 million links and not one completion seen.
Fixed¶
webstats: a file whose reader could not be built left the report and skipped the next one. The path recorded nothing insources, so the file vanished from a report that then read as a report of everything — the one thingsourcesexists to prevent. It also dropped the file from the pending list itself, andac_next_file()drops the head too, so the next file went with it, unread and unmentioned.The second half only became reachable with the change below: before it, the correlator guard abandoned the run before either removal. Found by forcing the path, which had never executed — the fix is that the file is recorded as unread and the list is left to its single owner.
Changed¶
webstatsuses posted events for its two continuations. The continuation between files isEV_NEXT_FILEand the one between chunks of a file isEV_READ_CHUNK, each named for what it does instead of arriving asEV_TIMEOUT. BothC_TIMER0children are gone;C_WEBSTATSkeeps itsC_TIMER, which measures a real time — the daily schedule.This also removes a way for a run to stall for ever. The reader-creation failure path armed the deferred continuation while the
reader_doneflag that told the twoEV_TIMEOUTs apart was still false, and the guard that read that flag abandoned the run:ST_READINGwith the schedule already disarmed, so no further report until the yuno was restarted, and everyreport-dayanswering “A run is already going”. With the event named, the flag and its guard have nothing left to do.
7.9.13¶
A release to carry one agent fix to the nodes. yuno_agent changed, so this is
a version bump and not a packaging revision.
It is a version bump for a reason worth stating: the agent is not a managed
yuno, so no install-binary reaches it. On a node with no SDK sources the
package is the only road to its binary — the same argument that cut 7.9.12.
yuneta_agent22 is untouched and needs no update: the escape hatch exposes tty
and consoles only, no config commands. That is what lets a node take the new
agent with its second agent still running.
Fixed¶
update-confignever refreshed thedescriptioncolumn, so a config row kept the description of its first version for the rest of its life while its content changed underneath.cmd_create_configwrote the column from the content’s__description__;cmd_update_configwrote onlyzcontentanddateand left the column alone.That column is not decoration.
list-configsis where an operator reads what a config row is for, and the batch convention exists to feed it: every edit rewrites__description__as that version’s changelog entry. The convention was resting on a field the update command did not maintain, andsync-configsdrives its wholeUPDATEpath through that command.Found on
e.com: thewebstatsrow still announced “1: initial load” long after its content had stopped reading the dead openresty tree. The content was right, the label was two edits stale, and nothing anywhere said so.How far it had spread: one row. The five nodes were swept afterwards, comparing every config row’s column against the
__description__its own content carries — 51 rows over 546 stored records — and thee.comone was the only mismatch. The reason is the convention itself: an edit normally bumps__version__, which takes thecreate-configpath and writes the column correctly.update-configoverwrites the row in use, which is the deliberate exception for changing one yuno without bouncing a node, and it is rare. The defect was real and silent; its blast radius was not wide.An update now writes the description with the content, unconditionally, as
create-configdoes — a content with no__description__leaves the column empty rather than keeping a text that describes content that no longer exists. Exercised on a live agent: changed, restored, and emptied.
7.9.12¶
A release with one reason to exist: to put the webstats binary in the
packages, so the other four nodes can run it. The yuno was written and proved
on one node, and a yuno that lives only in a build tree reaches nobody — the
.deb and the .rpm ship outputs/yunos/ whole, and that is the road to a
node that carries no SDK sources.
Nothing under kernel/, modules/ or utils/ moved, so no running yuno needs
a rebuild for this. @yuneta/gobj-js stays at 7.9.11 on npm: this release
carries no JavaScript.
Added¶
webstats, a new yuno: the daily report of the node’s web server logs. It reads the nginx (or openresty) access and error logs of a day, counts what they hold, keeps the numbers in TimeRanger2 and mails the result throughemailsender. One yuno per node.It exists because a sealed node has no SSH, and reading a log by logging in is a task that stops working the day the node is sealed. The five nodes had never had their logs read until somebody went looking.
The decision that keeps it small: the day of a line comes from the timestamp the line carries, never from the file it sits in. No read offset, no hook into
logrotate, so the report of any day still on disk can be rebuilt at will and a yuno that was down loses nothing.What the report leads with is visitors, and how many are new — an address that asked for a piece of the page and got it, whose user agent carries no crawler mark. Counting requests answers a different question: on the first real day the top user agent was one
curlmaking 3015 requests from a single address, and 1346 addresses claimed to be a browser while 71 ever fetched a script. Fingerprints are stored, never addresses.Then: totals and status classes, per hour, per vhost, the top paths, 404s, clients, agents and referrers, every 5xx whole, probes counted (banning stays
fail2ban’s job), a latency histogram with percentiles, and the error log grouped by signature — the half where the real findings were.Documented at
/webstatsand inyunos/c/webstats/README.md, which carries the design and the traps.
Changed¶
log_format vhostgained$request_timeand"$upstream_response_time", appended last, under the same rule as$host: everything before them stays byte for bytecombined, soawkcolumn positions, goaccess and the stock fail2ban filters keep working. The count of quotes now says which generation a line belongs to — 6, 8 or 10 — and the rotated files hold all three for 30 days, so any reader of these logs has to accept all three. Deployed on the five nodes; the configs live in their own operations repos.CLAUDE.md: thecommand-yunoname collision is not onlyid.cmd_command_yunopasses its whole kw as the filter that selects the yuno, so every field of the yuno record is a reserved parameter name. A command with adateparameter answers “Yuno not found” and names the yuno, never the parameter.CLAUDE.md: a bumped config version does not reach the yuno on its own.create-configappends a row and the primary does not move — not even a restart of the yuno picks it up. Only a node-widedeactivate-snappromotes it. The note now says when to bump and when to overwrite the row in use withupdate-configinstead of bouncing a node to change one yuno’s config.
7.9.11-3¶
A packaging revision, like the one before it: nothing under kernel/,
modules/, utils/, yunos/ or tests/ changed, so YUNETA_VERSION stays at
7.9.11 and only the RELEASE counter moves. The packages are rebuilt as
yuneta-agent-7.9.11-3 and attached to the existing 7.9.11 tag.
It exists for one reason: the fail2ban filter shipped in 7.9.11-2 could not ban anybody on a single-page app, and a package is the only way that fix reaches a node that installs from scratch. It was found by doing exactly that — imaging two nodes, installing 7.9.11-2 on them and probing them end to end.
Fixed¶
The fail2ban probe filter never fired on a SPA vhost. It matched only responses 404, 403 and 444, on the reasoning that a path some app really serves would stop matching once it answered 200. That holds for a static site and collapses on a single-page app: with
try_files $uri $uri/ /index.htmlevery unknown path answers 200 withindex.html, so/wp-login.phpcame back 200 and the filter matched nothing. Both yunovatios consoles were running the jail blind — enabled, healthy, watching the right file, and incapable of ever banning anybody.Found by installing 7.9.11-2 on a freshly imaged node and probing it end to end, which is the only way it could have been found: on the nodes it was developed against, every unknown path 404s.
The status is gone from the three patterns. It costs nothing, because the paths cannot be legitimate on a Yuneta node — no PHP, no WordPress, nothing serving
.envor.git. On one day of real traffic the restriction was also hiding 606 further lines on a mixed node: probes answered 301 by the http→https redirect and probes answered 200 by a SPA. Of all of them, none came from another node of the fleet and none from a real crawler.The path is now matched up to the query string, so a
.phpappearing only in a parameter is not a match.
7.9.11-2¶
A packaging revision, not a new version of Yuneta: nothing under kernel/,
modules/, utils/, yunos/ or tests/ changed, so YUNETA_VERSION stays
at 7.9.11 and only the RELEASE counter moves. The packages are rebuilt as
yuneta-agent-7.9.11-2 and attached to the existing 7.9.11 tag.
What it ships is the answer to a question nobody had asked of these nodes:
what is in their logs. nginx has no rotation of its own, so access.log and
error.log had been growing since the day each node was installed — on all
five, none had ever been rotated. Reading them for the first time turned up
that 99.9% of one node’s error.log was scanner noise, that 43% of all
requests were probes for /.env and /wp-login.php, and that nothing was
watching any of it.
Added¶
The packages rotate the web server logs —
/etc/logrotate.d/yuneta, a conffile in both the.deband the.rpm, withlogrotateadded toDepends/Requires.nginx has no rotation of its own: it only knows how to reopen its files when it gets
USR1. Nothing was sending that signal, soaccess.loganderror.loghad been growing since the day each node was installed — on all five, none had ever been rotated, and the busiest was writing 3.5 MB a day. Disk was never the problem. A log nobody can open is.Both trees are listed (
/yuneta/bin/nginx/logs/,/yuneta/bin/openresty/nginx/logs/): a node runs one or the other, andmissingokcovers the absent one. Daily, 30 kept, compressed withdelaycompressso the file the master still writes to is not cut. ThepostrotatesendsUSR1only to a pid that is alive, so a stale pid file cannot signal a process the kernel handed to somebody else. The logs of the certbot deploy hook (/var/log/yuneta/*.log) rotate monthly in the same file.The yunos are not touched: each one writes numbered files under
/yuneta/realms/<realm>/<yuno>/logs/and rotates them itself.The packages ship a fail2ban filter and jails for the web server —
/etc/fail2ban/filter.d/yuneta-nginx-probe.confand/etc/fail2ban/jail.d/yuneta-nginx.conf, both conffiles.The stock
nginx-botsearchfilter looks for webmail, phpMyAdmin and WordPress. Measured against 15 days of a real node’saccess.logit matched 1089 lines out of 224 645, while that log held 95 829 refused requests: the node was scanned all day with nothing watching. The new filter matches 52 383 of those lines, from 972 addresses.It does not ban on 404s — search engines collect those honestly, and a rate rule bans Googlebot first. It bans on what was asked for: any
.phppath (no node runs PHP), the dot-directories that hold source control or credentials, and/wp-*. Several hundred matching lines carry the user agent of Googlebot, GPTBot or ClaudeBot, and every one is an impostor: the addresses reverse togoogleusercontent.comand to Cloudflare, and the real Googlebot does not ask for/.env.backup.Both jails ship disabled, which is not timidity. If none of a jail’s
logpathglobs resolves to a file, fail2ban does not skip the jail — it refuses to configure and the whole server exits 255, taking every other jail down with it,sshdincluded. A node carrying this package whose web server has not run yet is exactly that case.packages/deb/README.mdcarries the two-line command to enable them, and the two ways a jail lies about its own health: watching nothing while reporting healthy (a node whose[DEFAULT]setsbackend = systemd, so the jails now pinbackend = auto), and recording bans that never reach the firewall (abanactionnaming a command that is not installed).The documentation site owns its
robots.txt—docs/doc.yuneta.io/robots.txt, installed bydeploy.shover the one mystmd builds.myst has no absolute address for the site (
site.options.base_urlis/), so theSitemap:line it wrote named the address of its own development server,http://localhost:3000/sitemap.xml. Every crawler had to discard that line, which left the sitemap myst does build — 12 KB of it — reachable by nobody.It also refuses six backlink and rank crawlers (Semrush, Ahrefs, MJ12, dotbot, DataForSeo, SERanking), which read the whole site to sell the numbers back and brought about 14 000 requests in 15 days. Search engines and the assistants are deliberately not on that list: those are how somebody finds Yuneta.
A 404 page for the documentation site —
docs/doc.yuneta.io/errors/404.html, installed at the root of the build bydeploy.shthe same way as the landing page, because the--deletersync erases anything dropped straight into the docroot. Five hostnames share that docroot and every one of their server blocks declarederror_page 404 /404.htmlfor a file that did not exist, so nginx logged a failed open on every 404 and the reader got the built-in page of nginx.
Fixed¶
install.shfetched the OLDEST package of a release, not the newest. It picked the asset withhead -n1, and the GitHub API lists assets in upload order — so on a release carrying more than one packaging revision it chose the first one uploaded. Caught the morning after this revision shipped: the installer downloadedyuneta-agent-7.9.11-1while-2sat next to it, and reported a clean install of a package with none of the logrotate or fail2ban configuration the revision exists to deliver.Now
sort -V | tail -n1, which compares the numbers as numbers: revision 10 sorts after 2, where a plainsortwould not.install.shis served frommain, not from any package, so this needs no rebuild and no new revision — it takes effect on the nextcurl | sh.
7.9.11¶
The version skips 7.9.10 on purpose: @yuneta/gobj-js shipped 7.9.10 and
7.9.11 on its own line while the SDK sat at 7.9.9, and the SDK catches up to
the package rather than the other way round.
Added¶
shutdowncommand onC_YUNO— the control plane can now ask a yuno to stop itself, orderly. It answers first and dies after: the response travels over the very event loop the shutdown stops, so a handler that calledset_yuno_must_die()on the spot would take the socket down with the answer still in it. The handler arms a timer andac_timeout_periodic()does the dying — the same shapetimeout_restarthas always used, so the wait is at most onetimeout_periodic(1 s by default). Measured end to end: asked att, dead att+999 ms.It exits with code 0, so the ydaemon watcher does not relaunch the yuno. It does not replace
kill-yuno, which stays the deploy path: that one is asked of the agent, which knows the yuno and deregisters it; this one is asked of the yuno, which is what you have when the yuno answers and the agent does not, or when the yuno runs under no agent.Pinned by
tests/c/command_shutdown/, which asserts the order that matters: the command answersresult=0, the yuno is still running when the answer comes back, and it dies on its own afterwards (a watchdog turns “it never died” into a failed check instead of a ctest timeout).@yuneta/gobj-js7.9.11 —C_TIMER: the two calls are the whole contract (submodule bump).set_timeout()arms andclear_timeout()disarms, as in C (c_timer.h); whether the gobj is running stops being the caller’s problem. The JS port had neither half: every view had to pair itsset_timeout()with agobj_start()and remember agobj_stop()on the way out, or get “Destroying a RUNNING gobj” when its route was left. The running state now follows the timeout inmt_writingon themsecattribute — not in the helpers, because those three PUBLIC functions are an escape from the gclass interface and must be sugar and nothing else: writingmsecby hand leaves the timer exactly asset_timeout()would.A second fix from the same reading: a periodic cleared from inside its own action now really stops. The re-arm ran after the action, undoing the clear and re-arming with the
msecthe clear had just written — a negative delay, whichsetTimeout()serves immediately, so the timer became a busy loop.BREAKING for callers that stop the timer themselves (
clear_timeout()followed bygobj_stop()now logs “GObj NOT RUNNING”). Every in-tree consumer was migrated in the same release; theyunos/jssubmodule carries its half. 7.9.11 finishes the job inside the runtime itself, wherec_ievent_cli.mt_stop()still stopped its timer by hand and logged that very complaint on every disconnect.@yuneta/gobj-js7.9.9 —gobj_set_gclass_no_trace()(submodule bump). The C kernel has it; the JS port did not, even though the field was there and already consulted. Without it the idiom every Cmain()uses to keep timers out of amachinetrace could not be written in JS at all, so the SPAs’ machine trace drowned in the yuno’s one-second periodic tick. Both JS yunos now carry that block. Also realigns the package withYUNETA_VERSION, which had drifted at 7.9.6.@yuneta/gobj-ui5.9.0 — shared clipboard helpers (submodule bump).yui_clipboard.jsgives every table one line to hand its rows over:yui_copy_table_json()copies what the user is LOOKING AT — the selected rows when there is a selection, otherwise every row the current filters leave on screen. Four views had each grown their own copy code and it had drifted; two wrote unindented JSON and two said nothing when the write failed.gui_agent’s tables copy as JSON (
yunos/jssubmodule bump). Until now the only way to pass one of those lists to anyone was a screenshot. Nodes and Statistics get a Copy JSON button; the console’s copy button — created disabled and only re-enabled for text answers, so it was dead fortop,list-yunosand most commands — now copies table answers too. Each history row also runs its command in one gesture instead of two.@yuneta/gobj-ui5.10.0 —yui_button_mark_done()/yui_button_unmark(), so a copy button can say “Copied” for a moment. The timing stays in the consumer’s FSM (aC_TIMER+EV_TIMEOUT), not in asetTimeoutinside the library.@yuneta/gobj-ui5.11.0 — the PWA install offer (submodule bump). Chrome advertises its install banner on a heuristic nobody can read, and goes quiet on an origin for months after a dismissal or an uninstall — the app then looks uninstallable when it is only unadvertised.yui_install.jsrefuses the banner, keeps the event and asks once per browser with the family dialog. The event arrives before the bundle is parsed, so each SPA catches it inpublic/install-prompt.jsloaded by<script src>— never inline, which theirscript-src 'self'drops in silence. Ported from yunomúsica.The JS yunos install as PWAs and their tables copy as JSON (
yunos/jssubmodule bump), and every action in them crosses the FSM: the Refresh buttons and the console’s copy flash were a direct call and a baresetTimeout, so a click and its consequences never reached themachinetrace.doc.yuneta.io reads offline.
pwa/sw.jskeeps the theme bundles (cache-first — their names are content-hashed) and every page the reader opened (network-first, so a redeployed page never reads stale), and falls back to aoffline.htmlfor a page that was never read. It does not precache: the build is 28 MB, which is not a cost to put on a mobile connection for pages nobody asked for.deploy.shstamps the deploy version intosw.js— the caches are named after it andactivatedrops every other, so a deploy cannot leave a mixture of an old page and new bundles. A page arrives under two urls (/<slug>, and/<slug>?_data=<route>when the theme routes on the client); this host answers the same bytes to both, so the query is dropped from the cache key and one entry serves both.pwa/offline-test.mjsasserts the lot against the deployed site — and does it by relaunching a profile behind a dead proxy, becausecontext.setOffline()is a no-op in Playwright’s Firefox and passes every assertion without cutting anything.doc.yuneta.io installs as a PWA. A manifest and its icons ship in
docs/doc.yuneta.io/pwa/, anddeploy.shinstalls them and injects the<link rel="manifest">into the<head>of every built page — the book-theme has no hook for the head. There is no service worker, so this buys a window and an icon, not offline reading. The manifest is served at/manifest.webmanifestby onelocationin thedoc.yuneta.iovhost alone, aliased to the/pwa/copy:yuneta.io,yuneta.com,yuneta.esandyunetas.comshare that docroot, but/on them is the landing page, so a manifest sitting at the root would offer five installs of a different app under one name.
Changed¶
c_yuno.hsheds 4 of its 9 public functions, and the two that stay say why.c_yunois one of the 4 gclasses inroot-linuxwhose header exports PUBLIC C functions. Four of them —add_allowed_ip(),remove_allowed_ip(),add_denied_ip(),remove_denied_ip()— had no caller anywhere: the add-/remove- commands that use them live inc_yuno.citself, so they arePRIVATEnow. Nothing outside changes; the operator interface (theallowed_ips/denied_ipsSDF_PERSISTattributes plus the six list/add/remove commands) was already complete.The rest of the header is kept on purpose, with the reason written next to it:
is_ip_allowed()/is_ip_denied()are asked once per accepted connection (c_tcp_s) and once per login (c_authz), where the canonical local method would cost akwper call; andyuno_event_loop()/set_yuno_must_die()are process-level, not gclass interface — every caller of the latter is ending its OWN process (a signal handler, a CLI told to quit, a test yuno that finished: 63 call sites across 49 test files), so an event would be a second door to the same thing. A justified escape that is written down is not the same as an oversight.c_authz’s two —authz_checker()/authentication_parser()— are not an escape at all: they implement kernel typedefs (authorization_checker_fn,authentication_parser_fn), are installed as process defaults byentry_point.cand swapped throughyuneta_setup(). The call goes downwards, from the kernel into the gclass, which no gclass-interface mechanism can express —authz_checker()does not even take a C_AUTHZ instance, it looks the service up itself.yuno_event_detroy()→yuno_event_destroy(). The typo had been in the header since the function was written. One in-tree caller (entry_point.c), no out-of-tree ones.C_TIMER/C_TIMER0: the timeout helpers are sugar now, and the behaviour lives behind the interface.set_timeout(),set_timeout_periodic()andclear_timeout()are three of the very few PUBLIC C functions a gclass header exports — 4 of the 33 gclasses inroot-linuxdo it, and CLAUDE.md says to treat those as defects. The arming therefore moved out of them and intomt_writing()on themsecattribute, which is the real interface:gobj_write_integer_attr(timer, "msec", 1000)now leaves the timer exactly asset_timeout(timer, 1000)does, where before it armed the countdown and left the gobj stopped. Same forC_TIMER0with its io_uring event.No caller changes and no behaviour change for anyone using the helpers — they write the same two attributes they always did, and every trace line is unchanged. What changes is that the gclass no longer has two doors with different behaviour.
c_timer0’s helpers had to swap their two writes (periodicbeforemsec), since themsecwrite is the one that arms.This is the C half of the same normalization shipped in
@yuneta/gobj-js7.9.10/7.9.11, where the split was doing real damage: JS had never started or stopped the gobj at all, so every view paired itsset_timeout()with agobj_start()and got “Destroying a RUNNING gobj” when it forgot the matching stop. The two ports are now the same contract, function for function.GOBJ.md§7 (the worked example isc_timer.citself) and the timer API page were rewritten to match.Not touched: the ESP32
c_timer(root-esp32), which arms withgobj_play()/gobj_pause()and cannot be built or tested here.
7.9.9¶
Fixed¶
A foreground process honours SIGTERM; a daemon still ignores it.
c_yuno.c’s signalfd handler marked SIGTERM// ignoredfor every process built on the framework, becausecapture_signals()runs unconditionally inmt_start. For a daemon that is deliberate — its watcher parent is deaf to everything,--stopkills with SIGQUIT then SIGKILL, and a stray SIGTERM frominitat shutdown must not take a node’s yunos down. For a CLI it was inherited by accident, and it breaks the Unix contract:timeoutnever escalates to SIGKILL on its own, sotimeout N ycommand …hung forever and left an immortal process behind — two were found alive after two and a half days, with the password still visible inps.SIGTERM now takes the same orderly path as SIGQUIT/SIGINT unless the process runs as a daemon, reported by the new
yuneta_is_daemon()(entry_point.h), which is true exactly when--startwas given. The agent launches every managed yuno with--start(c_agent.c), so no daemon changes behaviour. Verified on both branches:ycommandnow dies on SIGTERM andtimeout 8returns in 8 s, while a--started yuno ignores SIGTERM in both watcher and child and still stops cleanly with--stop.The standalone timeranger tools (
tr2list,tr2keys,tr2search,tr2migrate,treedb_list,msg2db_list,stats_list,fs_watcher) do not useyuneta_entry_point, so they carried their own copy of the same v6-era boilerplate; each now routes SIGTERM to its orderly quit handler instead of ignoring it.
Added¶
@yuneta/gobj-ui5.6.0 → 5.8.2 (submodule bump; its ownCHANGELOG.mdcarries the detail).The bottom toolbar of
C_YUI_FORMis configurable, and a toolbar with a single group is centred. A button the caller drops no longer breaks the form.yui_shell_confirm_danger()— a destructive confirmation whose button is RED.yui_shell_confirm_yesno()puts its yes inis-link, the right colour for “do you want to continue” and the wrong one for “this deletes an account”: the two read the same at a glance, and the destructive one is the one that must not be clicked by reflex. The safe answer is the last button, so Escape, the backdrop and the X all resolve to it.The treedb table’s search stretches on a phone, where its row used to keep its natural width and leave the most used control the narrowest thing on screen. Its placeholder was the literal
'search...'— a placeholder is not a text node, sorefresh_language()could never reach it and it stayed English in every language; it carriesdata-i18n-placeholdernow.Edit is a mode toggle in that table, not one more action.
A toolbar dropdown no longer opens off-screen. The panel was anchored to one edge of its trigger and only that edge was guarded against the viewport, so a right-aligned panel near the left of the bar hung off the screen — which is what every
navbar-endtrigger does underdir="rtl". It is clamped on both edges now; LTR positions are unchanged.
Changed¶
The JS yunos consume the libraries from npm, not by
file:(yunos/jssubmodule bump).gui_agentandgui_treedbpointed@yuneta/gobj-jsand@yuneta/gobj-uiat thekernel/js/*submodule checkouts, which tied a build to the superproject and to whatever those working trees held. They now resolve^7.9.6/^5.8.2from the registry, like wattyzer — so there are nofile:consumers left, and a local edit underkernel/js/**reaches an app only afternpm publishplus a range bump. Thefile:deps were symlinks and forcedresolve.preserveSymlinks, which loads duplicate module instances; that flag and thesrc/aliases are gone, anddedupenow also lists@yuneta/gobj-js.Both JS yunos install as a WebAPK. Each ships a complete manifest (
display: standalone,start_url,scope, 192/512 PNG icons and a maskable 512 variant) rendered from its existing SVG mark.gui_treedbhad asite.webmanifestthat was never installable: nodisplay(so it defaulted tobrowser) and an SVG as its only icon, which Chrome does not accept. No service worker is involved — Chrome no longer requires one for installability. The manifests declare noorientation:orientation: "any"overrides the device’s rotation lock, so the app rotates even when the user locked it. Both are namedmanifest.webmanifest, the name every SPA in the family uses, so the nginxlocationthat declares their MIME type is the same line everywhere —.webmanifestis absent from nginx’s stockmime.types, and without that block a manifest goes out asapplication/octet-stream.
Documentation¶
doc.yuneta.io links back to the landing page. The landing is raw HTML installed at
/landingand is the front door served atyuneta.io, but it sits outside the myst toc, so nothing in the built site pointed at it: a reader who entered throughdoc.yuneta.iohad no way to reach it. It is aHomeentry insite.navnow. The book theme hides nav items below 1024px and its hamburger only opens the toc, so two rules in_static/custom.csskeep the link reachable on a phone.A new standalone page at
/high-semantics— “High-level semantics, low-level language”: Yuneta in one page, with themachinetrace in the format it really prints, the same state table in C and in the browser, and a bill next to every decision. It argues rather than instructs, so it opens a third band on the landing (“The idea”,essay-card), with its ownESSAYSlist andcheck_bandguard indeploy.sh.The standalone pages moved above the argument on the landing. They were the last three bands before the site map: ten screens down on a 390×844 phone, on a page 13.6 screens long. A card nobody scrolls to is a card that does not exist.
/navigationwas also missing from the site map.
7.9.8¶
Fixed¶
A deferred answer never reached a browser:
input_servicedid not name a service.C_IEVENT_SRVrecorded it asgobj_name(gobj_parent(gobj)). That parent IS a service in the agent’s topology, and it is the CHANNEL when a SPA connects straight to a yuno through an iogate (C_IOGATE^__top_side__→C_CHANNEL^tcps-N→C_IEVENT_SRV^tcps-N). Anything that read the field back to route a deferred answer then looked the requester up under a name that is not a service, found nothing, and dropped the answer.That is why
register-idp-useransweredcommand-yunoand never answered a SPA — the browser sat waiting for ever, with no error anywhere. It is now the nearest service ancestor, which is the same value as before wherever the parent already was one.kc_answer()dropped an answer in silence when it could not resolve the requester. Silence there is indistinguishable from a request that never arrived, and that ambiguity is what hid the bug above. It logs a warning withop,req_serviceandreq_channelnow — a client that disconnects mid-flight is normal, hence warning and not error.list-idp-usersandget-idp-userasked Keycloak for the user-profile schema of every row. Measured against a real realm: 1.5 kB per account of which 230 bytes are the account, and the sameuserProfileMetadatablock repeated for each one — a page of 50 came to ~77 kB instead of ~12 kB.briefRepresentation=truedoes not suppress it;userProfileMetadata=falsedoes, and both commands send it now.
7.9.7¶
Added¶
Five commands to manage the accounts of the IdP, so an operator does not have to open the Keycloak web console:
list-idp-users(search + paging),get-idp-user,update-idp-user(name,enabled,emailVerified,requiredActions),delete-idp-userandsend-idp-user-actions(the invitation email again, with the actions you choose). All on theidpservice, all asynchronous, all through the same bounded queue and the same shared connection as the registration.They share one gated permission,
manage-idp-users, kept apart fromregister-idp-useron purpose: an operator who registers people does not have to be able to delete them.delete-idp-userdeletes in the IdP only. Thetreedb_authzsrecord stays. The two planes are deleted one by one so nobody loses an authorization record to a cascade they did not ask for;delete-userof theauthzservice is the other half.The page is bounded (default 50, ceiling 500). The realm is shared with every other product of the organization, so one call must not be able to pull all of it.
default_roleinC_AUTHZ(persistent, writable). When it reacts toEV_IDP_USER_CREATEDand the event carries no role, it links this one. Empty by default, so the user enters and can do nothing, which is visible and safe. No role is hardcoded, because the roles come from theinitial_loadof each realm and none is guaranteed to exist — and for the same reason the attr is checked before use: adefault_rolethat does not exist creates the user with no role and logs an error, instead of failing in silence for ever.A
messagestrace level inC_IDP_KEYCLOAK: the operation, the method, the resource, and the status of the answer. Deliberately not the request and not the body — the request carries theAuthorization: Bearer <admin token>header, which opensmanage-usersover the whole realm until it expires, and the body of a read is the account data of the realm. A trace turned on to debug a request must not leave either in the log.
Changed¶
The IdP request queue is generic.
KC_PENDINGcarried the fields of the registration in the struct (email,first_name,last_name,role); adding five operations that way would have been five copies of the same shape. It is nowIDP_PENDINGwith{op, params}, the job list is chosen fromop, and the five one-round-trip operations share one job pair instead of five near-identical ones. A new operation is a case in two functions and nothing else.BREAKING: IdP account provisioning leaves
C_AUTHZforC_IDP_KEYCLOAK.C_AUTHZanswers one question — does this user hold this permission, againsttreedb_authzs. Creating accounts in an external identity provider is a different responsibility, with different credentials (a confidential admin client withmanage-users), a different transport (C_PROT_HTTP_CLoverC_TASKwith its own queue) and different failure modes (the IdP down, a 409, an expired token). They shared a file by accident of how it was written.set-kc-config,view-kc-configandregister-idp-usermove unchanged to the new gclass, together with thekc_*attrs, the pending queue and the three-job Keycloak pipeline — about 600 lines out ofc_authz.c. The commands and the service name stay neutral (idp,register-idp-user), so a second provider enters as a sibling gclass serving the same vocabulary, selected in configuration: theytlspattern with OpenSSL / mbedTLS, and not an abstraction invented against a single implementation.What callers must change. Send
register-idp-user,set-kc-configandview-kc-configto the serviceidp, notauthz, and declare the service in the yuno config:{'name': 'idp', 'gclass': 'C_IDP_KEYCLOAK', 'priority': 0, 'default_service': false, 'autostart': true, 'autoplay': false, 'kw': {}}What operators must do. Persistent attrs live in
<GCLASS>-<service>-persistent-attrs.json, so whatset-kc-configwrote forC_AUTHZ-authzis not found byC_IDP_KEYCLOAK-idp, and the firstregister-idp-useranswerskc_unavailable. Move the sixkc_*keys to the new file with the yuno stopped — move, not copy:C_AUTHZno longer declares those attrs, so a key left behind logs “GClass Attribute NOT FOUND” at every start. The recipe is inYUNO_AUTH.md§7.3; runningset-kc-configagain also works, at the cost of the client secret on a command line. The permissions moved with their commands and are now permissions of theidpservice.New event
EV_IDP_USER_CREATED, the seam between the two planes. The provisioner publishes it after a created account and each plane records its own user:C_AUTHZwrites thetreedb_authzsnode it used to write from inside the Keycloak pipeline. The dependency points one way — the provisioner knows there is an authz plane to notify, and the authz plane knows of no provider — so the event is declared inc_authz.hand no authz code includes a provider header. A subscriber action returns 0 when it recorded the user; a negative return reaches the caller as the existingauthz_write_failedwarning. TaggedEVF_NO_WARN_SUBS: a yuno may provision accounts with no authz plane at all.New local method
has_roleonC_AUTHZ. The roles belong to the authz plane, and the provisioner has to refuse an unknown one before it creates anything in the IdP — which is whatregister-idp-userdid when both lived in the same gclass. A local method and not a public C function, because a gclass exposes itself through attributes, commands, events, local methods and statistics, and nothing else.
7.9.6¶
Added¶
--ssl-server-nameinycommand,ybatchandystats, and as a parameter ofconnectinycli. The name to check the server certificate against, used for SNI too. It travels toytlsasssl_server_name, which both TLS backends already implemented andc_tcpalready honored (“a config-supplied ssl_server_name wins”); no tool exposed it.This is what a pinned certificate needs. The agent serves one long-life certificate of its own on every node (
CN=yuneta_agent.yuneta.io), so its name never matches the host dialed, and every tool rejected the handshake for a hostname mismatch and then retried in silence. Nothing is weakened: the chain is still validated, against the PEM given in--ssl-trusted-certificate.ybatchandystatstake the whole TLS family (--ssl-use-system-ca,--ssl-trusted-certificate,--ssl-server-name,--ssl-allow-insecure-client). Their crypto was a literal{"ssl_use_system_ca": true}in the code, so a private CA was out of reach and awss://agent answered “unknown CA”.yclitakes the same four as parameters of itsconnectcommand, next to the url they belong to.
Changed¶
Submodules.
kernel/js/gobj-uimoves to the maplibre-gl floor raise (peer^6.0.0→^6.1.0, the version this SDK ships), andutils/python/tui_yunetasto the yunetas CLI 0.19.1, which stores the TLS values of a node and carries the node identity on everyycommandcall it makes.
Fixed¶
C_AUTHZhammered the IdP with a connection nobody asked for. The outbound client to Keycloak is created with manual start,process_next_kc()starts it, and nothing stopped it:C_TCPretries for ever, so the yuno reconnected once a minute for its whole life. One run of a deployed yuno had 759 of those round trips, every one publishing to nobody.ac_end_task()now stops the client when no request is left, andprocess_next_kc()starts it again for the next one. Beyond the noise, this gives each registration a fresh connection, so a late answer can no longer land on the following task. The stop on a timed-out round trip stays where it was — before the next task starts — because there the connection is poisoned even when more requests are queued.Measured against a stub of the IdP: three idle minutes, zero connections, where the same window used to open three.
set-kc-configdestroyed a running gobj. It dropped the cached client without stopping it first, which the framework reports as “Destroying a RUNNING gobj”. It stops it now.gobj_local_method()matched a prefix, not the name. The dispatcher compared withstrncasecmp()over the length of the TABLE entry, so a gclass whoseLMETHODtable helddo_itanddo_it_resultanswered BOTH lookups withdo_it: the second method was unreachable, and the job that asked for it silently ran the first one again. It now compares the whole name.This is what kept
register-idp-userfrom ever working.C_AUTHZnames its jobskc_create_user/kc_create_user_resultandkc_send_email/kc_send_email_result, so the create-user request went out twice and the handler that reads the answer of Keycloak — the status, and theLocationheader that carries the id of the new user — never ran. The task then failed on an empty user id and answered the catch-allkc_unavailable, which reads as an unreachable IdP. Verified end to end against a stub of the Keycloak admin API: token, then 201 withLocation, thenexecute-actions-email, and the caller reads “User registered”.BREAKING for a gclass that leaned on the prefix behavior: a lookup that used to reach a shorter entry now answers “internal method NOT EXIST”. A sweep of every
LMETHODtable in this repo found exactly two prefix pairs, both inC_AUTHZ, and both were the bug.A scalar answer was invisible in
ycommandandybatch. Their display handleddataonly as an array or an object, so a command that answers with a string, a number or a boolean printed nothing at all — legal JSON, read by nobody.node-uuidis how it showed: the uuid appeared inycli, which has its own display, and vanished throughycommandand therefore through the controlcenter scripts, which runycommand. The answer had left the agent well formed.node-uuidanswers like every other command. It put the uuid indataas a bare string. It now carries it incomment, prefixed with the yuno identity as the othercmd_*handlers do, and indataas{"uuid": "..."}. Clients no longer need to special-case a scalar to show it, which is the half of the fault that lives in the agent.ystatsnever sent its token. Its gobj tree defined__jwt__and no field used it, so every remote connection went out anonymous and the agent refused it with “Without JWT/passw only localhost is allowed”. Thejwtfield is now in theC_IEVENT_CLIkw, as inycommandandybatch.ystatsswallowed a failed open. Its FSM declared neitherEV_ON_OPEN_ERRORnorEV_ON_ID_NAK, so a refused connection was lost in “Event NOT DEFINED in state” and the tool retried for ever without a word. Both now reachac_on_close, which is whatycommanddoes.ystatsignored-o,-Oand-S. The connection was built with literals (yuneta_agent,__default_service__), so the three options existed and changed nothing.ycommandapplied the pinned name to the IdP too.build_client_crypto()serves two different peers, andssl_server_namenames ONE certificate: given to the token endpoint as well, the login died with a hostname mismatch before the agent was ever dialed. The builder now takespin_server_name, true only for the agent.C_AUTHZ:register-idp-usercould never work. The outbound client to Keycloak (C_PROT_HTTP_CL) is a CHILD gclass, so it subscribed its parent to every event it publishes, andensure_kc_client()never dropped that subscription. Three faults at once, all from the one cause:EV_ON_MESSAGEis not in the FSM ofC_AUTHZ, so every answer of Keycloak was lost in “Event NOT DEFINED in state”. The task ended without an answer and the caller got the catch-allkc_unavailable, “Keycloak register-idp-user failed”, which reads as an unreachable IdP.EV_ON_OPENis not in the FSM either, so the connection logged a second error.EV_ON_CLOSEis in the FSM, for a USER channel that closes. The close of this outbound socket therefore enteredac_on_closeand ran the logout of a user that does not exist (“session_id not found”, “User not found”, and “kw must be list or dict” from the treedb).
ensure_kc_client()now unsubscribes the parent right after it creates the client, which is whatC_AUTH_BFFalready does with its IdP client. The audience is theC_TASK: it subscribes togobj_resultsby itself.The path was never exercised before, because it needs a confidential IdP client that no deployment had. Verified against a real Keycloak with a
kc_admin_client_idthat does not exist: the answer is nowkc_token_refusedwith “Keycloak admin authentication failed (status 401)”, and the log carries neither the undefined events nor the phantom logout.C_AUTH_BFFnow logs what the IdP answered. An IdP failure that did not match a specific mapping was reported to the operator asauth_unexpected_errorand nothing else, becausesend_error_responselogs the browser-facing code, not the cause. A deletedclient_idwas therefore indistinguishable from an IdP outage, and a real deployment lost its login to exactly that (Keycloak answers 401 +invalid_clientfor a client that no longer exists).send_token_to_browsernow emits one line withaction,idp_status,idp_error,idp_description,client_idand the browser code it mapped to. It fires only for the three generic mappings (auth_unexpected_error,auth_config_error,auth_service_unavailable); the specific ones (invalid_credentials,session_expired,account_disabled,auth_rate_limited) already name the cause and stay on one line, so a wrong password does not double its log volume. Nothing that reaches the browser changes.C_AUTH_BFFmapsinvalid_clienton 401 too. The mapping accepted that IdP error only with status 400, and Keycloak answers 401 for aclient_idthat does not exist or whose secret is wrong. The commonest misconfiguration of all — a client deleted or renamed in the IdP — therefore reportedauth_unexpected_error(502), which reads as an outage of the IdP. It now answersauth_config_error(500), the code whose catalogue entry inc_auth_bff.halready said “BFF misconfigured (bad client_id, ...)”.BREAKING for a client that switches on the code: this case moves from
auth_unexpected_error/502 toauth_config_error/500. Both are already in the published catalogue, and no SPA in this repo branches on either.
7.9.5¶
Security¶
llhttp 9.4.3 (vendored,
kernel/c/gobj-c/src/llhttp*). Upstream fix: do not allow an empty transfer-encoding. The parser now rejects a request whoseTransfer-Encodingheader is present but blank, instead of accepting it, which is the class of ambiguity that request smuggling is built on.c_prot_http_srparses requests from the network, so this reaches every yuno that serves HTTP.The four vendored files were pristine 9.4.2 with no local modification, so 9.4.3 is a verbatim drop-in. Only the generated state machine and the version constant change:
api.candhttp.care byte-identical between the two releases. Verified with a clean SDK rebuild and the full suite, 119 of 119 passing,test_c_llhttp_parserincluded.
Changed¶
maplibre-gl 6.1.0. No breaking change, no worker or bundling change, so the v6 worker handling stays as it is: the bundle still emits
maplibre-gl-worker.jswith a.jsextension, which is what keeps the MIME type servable. Worth having are the renderer fix for raster tile sources with errored tiles, the terrain resource leak when switching configurations, and the tile/image race with an undefinedAbortController.gui_treedbpins^6.1.0, because an application takes the floor it was tested against.gobj-uikeeps its peer range at^6.0.0, because a library must not force the update on estadodelaire, hidraulia or wattyzer for fixes it does not itself depend on.
Added¶
@yuneta/gobj-js7.9.5: themachinetrace is back, aligned withgobj.c. The JS port had the trace lines written but disconnected —traceacame from a yuno attr and the calls ingobj_change_state, start/stop and create/delete were commented out — so the runtime that the browser SPAs are built on could not answer “what happened?” the way a node does. The C kernel’s level model is now in place, with the same names and the same bits:gobj_set_global_trace("machine", true),gobj_set_gclass_trace(...),gobj_set_gobj_trace(...), plus the no-trace veto by SOURCE, the union of global|gclass|gobj, andtimer/timer_periodicfiring for their own event only. Read it withset_log_callback(). See gobj-js’s own CHANGELOG.Fixed on the way:
log_error/log_warningreached forwindow.consoledirectly, so any error logged outside a browser threwReferenceError— the failure path replacing the failure it was reporting.Published to npm as a patch ahead of the SDK; consumers on the registry (wattyzer, estadodelaire, hidraulia) get it by bumping their range.
The Spanish artifact of
/navigationis in the repo, beside the page it translates (docs/doc.yuneta.io/navigation/artifact/), with the script that assembles it: content, the demos’ CSS from the English page,demos.jswith its strings translated, and gobj-js’s IIFE build inlined — an artifact serves no sibling files.deploy.shexcludesartifact/from the install, since it is source for a page published elsewhere, not part of this site.
Fixed¶
gui_treedbtalks to the shell throughyui_shell_of(), not through its parent.ac_child_selected(mirror the selected topic into the url) andac_remove_conn(the confirm dialog) took the parent to be the shell, which only holds while a view hangs off a route the shell itself declares — under aC_YUI_NODEtree the parent is the NODE, with nouse_hash, noitem_indexand noEV_ROUTE_REQUESTED. Latent here (gui_treedb mounts on declared routes); it is what actually broke yunovatios, whose treedb views moved under a node tree. TheEV_ROUTE_CHANGEDsubscription keeps using the parent on purpose — a subscription goes to whoever PUBLISHES — and its variable is now namedhostto keep the two apart.
Tools¶
scripts/check_ste.py— a linter for the documentation’s English. The docs are written to the structural rules of ASD-STE100: short sentences, one word one meaning, simple tenses, active voice, condition before command. Most of those rules need judgement, but a useful subset does not, and that subset is what this reports: contractions, semicolons, the banned modals (should/would/may/might/could), British spellings, Latin abbreviations, perfect tenses, phrasal verbs and the usual filler.It reads markdown table cells, which is where reference material lives, and it skips the Untouchables — code fences, inline code, link targets, URLs and double-quoted text, that last one because a quoted log line or command description is quoted material under rule 8.6 and must stay exact. Quote state carries across a line break, so a quotation that wraps is not scanned as prose.
--summarygives one line per file,--rule <id>filters to one category, and--strictadds the judgement calls that are off by default: a possessive is legal when it is correct, and “just” is filler in “just run it” but temporal in “the yuno you just built”. Exit 1 when anything is reported, so it fits a pre-commit hook.A clean run is not compliance. It is a spell-checker for the rules that need no judgement, and it says so.
Documentation¶
The declarative shell has a public demo: demo.yuneta.io. The
gobj-uitest-app was only reachable atniyamaka.com, a domain whose name says nothing about Yuneta, and no page linked to it. It now has its own host on the documentation box, with its own certificate, anddoc.yuneta.iolinks to it under See it run. It shows everyC_YUI_NAVlayout and the per-zone responsive model, with no backend and no login.The script that deploys it was untracked, listed in
.git/info/exclude, so nobody could reproduce the deployment from a clone.test-app/deploy.shreplaces it: the host is an argument,demo.yuneta.iois the default,niyamaka.comstays for mobile testing, and it curl-verifies the result instead of reporting success on the rsync alone./guide-folders: thetoolsentry linked to the CHANGELOG. The list of top folders links each name to its section below, and[tools](#tools)resolved to/changelog#toolsinstead — a reader who clickedtoolsleft the guide entirely.#toolshad no owner on the page as far as myst was concerned, so it matched an implicit heading elsewhere in the site and took the first one it found.The file already carried the answer: its
yunosentry is(folders-yunos)=, renamed at some point for the same reason.toolsis now(folders-tools)=and resolves on the page. myst reported this as “Linking ‘tools’ to an implicit heading reference”, which reads like a style note and is why it survived so long.The whole documentation set is rewritten in Simplified Technical English — 194 files:
docs/doc.yuneta.io/**and the eleven onboarding chapters underyunos/c/yuno_agent/. The content did not change. Every fact of the previous text survives, and the sentences that carried it are shorter.check_ste.pyreports zero on all of it.Two defects turned up that were not style.
DEBUGGING.md§1 announced “Two destinations” above a diagram of four, with a sentence below it that already said “all four destinations”. Anddeploying-yunos.mdandNODE_SEALING.mdboth buried a node-wide SIGKILL under the command block that causes it, so a reader who followed a recipe top-down had already killed the node before reading the warning. Both now carry a CAUTION above the commands, which is the rule for a safety instruction: the risk level first, then the command, then the result.What changed in the prose. British spellings became American. Contractions are expanded.
shouldis gone where it stated a requirement, because a reader treats it as optional. Each concept keeps one term, soenableanddisablereplaced a rotation ofturn on,activate,armandswitch. Semicolons became sentences, phrasal verbs became plain verbs, ande.g.,i.e.andetc.are written out. The alt text of every diagram was re-punctuated too: a screen reader reads it as continuous prose, and every one of them was a run-on joined by semicolons.Untouched throughout: commands, code blocks, identifiers, quoted log lines and quoted source strings. The command description “WARNING: Don’t use in production!” keeps its contraction, because it is the literal string in
c_agent.c.The landing’s live trace INDENTS, like the kernel’s. Every line sat at column zero, so the one thing a machine trace exists to show — that this event was fired from inside that action — was invisible.
tab()ingobj.cwrites 2 spaces per level of__inside__; the panel now does the same, and the sequence carries the two publicationsC_TCPreally makes (EV_CONNECTED/EV_DISCONNECTEDto its iogate) so there is a nested level to see. The state chips keep trackingC_TCP: a nested line belongs to another gobj and leaves them alone.@yuneta/gobj-js7.9.6:tab()indents 2 spaces per level, likegobj.c— it was2n - 1, one short at every level, and its floor was zero where the C version’s is one, so a JS trace read beside a node’s did not line up. Published to npm./login-flow: the transport controls moved INSIDE the graph box, under the canvas. They drive what the canvas shows, and sitting below the stage meant reaching past a screenful of text to pause the thing you were watching. The row of twelve numbered step buttons is gone with them: it was longer than the rest of the bar, and the arrows already walk the sequence. In its place,5|12between the arrows — the number belongs to what the arrows move, and← 5|12 →is one glance — with a fixed width and tabular figures so a number changing twelve times a cycle does not shove the→button. Speed is a setting rather than transport, so it wraps to its own row: on a phone the bar reads as two lines, on a desktop as one./navigation: the three mechanisms are live, and they run on gobj-js. Each demo is a real GCLASS on the real runtime (the 7.9.4 ES build ships next to the page): a card click sendsEV_OPEN_CARD, the tree sendsEV_GOandEV_SET_NAV_MODE, the pager sendsEV_PUSH_PAGE/EV_POP_PAGE, and a panel under each demo shows which event landed in which state. A page whose argument is a click IS an action had no business making it out of DOM callbacks. Thenav_modedemo is the one that earns its keep: the same tree at the same depth, redrawn as stacked strips, a single← parent, or one breadcrumb — the table’s three rows, live.The panel under each demo is the kernel’s own
machinetrace, not a log the page writes:gobj_set_global_trace("machine", true)+set_log_callback(), the same two calls a yuno makes on a node. That is what the gobj-js fix above was for — the trace did not exist when the demos were written.New walkthrough:
/navigation— “Getting back”. Navigation in a gobj-ui app is one decision — where the reader’s position lives — and the page is built on that axis: the url (C_YUI_NAV, cards + subpath), the node tree (C_YUI_NODE, one declared route and free depth), and the pager’s in-memory stack (C_YUI_PAGER). It exists becausestack/back/pathread like three ways to navigate when they are the three values ofnav_mode, and all three live inside the url: they choose how the way back is drawn, not where the position lives. Taking"stack"for the pager’s stack inverts the one distinction that matters./login-flow: the player was stuck on step 1 for anyone with “reduce motion” on.restart()returned before scheduling anything whenprefers-reduced-motion: reducematched, soplaying = truehad no effect and Play did nothing — the walkthrough only moved with the arrows. Honouring the preference means not animating a packet along a wire nobody asked to watch; it does not mean refusing to turn the page the reader pressed Play to turn, so the step now advances on a timer with no travel./login-flow: the graph fits a phone. Below 700 px it is drawn turned — the same topology with its axes swapped, its five columns spread as five rows — instead of keeping the 700 px floor and letting the container scroll sideways, which on a phone meant watching a third of the graph while the packet animated off-screen.
7.9.4¶
Ships with @yuneta/gobj-js 7.9.4 and @yuneta/gobj-ui 5.4.0.
A kernel fix that had been paid for three times at the call site, and the three ways of showing depth in a node tree become one runtime knob.
Fixed¶
gobj_destroy()now actually stops the gobj it complains about (kernel/c/gobj-c/src/gobj.c, and identically inkernel/js/gobj-js). Destroying a live gobj is the caller’s bug and the kernel has always said so out loud — “Destroying a RUNNING gobj”, “Destroying a PLAYING gobj” — and then tried to repair it withgobj_stop()/gobj_pause(). That repair could never run:obflag_destroyingwas raised first, and both entry points refuse a destroying gobj. That refusal is correct and stays — nobody OUTSIDE may stop something already being dismantled — but it meant the rescue died on its own guard, logging a second, misleading “hgobj destroying”, andmt_stop/mt_pausenever ran: the gobj was taken apart still holding its timers, subscriptions and children.The pause/stop now happen before the flag goes up, so
mt_stopsees exactly what an orderly stop sees. That is not cosmetic ordering: a gclass that stops its children inmt_stopgoes throughgobj_stop_children(), which carries the same guard — with the flag already up, the whole subtree would have stayed running. Both complaints now also carryLOG_OPT_TRACE_STACK, since the useful information is who destroyed a live gobj.The fix at the call site is unchanged:
gobj_stop_tree()beforegobj_destroy(). What changed is that the framework no longer pretends to repair it when you forget. The trap had been diagnosed and fixed at the caller at least three times in the JS GUIs alone. New unit tests in gobj-js (tests/destroy_stops.test.js) pin the order and fail against the old one.scripts/check_doc_line_refs.py --repinreaches the site’s “Current version” stamp. The regex required a path after the tag, so a link to the repo ROOT (/tree/7.9.2) never matched — and that link is the front page’s version stamp, the most visible version number the site has. It was left behind on every release and corrected by hand, which is the exact failure mode the script exists to remove. The path is optional now, and a link whose TEXT is just the tag gets the text swapped too.
Changed¶
@yuneta/gobj-ui5.3.2 → 5.4.0:nav_mode, the three shapes of a node tree as one knob. AC_YUI_NODEtree can be asked for stacked strips ("stack", the default and what it declares), a single← parent("back") or the trail as one breadcrumb ("path") —yui_node_set_nav_mode(root, mode)at runtime, or"nav_mode"in the root’s declaration. All three shapes were expressible before, but only by rewriting the declarations:"back"is not a root-level edit (projections do not inherit, so every branch had to be rewritten) and going back was lossy, since restoring “stacked” meant imposing a canonical shape on branches that had declared their own. A mode is now a filter applied when the renders are asked for, so"stack"is an exact restore, per branch. The knob is per tree, so an app can run one section as a breadcrumb and another as a backbar.With it, the rule that was missing from the docs: chrome belongs to the node that declares it, so every branch declares its own. The library was already right — a node’s backbar goes back to its route — but the README’s own example showed the shape that breaks it (the pair on the root, a child with only an
index), and an app that copies it ends up with a single ← that reads “← root” at every depth. 5.3.3 also fixed alinknode’s viewer being sized by its own content instead of by the body.
7.9.3¶
Ships with @yuneta/gobj-js 7.8.7 and @yuneta/gobj-ui 5.3.2.
Two auth_bff fixes finish the session-restore path opened in 7.9.1, and the
UI library turns navigation into a tree of gobjs.
auth_bff:/auth/refreshanswers with the identity too. It returned only{success, expires_in, refresh_expires_in}, which made session restore impossible to finish: the tokens are httpOnly, so after a reload JavaScript has no other way to learn who the user is — the SPA came back authenticated but anonymous, avatar on?, until the next full login.usernameandemailare already decoded from the access token for every action, so this only ADDS fields.auth_bff: oneerror_code, one HTTP status. The discovery-drain path added in 7.9.1 answered503withauth_service_unavailable, butc_auth_bff.hpins that code to502and503is alreadyserver_busy. Clients branch on the status as well as the code, so one code answering with two statuses is a defect however transient both happen to be.@yuneta/gobj-ui5.2.1 → 5.3.2: navigation becomes a tree of gobjs.C_YUI_NODElets a section declare its own subtree, so depth stops being a flat route table: a node projects its children as anindexor aschrome(tabs, cards), and the newprojection.pathadds a third mode — a breadcrumb drawn from the tree ROOT, for branches where one strip per level becomes a wall.yui_node_set_chrome_depth()makes that cap reachable at runtime, which the config could already declare but the API could not change. The shell root is itself a node now (config.shell.tree).Around it: the site map became a navigation PANEL instead of a transient overlay — it joins the window manager when the app has one, draws each subtree once (a route reachable from three surfaces repeated its whole branch three times), and hides reference rows behind a toggle without ever emptying a menu.
remember_section_position(opt-in) returns a menu click to where the reader was inside that section, andC_YUI_JSONgrew depth guides. Indentation is four spaces wherever structure is shown as indentation, rendered trees indenting inchso the guides stay lined up with the text at any zoom.The same line shipped broken twice, for a reason worth recording:
keep_on_navigatewas verified only in an app with a window manager, where a dock-managed window registers no overlay at all — the one configuration in which the flag is not used. 5.3.1 fixed a different bug on it (gobj_find_service()answersundefined, not null, and handing that to aDTP_POINTERattr fails the whole kw), 5.3.2 the flag itself, and_qa_routingnow drives the shell contract directly instead of the app.Also in 5.2.1: clicking a node in the schema graph threw
ReferenceError: gobj_send_event is not defined— the module publishedEV_NODE_CLICKwith it and never imported it, breaking the schema landing in every consumer.yunos/js: the console dump indents with four spaces, the last two-spaceJSON.stringifyleft in either SPA. Both also carry the site map’s two new i18n keys: the library translates through the APP’s i18next, so a key missing there renders as the key itself, in lower-case English, and never changes language.Docs: the 7.9.2 packaging incident is written up at
/package-transition, withverify-package-transition.sh— a harness that reproduces the delete with real dpkg in a temp root and then runs the SHIPPED hooks against a fake one. The login exchange is a running graph at/login-flow.
7.9.2¶
Install this instead of 7.9.1. On a node coming from 7.9.0 or older, 7.9.1 deletes the web server configuration it was meant to protect. If a node already runs 7.9.1 it is past that transition and is not affected.
The upgrade no longer DELETES the configuration it stopped shipping. 7.9.1 removed
nginx.conffrom the payload so an upgrade could not overwrite it. It cannot — but dpkg and rpm delete files that an upgrade no longer provides, so the very transition that fixed the overwrite removed the operator’s config instead. It took down the GUIs on two nodes, and the two managers failed differently:deb — dpkg dropped the file, then
postinstseeded the stock default, so nginx served a configuration with no vhosts.rpm —
%postruns before the old package’s files are erased, so the seeding was skipped and the erase left nonginx.confat all; nginx would not even start.
The configuration is now saved before the manager can touch it and restored afterwards: a new
preinst(deb) /%pre(rpm) copies it tonginx.conf.pkgsave, and the restore prefers that copy over the stock default. On rpm the restore moved from%postto%posttrans— the only hook that runs after the erase; anything%postwrites is wiped moments later.Both webservers are covered,
nginxandopenresty, each keeping its own file. Verified with real dpkg (the delete on upgrade is reproducible), then by running the shippedpreinst/postinstand%pre/%posttransagainst a bind-mounted/yuneta, for the upgrade case and the fresh-install case.preinstshipped non-executable. The blanketfind DEBIAN -type f -exec chmod 0644re-flattened it after its ownchmod 0755, anddpkg-debrefuses a maintainer script that is not executable — the package would not build at all.
7.9.1¶
A recovery release: after a machine reboot, a node came back with its login wedged and stayed that way. Three defects were behind it, at three layers.
c_prot_http_clsentEV_DROPas a string literal, so the event never matched. Events are compared by POINTER identity (_find_event_action), never by text, so the interned symbol and a"EV_DROP"literal are two different things.ac_timeout_inactivityused the literal, which produced the misleading “Event NOT DEFINED in state / C_TCP / ST_CONNECTED / EV_DROP” — naming an event the table does declare — and left the outbound to an unreachable peer never dropped. Present since80e6f7ad3and in 7.8.7 too: it only bites once the inactivity timeout actually fires, which is precisely what an unreachable IdP causes. It was the only such literal in the tree.c_auth_bffnow bounds the wait for an IdP it cannot reach, and retries.C_TASK’sexec_timeoutdoes not cover this: it is armed insideexecute_action, andexecute_actiononly runs once the channel is connected (c_task.c,mt_start). While the IdP is unreachable the task waits for a connection that never arrives — unarmed and unbounded — so noEV_END_TASKis ever published,discovery_donestays FALSE, and every browser request queues behind it forever with nothing logged. A reboot produces exactly that: the yuno starts before the network is usable.The BFF now owns the deadline, since it is the one with browsers waiting:
a
C_TIMERwatchdog over the connect gap only, with its own budget (idp_connect_timeout_ms, default 30 s).idp_timeout_mskeeps timing the round-trip once connected — conflating the two failed an IdP that is merely slow to appear.ac_on_opendisarms the watchdog the moment the outbound connects, so the two never time the same thing twice.the stuck task is failed through its own path (
EV_TIMEOUT→stop_task(-2)), so the existingEV_END_TASKhandling applies unchanged: 504 for a request, a drained queue for discovery.a failed discovery is no longer terminal:
process_nextre-arms it on the next request, so recovery rides on a retry rather than on a background timer — there is nothing to poll for, the endpoints are only needed at login.queued requests are answered 503
auth_service_unavailableinstead of being left on a socket that will never produce bytes.on fire, a connect that burned its whole budget is aborted, so the next attempt starts fresh instead of waiting out the kernel’s SYN ladder (
tcp_syn_retries=6→ ~127 s); and when a request needs an outbound that sits disconnected, it is nudged withEV_CONNECT—c_tcp’s own on-demand entry point — instead of waiting out a backoff already grown to its 30 s cap.
Verified against a blackholed IdP: login answered 503 in 22 s instead of hanging, back to 200 in 0.6 s once the IdP returned without restarting the yuno, 0.3 s in steady state.
The
.deb/.rpmno longer overwrite the node’s web server configuration. The packagers stage/yuneta/bin/nginxwholesale, andconffiles/%config(noreplace)only cover/etc, so every upgrade replaced the operator’snginx.conf— with the stock one built in CI — and shipped the build machine’sconf.d/on top.%fileslists/yunetaas a directory, so the file cannot even be tagged%config(rpmbuild: “file listed twice”). Both packagers now stripnginx.confandconf.d/from the payload — a file the package does not contain cannot be replaced — andpostinst/%postseednginx.conffrom the pristinenginx.conf.defaultonly when there is none. Verified on a built.deb.CI: the packaging actions moved onto the Node 24 runtime. Every run of
release-packages.ymlcarried the annotation “Node.js 20 is deprecated … being forced to run on Node.js 24: actions/checkout@v4, softprops/action-gh-release@v2”. GitHub was compensating; when it stops, both steps fail, and since this is the repo’s only workflow that means a release with no.deband no.rpm— found while cutting one. Bumped to the latest majors that declareusing: node24(checkoutv4→v7,action-gh-releasev2→v3); neither breaking change touches this workflow. Verified with aworkflow_dispatchagainst a throwaway tag — not against7.9.0, whose packages must keep matching the tagged tree — both jobs green and zero Node 20 annotations.
7.9.0¶
Ships with @yuneta/gobj-js 7.8.7 and @yuneta/gobj-ui 5.2.0.
The headline is a BREAKING change in the agent: delete-yuno no longer
deletes a whole yuno by omission. See the entry below for why the old default
was the wrong way round.
BREAKING (agent):
delete-yunono longer deletes the whole yuno by omission — it requireswhole=1. Withoutyuno_release=the command targeted the in-memory primary, i.e. the yuno and every release behind it. That put the destructive reading behind the SHORTER command line, and the safe one behind the longer: an operator reaching for “drop this release” and forgetting the parameter got “drop the yuno”. It has already cost a realm itsauth_bff, withdelete-yunocascading onto the config row.The three forms are now explicit:
delete-yuno id=<id> # refused, with both options named delete-yuno id=<id> yuno_release=<rel> # one release delete-yuno id=<id> whole=1 # the yuno and all of its releasesforce=1does not stand in forwhole=1— it bypasses the snap-tag guard, a different question, and its help line now says so. Nothing in the tree called the bare form (not theyunetasCLI, the sync tools or the agent SPA); the callers were operators followingREALMS.mdandYUNO_LIFECYCLE.md, both updated.create-yuno/delete-yuno/find-new-yunoshelp lines name what they actually do. All three act on a yuno or on one of its releases, with a parameter as the discriminator, and the old one-liners (“Create yuno”, “Delete yuno”) hid it — the flow to adopt a new binary/config version isfind-new-yunos create=1+deactivate-snap(bundled asyunetas upgrade-yunos), and readingcreate-yuno/delete-yunoas a symmetric pair is what leads to inventing a destructive substitute for it.Also dropped three alias arrays that pointed each of those commands at its own name: they add nothing to
command_get_cmd_desc(which matchesnamefirst whenever the command has ajson_fn) and surfaced in the help JSON as an alias identical to the name, advertising a spelling that does not exist.delete-user(C_AUTHZ) can now delete a role-holding user, and reports its result truthfully. Two problems fixed together: (1) the command always returned the comment"User deleted"while passinggobj_delete_node’s code through verbatim, so a failed delete surfaced asERROR -1: User deleted; (2) it could not delete any user linked to a role, even though roles are not a deletion boundary — a local password user (e.g. an MQTT/IoT gate identity) can hold roles too, and the seed admins are protected by immutability, not by having roles.delete-usernow: checks immutability up front and refuses cleanly (force-proof, no session reject); refuses a role-holding user unlessforce=1is passed (withforceit unlinks the roles first, viagobj_delete_node’s force option); and returns an accurate result in every path.pm_delete_useradds theforceboolean. A new regression testtests/c/command_delete_user/drives a real C_AUTHZ over a temp store and covers all four cases. This also fixes a latent use-after-free incmd_delete_user: it passed the user node toEV_REJECT_USER(whose handler readsusernameand frees the kw) and then reused that freed node ingobj_delete_node— a double consume that also made session-rejection a silent no-op (the node keys onid, notusername). It never crashed in production (freed-but-unreused memory) but the new test caught it underCONFIG_DEBUG_TRACK_MEMORY.EV_REJECT_USERnow gets its own{username}kw, so it actually kicks sessions andgobj_delete_nodeowns the single node ref.dns_warn_due()(static_resolv.c) silences-Wformat-truncation. The helper took aconst char *nsGCC could not prove NUL-terminated within its 46-byte source array, sosnprintf("%s")worst-cased a read spanning the wholechar[3][46]into the 46-byte slot.parse_resolv_conf()already bounds every entry, so no real truncation was possible; the copy is now precision-capped (%.*s) so the bound is provable. No behaviour change.Scaffolding: new projects and yunos get a
CHANGELOG.md, not aCHANGES.txt. Theyuno-skeletontemplates (c_project,yuno_citizen,yuno_standalone) emitted a bareCHANGES.txtthat nothing in the toolchain read and no project kept up. Every repo in the ecosystem keeps a Keep-a-ChangelogCHANGELOG.mdinstead, so the skeletons now start one.is_service_authorized()no longer trusts everyauthorized_servicesentry to be a string. The list is filled byac_identity_cardfromjson_object_foreachkeys, so entries are always strings today — but a non-string entry would have madejson_string_value()return NULL andstrcasecmp(NULL, …)crash the yuno. A non-string entry now logs aMSGSET_INTERNALerror (broken invariant) and is skipped.
7.8.7¶
A release the toolchain asked for. Ubuntu 26.04 brought gcc-15 and clang-21,
whose <string.h> hands back a const char * where the code expected a
writable one, and the fourteen warnings that surfaced pointed at one place
that really was editing memory it did not own. Following that thread through
the inter-event layer turned up four identity checks that had drifted apart
from each other and from how the framework matches names everywhere else.
The ievent identity checks agree with each other, and with the framework.
dst_roleanddst_yunowere checked in four places that all disagreed.C_IEVENT_CLIcompared only the part before a^, and it obtained that part by writing a NUL over the^and restoring it afterstrcmp()— over a buffer owned by jansson, sincekw_get_str()returnsjson_string_value().C_IEVENT_SRVdid no^handling at all, and compared case-insensitively while greeting (identity_card) but case-sensitively for every message after it, so a peer could pass the handshake and be rejected by its first event.Yuno, role and service names are a case-insensitive namespace —
gobj.cregisters services throughstrntolower()and matches names withstrcasecmp(). All four checks now do the same, the^handling is gone (nothing in the tree ever put a^indst_yuno; it came from the pre-v7 wire format), and no one edits a jansson string any more.The cross-service authorization gate stops rejecting on letter case.
is_service_authorized()matched the names inauthorized_services— which come from the roles written in the treedb — against the live service withstrcmp(), whilegobj_find_service()resolves them lowercased. A role grantingTreeDBtherefore resolved the service and then failed the gate. It now usesstrcasecmp(), like the rest of the naming. This only turns false denials into the access the role already granted; it cannot widen a grant, since the name still has to be in the list.The string helpers stop discarding
const. glibc’s<string.h>now declaresstrchr/strrchr/strstras the C23 const-generic macros, so on aconst char *they returnconst char *— and thirteen call sites that stored that in achar *started warning under gcc-15 / clang-21. All of them only did pointer arithmetic or printing, so nothing changed at runtime; they are now declaredconst char *. The one site that does write through the pointer (gobj_search_path(), splittinggclass^name) keeps an explicit cast, since the string it edits issplit2()'s own heap copy.install.shrefreshes the apt index before installing the.deb. It never did, and a freshly imaged Debian node carries whatever index its image was built with. Debian keeps only the current version of each package in the pool, so apt asked for superseded files and got 404s —rsync 3.4.1+ds1-5+deb13u1while the mirror already served+deb13u4,libpython3.13 3.13.5-2against3.13.5-2+deb13u3— the dependencies went unmet andyuneta-agentwas left unpacked but unconfigured. Re-running the installer did not recover: nothing in it refreshed the index, so the node stayed stuck until someone ranapt updateby hand. Two nodes installed minutes apart from the same release differed only in how old their image was. (This one reaches every node immediately:install.shis fetched frommain, not from a release asset.)
7.8.6-4¶
Another packaging revision — YUNETA_VERSION stays at 7.8.6, only RELEASE
moves. Everything here is about what an operator is told when something goes
wrong: a clean install on a fast Debian 13 dedicated server failed to install
certbot and reported three symptoms and no cause, and a Rocky node reported its
firewall work in words that sent the reader to the command that says
“FirewallD is not running”.
The Debian certbot helper stops hiding why it failed, and retries. It ran
snap wait system seed.loaded 2>/dev/null || trueandsnap install core || true: the one safeguard against snapd’s start-up race, silenced, plus the first two real errors swallowed — so a failed run showed only the third line and no cause. snapd restarts itself right after being installed, and a store request in flight when that happens dies withcontext canceled, which reads like a network fault and is not one. The seed wait is now reported (and bounded at 180s, sincesnap waithas no timeout of its own and a broken snapd would hang the installer), the installs retry three times, and a genuine failure prints what each error class means plus the commands that tell them apart —uname -r,snap debug sandbox-features,snap changes,journalctl -u snapd.Both helpers now also document how to add a DNS-01 provider plugin, which neither installs. On RHEL it is one
dnf install python3-certbot-dns-<provider>. On Debian the plugins are separate snaps and the first command is the one nobody guesses —snap set certbot trust-plugin-with-root=ok— without which the install dies complaining about trusting the plugin author, inside a box snap prints in a way that is easy to miss entirely.It also warns instead of staying quiet when
/usr/bin/certbotis a real file rather than the snap symlink. Debian’s owncertbotpackage is not a drop-in replacement: it has nodns-ovhplugin (the archive ships cloudflare, desec, google, infomaniak, rfc2136, route53 and standalone only), so a node whose certificates carryauthenticator = dns-ovhcannot renew them with it — and installing it takes over/usr/bin/certbotand adds a second renewal timer beside snap’s, both aimed at the same/etc/letsencrypt.The
.rpm’s firewall message names the branch it took, and the command that checks it. It saidopened …whether the ports had been applied to a running firewalld or written to the permanent config of one that has not started yet. An operator reading that reaches forfirewall-cmd --list-ports, which talks to the daemon: on a fresh node — firewalld enabled but not started, which is the normal case right after installing — that answers “FirewallD is not running” and reads as a failure when nothing failed. Each branch now says what it did and points at the matchingfirewall-cmd/firewall-offline-cmdverify command.
7.8.6-3¶
A packaging revision, not a new version of Yuneta: no source under kernel/,
modules/, utils/ or yunos/ changed, so YUNETA_VERSION stays at 7.8.6
and only the RELEASE counter moves. The packages are rebuilt as
yuneta-agent-7.8.6-3 and attached to the existing 7.8.6 tag.
Revision 2 was withdrawn without being announced: its .deb still came off
the ubuntu-22.04 runner, so the archives it ships were stamped glibc 2.35
while Debian 13 runs 2.41. libc_guard.cmake compares the two as an exact
string, so that package installs and runs but leaves the node unable to build
anything against the SDK it just dropped there. The number is burned rather
than reused: two different payloads must never share one version string.
The AMD64
.debis built on Debian 13. It came off anubuntu-22.04runner while the.rpmhad already moved into arockylinux:9container — the asymmetry that made the whole glibc-provenance rule necessary in the first place. Both jobs are now containers with the same step order. Droppinglibpcre3-devwas part of it: PCRE1 no longer exists in Debian 13, and nothing links it — the build vendors PCRE2 and openresty wants PCRE2 too, so it only survived because it still existed on Ubuntu.installation.mdtold Debian users to install it, which fails on trixie.
What prompted it: clean installs on Debian 13 and Rocky 9 both finished on a green tick while leaving something broken behind them — a node that would stop answering at its first reboot, certificates that would not renew, and an apt that could not tell whether the payload fit on the disk.
The two install.sh fixes below are listed for the record but were live the
moment they were pushed: that script is fetched from main, not from a
release asset.
The
.rpmopens the ports the node serves.fail2ban— a weak dependency — pullsfail2ban-firewalld, which pullsfirewalld, which its own scriptlet enables. firewalld is not started during the transaction, so a fresh Rocky node worked right after installing and would have started refusing connections at the first reboot — the very thing the installer recommends doing. Nothing in the package touched the firewall. The%postnow opens 1993/tcp (the agent’s external control plane) and 80/443 (the bundled web server), viafirewall-offline-cmdwhen firewalld is enabled but not yet running. A node whose ports were moved from the defaults still needs them opened by hand; the failure path says so instead of passing silently.The
.rpmcertbot helper starts the renewal timer. certbot’s scriptlet enablescertbot-renew.timerwithout starting it, and says so — meaning renewals did not begin until the node rebooted, while the Debian side gets a live timer from the snap in the same run. The helper printed the command to fix it instead of running it.The
.debdeclaresInstalled-Size.dpkg-deb --builddoes not compute the field — onlydpkg-gencontroldoes, and the packager writesDEBIAN/controlby hand — so apt counted the package as taking zero bytes: a 294 MB.debannounced “After this operation, 49.0 MB of additional disk space will be used” (the sum of its dependencies alone). apt’s disk-space check therefore passed on a box with no room for the payload, and the install filled the disk instead of refusing up front.install.shverifies the agent is running before declaring success. The postinst starts the service throughinvoke-rc.dand ignores the result; on a systemd box the init script’s output goes to the journal, so nothing about the start ever reached the terminal. Acurl | shrun ended on a green tick whether or not anything came up. It now waits for the process, names what is running, and exits non-zero with where to look when the agent is not. The init script’s ownstatusexit code is unusable for this: it reports the web server — absent on a fresh node — not the agent.The Debian developer toolchain installs
wget.install.shpromised it in its own header and the.rpmlist already carried it; only the.deblist did not.install.shno longer makes apt drop its download sandbox. apt fetches as_apteven from a local file and could not traverse the 0700 root-ownedmktempdirectory, so it fell back to an unsandboxed download and said so. Cosmetic, but it was noise introduced by the 7.8.6 move toapt-get install.
7.8.6¶
A release about the first minutes of a node’s life. Installing on a clean Ubuntu printed a wall of dependency errors and rescued itself; the two distro families named the same helper differently; and nothing in the docs said that the SDK a package drops on a node can only compile there when the glibc matches — which, on Ubuntu 26.04, it does not.
install.shinstalls the.debwithapt-get, notdpkg -i.dpkgdoes not resolve dependencies, so every clean install printed a wall of dependency errors and left the package unconfigured until theapt-get -ffallback rescued it —gdb, new in 7.8.4, made this fire on every node.apt-get install ./pkg.debresolves deps from the file directly; thedpkgpath stays only as a fallback for an apt too old to take a file argument.The certbot helper is
install-certbot.shon both families. Debian shipped it asinstall-certbot-snap.sh, so operators andinstall.shhad to know which distro they were on to name the same helper. The old name remains as a symlink.Docs:
installation.mdgains the glibc-provenance rule and a verify-a-fresh-install checklist. The page promised that the sparse SDK lets projects “compile against the published runtime” without saying that this only holds when the node’s glibc matches the one the package was built against — a mismatch links silently and corrupts the heap at run time. Since the AMD64.debis built on ubuntu-22.04 (glibc 2.35), Ubuntu 26.04 nodes are runtime-only.yunetasCLI 0.18.0:--helpis now an extended help. It printed the same one-line-per-command summary that bareyunetasalready shows, so the flag told you nothing you had not just seen. It now documents every command and option, grouped by job. Bareyunetaskeeps the compact listing andyunetas <command> --helpkeeps click’s rendering.Docs: the CLI page documents the node registry and secret overlays. Both shipped in 0.15/0.16 and the command map still listed neither.
7.8.5¶
A release about where things live. The deploy tooling was split across two
release channels — the CLI on PyPI, the Python tools inside the packages — and
the halves drifted until a pipx upgrade alone could break a deploy. They are
one package now. The .rpm likewise stops being built on Ubuntu, and configs
stop carrying credentials into git.
The Python deploy tools move into the CLI.
sync_binaries,sync_configsandset_start_prioritiesship inside theyunetaspackage asyunetas.agent_tools.*(CLI 0.17.0) instead of being read fromtools/agent/at run time. One tool released through two channels meant the halves drifted: a node found running CLI 0.14.0 against scripts from 7.8.4 would have broken on apipx install --upgrade yunetasalone, handing--secrets-dirto a script that had never heard of it. The files undertools/agent/are now deprecated forwarding shims — operator runbooks reference those paths — and go away in a release or two. The CLI still shells out toycommand, so this removes the version skew, not that dependency. CLI 0.17.1 also fixes the dependency declaration (typer[all]names an extra typer no longer provides;richis now declared, since the CLI imports it).Secret overlays for config deploys (
sync_configs.py --secrets-dir, andyunetasCLI 0.16.0 wiring it to~/.yuneta/secrets/<node>/). A committed config declares a credential with the value"__SECRET__"; the value lives only on the deploy machine and is deep-merged in just before the push. The shape of the config stays versioned in git — which is what makes it reconstructable — and only the value is withheld. Fails closed: a surviving"__SECRET__"refuses the push rather than shipping an empty password, which for SMTP means an auth failure or an unauthenticated send.Prompted by an SMTP password committed in cleartext in a project repo. It does not fix that one — that is still in git history and needs the credential rotated.
The merged copies hold real credentials, so they are written 0600 into a 0700 temp dir removed in a
finally, on SIGINT/SIGTERM/SIGHUP, and swept at the start of the next run. That last one is not belt-and-braces: a SIGKILLed run (timeout, OOM) skips every handler, and testing this left exactly such a plaintext copy behind.yunetasCLI 0.15.0: a node registry, so deploying to another machine is--node <name>instead of a url plus four OAuth2 flags re-derived from a config file each time.register-node/list-nodes/unregister-nodeback it with~/.yuneta/nodes.json(0600), and--node/-Nworks onsync,sync-binaries,sync-configsandupgrade-yunos. It stores where a node is and which identity you present, never a credential — those come from the environment at call time. A node registered with--sshinstead of--urlis reached by forwarding a free local port to its loopback-bound agent, torn down after the command. See the deploy guide.The
.rpmis now built on Rocky 9, not on the Ubuntu runner. Both packages came out of a singleubuntu-22.04job, on the stated assumption that fully static binaries run on RHEL too, so a separate EL9 build was unnecessary. That is true of the binaries and false of the archives: each package also shipsoutputs/lib+outputs_ext/libfor nodes that compile their own yunos, and those are tied to their building glibc. So the.rpmshipped glibc-2.35 archives onto EL9 nodes running glibc 2.34 — the same defect that destroyed an Ubuntu 26.04 node, surviving on Rocky only because the two versions are adjacent. Verified on a live Rocky 9 node: itslibyunetas-gobj.areportedGCC: (Ubuntu 11.4.0).The workflow is now two jobs, each building its package on its own distribution, and each asserts its payload’s glibc stamp before packaging — the EL9 job additionally requires the stamp to equal the container’s own glibc, so this cannot silently regress.
7.8.4¶
A release about a lie the build was telling. A node with a compiler could
compile its own yunos against the archives shipped in the package, the link
would succeed, and the binary would corrupt its heap the moment it ran. It
cost a day of chasing a refcount bug that did not exist — the stack traces
pointed at kw_decref and jansson, and both were innocent bystanders. The
build now refuses that link instead of producing the binary, and the CLI stops
reporting success when the refusal happens.
The build refuses to link prebuilt archives against a different glibc. The packages ship prebuilt static archives (
outputs/lib,outputs_ext/lib) built by CI onubuntu-22.04, i.e. glibc 2.35. A node that has a compiler and project sources can compile its own yunos against them, and withCONFIG_FULLY_STATICthe link then mixes those archives with the node’s static glibc. That is not portable: the archives reach into glibc internals such as_dl_x86_cpu_features, which back the ifunc resolvers selecting the CPU-tunedmemcpy/strlen. A dynamic link fails loudly with an undefined reference; a static link resolves silently against a layout that changed between releases, the resolver picks a wrong routine, and the heap is corrupted at run time.The failure is brutal to diagnose because it looks like someone else’s bug: SIGABRT/SIGSEGV inside
unlink_chunk/_int_free_merge_chunk/_int_malloca couple of seconds after start, no Yuneta error logged first, and a stack trace blaming whatever innocent code happened to free next. Found on an Ubuntu 26.04 node (glibc 2.43) where four yunos crash-looped ~130 times from their very first launch and had never once run.tools/cmake/libc_guard.cmakenow records the building glibc next to the archives (outputs/lib/yuneta_libc.stamp, written by thegobj-cbuild and shipped in the package), and every other build compares it against the glibc of the machine doing the linking, failing at configure time with the three ways out: build off-node and ship binaries, build the whole SDK from source locally, or install a package built for that distribution. Archives older than the stamp warn instead of failing. Both packagers now refuse to build a package whose payload has no stamp.Note this guards glibc only. The compiler is deliberately not checked: rebuilding the same sources on the affected node with clang instead of GCC still crashed, so the compiler is not the variable — only the libc is.
gdbis now a hard dependency of the agent package (Depends:in the.deb,Requires:in the.rpm). A Yuneta node is expected to produce core dumps in/var/crash(7.8.2/7.8.3 went to some length to make sure it does), and a core is useless on a node with no debugger: analysing one meant copying a multi-hundred-MB file off the node, or installinggdbby hand at exactly the moment something is already broken. It is a dependency rather than a recommendation because--no-install-recommendsis common on server installs, which is precisely where the crash will happen.yunetasCLI 0.14.0:initno longer printsinit doneand exits0after a cmake it just watched fail. It now reportsinit FAILED, lists the directories, and exits1. The glibc guard above fired correctly on a node and the CLI declared success anyway, which is exactly the combination that makes a real error invisible.build’s failure path also stops using the bareexit()(absent underpython -S) and exits1instead of255.
7.8.3¶
A packaging release, and a lesson about who else wants your kernel knobs.
7.8.2 fixed /var/crash losing a tug-of-war with kdump; this one fixes
core_pattern losing the same kind of fight to Ubuntu’s apport. Validated
end to end on an Ubuntu 26.04 node: a deliberate SIGSEGV now leaves a real
core in /var/crash, which it did not before.
apportis disabled on.debnodes, andcore_patterntaken back. On Ubuntu,apport.servicestarts aftersystemd-sysctl.serviceand overwrites/proc/sys/kernel/core_patternwith a pipe to itself, so the value from/etc/sysctl.d/99-yuneta-core.confsurvived the install (the post-install runssysctl --system) and was lost at the next boot. Since apport discards cores from binaries that did not come from a distro package, yuno cores then stopped existing with nothing logged — found on a node reading|/usr/share/apport/apport …while the Rocky node next to it still had/var/crash/core.%e. The package now shipsyuneta-core-pattern.service(After=apport.service, ordering-only so it is harmless where apport does not exist) which re-applies the file, and the post-install warns vialoggerifcore_patternstill is not ours — a generic check, so it also catchessystemd-coredumpor RHEL’sabrt. The post-install now also turns apport off (enabled=0+systemctl disable --now): apport is Ubuntu’s crash telemetry client, shipping deduplicated reports toerrors.ubuntu.comfor Canonical’s benefit, and it deliberately discards crashes of binaries outside its packaging allowlist —/yunetais not on it. On a dedicated appliance node that trade is all cost. Order matters and is load-bearing:apport --stopdoes not restore our value, it writes the bare wordcore, which drops dumps into the crashing process’s CWD — so the sysctl is re-applied after apport is stopped, never before. Same shape as the/var/crashfix in 7.8.2: a one-shot setting losing to a later actor, so re-assert instead of fighting. The RHEL/Rocky equivalent (abrt-addon-ccpp) is not covered yet — it is not installed on our nodes, but it would do exactly the same thing.
7.8.2¶
A release about a failure that could not be seen. A node came up with a
black-holed nameserver first in /etc/resolv.conf; every name resolution paid
~6 s, resolution runs synchronously inside the event loop, and a yuno building
25 channels spent ~2 min 40 s in start up — long past the agent’s handshake
timeout, so the agent reported it as not running while the process sat there
alive and listening. Nothing in any log said any of that. The fixes below are
in the order they were needed: a way to trace start up at all, then the fix,
then the two warnings that would have made the whole hunt a one-line grep.
Core dumps survive a
kexec-toolsupdate (.deband.rpm)./var/crashis co-owned: on RHEL/Rocky kdump’skexec-toolsalso declares it, asroot:root 0755. The packages’ one-shotchmod 0775/chown root:yunetain the post-install therefore held only until the next transaction touching that package, which reverted group and mode — and then cores silently stopped being written, withrpm -V yuneta-agentreporting.....UG.. /var/crashfor anyone who thought to look. A/usr/lib/tmpfiles.d/yuneta-crash.confdrop-in now re-asserts it on every boot (and immediately, viasystemd-tmpfiles --create).kernel.core_patternis unchanged.getaddrinfo()reports when it blocks the event loop (yev_loop.c). Resolution is synchronous and the loop calls it while arming a connect, a source bind or a listen, so a slow resolver does not delay one socket — it stops every gobj, timer and pending completion in the process. All three call sites are timed, and over 1 s they emit agobj_log_warning(msgsetOS) carrying the host and the elapsedmsec, attributed to the gobj that paid it. This is what makes “the yuno is slow” legible as “resolving X stopped us for 7032 ms”.The static resolver leaves a trail in syslog (
static_resolv.c). It sits below the gobj log and cannot reach it, so it now writes tosyslog(3)directly — a stack buffer and no allocation, unlikeprint_error(), whichmallocs its message and is therefore the wrong tool for reporting thatmallocfailed. It reports an unresponsivenameserverwhen there is another to fall back to (rate limited to one per nameserver per 5 min, since this runs on every connect), any resolution over 1 s, an empty/etc/resolv.conf, and allocation failures that used to return a silentEAI_MEMORY.The static resolver caches its DNS answers (
static_resolv.c, soCONFIG_FULLY_STATICbuilds). Every connect used to re-query, so a yuno building N channels to the same host paid the full round-trip N times — invisible while DNS answers in a millisecond, fatal when it does not. A node whose/etc/resolv.conflisted a black-holed nameserver first cost the A and AAAA timeouts (~6 s) on every connect, andyuneta_getaddrinfo()runs synchronously inside the event loop, so it blocked the whole process: anauth_bffwith 25 channels spent ~2 min 40 s in start up, never sent itsagent_clientWebSocket handshake in time, and the agent gave up on it (1 raised, 0 reached) while the process sat there alive and listening. With the cache the same yuno starts in ~7 s and the agent reportsyuno up. Entries are held for the answer’s own TTL (now parsed instead of skipped), clamped to 5..300 s; the table is fixed-size and allocation-free; only the DNS step is cached, not numeric literals or/etc/hosts; failures are not cached, so a recovered IdP is picked up at once. This mitigates the blast radius, it does not make a brokenresolv.conffree: the first lookup still pays it.--global-trace=LEVELenables global traces from the command line. Repeatable and comma-separated (--global-trace=machine,create_delete), with--global-trace=listprinting the available levels. Levels are applied right after every gclass is registered, so they are already live when the first service starts — which is the whole point: until now a yuno that failed before it could reach the agent could not be traced at all, becauseset-global-tracetravels over the very control channel that is missing. The fallback waskill -10 <pid>(SIGUSR1 cycling the global mask), which cannot catch anything that happens during start up. Unknown levels are rejected with a pointer tolistrather than being ignored.--verbose-logis not a trace switch — it only overrides the stdout log handler’s field bitmask — and its help text now says so.memlockadded to the packaged resource limits (.deband.rpm). io_uring rings are pinned memory charged againstRLIMIT_MEMLOCK, and the budget is per user, shared by every yuno running asyuneta. A yuno withio_uring_entries=32768pins ~3.1 MB, so the usual 8 MB default admitted only two of them: from the third on,yev_loop_create()failed withENOMEM, the yuno aborted at startup and the ydaemon watcher relaunched it indefinitely. Freshly installed nodes with gigabytes of free RAM could not bring up their full yuno set. The limits drop-in, the init script andprofile.dnow all raise it.yev_loop_create()names the culprit onENOMEM: the warning and the critical now carry the ring’s estimated footprint (ring_bytes) and the effectivememlocklimit, and the critical adds a hint pointing atulimit -landio_uring_entries. The warning no longer calls the condition “transient pressure” — a fixed ceiling is not transient, and the wording sent readers looking for a memory leak.
7.8.1¶
A bugfix release. Most of it came out of watching a node come back from a cold machine reboot with its IdP (Keycloak) still booting — the state where half of these paths run for the first time. Two crashes and several noisy-but-harmless log storms were fixed, plus one refcount contract that had been quietly wrong.
Two crashes on the login path:
controlcenterdereferenced a not-yet-open treedb for any login arriving before it played, andprot_tcp4hre-sentEV_RX_DATAto itself after a synchronous publish had already dropped the connection.authz’s login publish became a veto point so a service can refuse a user it cannot yet register (see YUNO_AUTH.md §4.9).A gbuffer double-free: a kw carrying a serialized gbuffer must be refcounted with
kw_incref/kw_decref, never thejson_*pair —c_task’s lmethod forward andc_iogate’s broadcast both got it wrong. The rule is now in CLAUDE.md and the tree was swept.The agent launched a slow-to-open yuno twice at boot (and
run-yunocould too), which a controlcenter lost to its treedb lock. It now marks a launch until the yuno registers, expiring the mark on a monotonic timer.Log-noise fixes for the IdP-down window: the bad-json cascade on ievent ports collapsed to one capped warning, and
auth_bffstopped logging two stack traces per request when a 5xx body isn’t the RFC-6749 JSON envelope.
No BREAKING changes. Ships with the same @yuneta/gobj-js 7.8.0 and
@yuneta/gobj-ui 4.0.0 as 7.8.0 — this release is C-only.
- **fix(auth_bff): a 5xx from the IdP with a non-JSON body logged two errors
per request.** `send_token_to_browser()` read `error` / `error_description`
off the response body to map the RFC-6749 error envelope, but that envelope
only exists when the IdP itself answered. When the IdP is **down**, the 5xx
body is a reverse proxy's HTML 502 or empty, so `kw_get_dict(kw, "body")`
returns NULL and the two `kw_get_str()` calls path into NULL —
`kw_find_path()` logs *"kw must be list or dict"* with a stack trace, twice
per failed request (seen as 10 of those against 5 `👤BFF server error`).
The envelope is now read only when the body really is a dict; otherwise
`idp_err`/`idp_desc` stay `""` and the generic `auth_unexpected_error` (502)
mapping applies, which is the correct outcome for an IdP outage. The 5xx
`👤BFF server error` itself is legitimate and unchanged — the operator does
want to know the IdP is down.
- **fix(iogate): broadcasting to two or more channels double-freed the
gbuffer.** `send_all()` handed each child `json_incref(kw)` while every
child `KW_DECREF()`s it, and `kw_decref()` drops the serialized binary
fields on every call. So the gbuffer took one decref per child plus
`send_all()`'s own, against the single `kw_incref()` of the caller: the
arithmetic only worked out with exactly one open channel, and from the
second on it was a double free (*"BAD gbuf_decref()"*). Now `kw_incref()`.
`send_one_rotate()` was never affected — it hands its own reference to a
single channel. Same defect as the `c_task` one below; the rule ("a kw is
refcounted with `kw_incref`/`kw_decref`, and events always carry a kw") is
now in CLAUDE.md's API footguns. The remaining `json_incref(kw)` event
sends in `root-linux` and `utils` (timer, ievent_cli, yuno, authz, ota,
the three CLIs) are corrected too: none of their kws carries a gbuffer
today, so they were latent, not live.
- **chore(gobj-c): `load_persistent_json()`'s cannot-open critical now
carries a stack trace.** It names the file but not who was opening it,
which is exactly what you need when two processes race for the same
exclusive lock.
- **fix(agent): a yuno slow to open was launched twice at boot.** The boot
runs in two sweeps — `run_util_yunos()` for the `yuno_tag=util` yunos,
then `run_enabled_yunos()` for everything else, `timerStBoot` later — and
both skip a yuno only when its `yuno_running` is TRUE. That flag is set in
`ac_on_open()`, i.e. when the yuno registers back, so a yuno that takes
longer than `timerStBoot` to open is still marked not-running when the
second sweep arrives and gets launched a second time. `run_yuno()` now
marks the yuno as launching and both sweeps honor the mark, which
`ac_on_open()` clears. The `run-yuno` **command** honors it too: it
guarded on the same `yuno_running`, so an operator launching a yuno that
was still coming up got a second instance the same way. The mark carries
its launch time and expires after `timeout_expiration` (30s, the window
the command counters already give a yuno to connect back): a yuno that
dies before opening never clears its mark and the agent doesn't watch
pids, so without the expiry a failed launch would block `run-yuno` for
that yuno until the agent restarted. This only ever bit at machine boot: on a warm
`yshutdown` + `restart-yuneta` every yuno registers well inside the
window. Seen on a controlcenter (a `util` yuno that loads a treedb, unlike
logcenter/emailsender): the second instance died on timeranger2's
exclusive `__timeranger2__.json` lock, logging a CRITICAL. That lock is
what kept two masters off the same store — a `util` yuno without a tranger
would simply have run twice, with nothing to log.
- **fix(task): forwarding a kw with a gbuffer to an lmethod double-freed the
gbuffer ("BAD gbuf_decref()").** `C_TASK`'s `ac_on_message()` handed the kw
to the job's lmethod with `json_incref(kw)`. A kw carrying a serialized
binary field has its own symmetric pair — `kw_incref()`/`kw_decref()` —
because `kw_decref()` drops the binary on *every* call, not only on the
last one. `json_incref()` bumps the JSON refcount but not the gbuffer's,
so the two `KW_DECREF()`s that follow (the lmethod's and the action's own)
decref the gbuffer twice against a single incref: it is freed early, and
the publisher's final `KW_DECREF()` then reads a freed header. Now uses
`kw_incref()`, like `c_iogate`/`c_qiogate` do when forwarding.
Only reachable when the kw actually carries a gbuffer, which for the HTTP
task path means a non-`application/json` response body (`ghttp_parser`
parses JSON into `kw["body"]` and keeps the gbuffer instead) — in
practice, an error page from a reverse proxy. Seen in production as 75
"BAD gbuf_decref()" matching 75 OIDC discovery failures 1:1, while
Keycloak was still booting and nginx answered 502 + text/html.
- **fix(prot_tcp4h): parsing a buffer whose frame dropped the connection
raised "Event NOT DEFINED in state".** `ac_process_payload_data()` ends by
re-sending `EV_RX_DATA` to *itself* to parse whatever is left in the
buffer. Just above, `frame_completed()` publishes `EV_ON_MESSAGE`
synchronously — and a subscriber may drop the whole chain from inside that
publish (an authz NAK, or a peer sending bad json). The FSM is then in
`ST_DISCONNECTED`, where `EV_RX_DATA` is not defined, so the self-send
logged an ERROR plus a full stack trace for every such connection.
`frame_completed()` already knew about this cascade — it guards its own
`start_wait_frame_header()` with the same state check — but it returns 0
regardless, so the caller had no way to tell it was already dead. The
leftover self-send is now skipped once we are `ST_DISCONNECTED`: nobody
upstream wants the rest of the buffer. The sibling self-send in
`ac_process_frame_header()` needs no guard — it follows a plain
`gobj_change_state()` with no publish in between, so the connection cannot
have died under it.
- **fix(ievent): one garbage packet no longer costs four ERROR entries and
two stack traces.** A peer sending non-JSON to an ievent port (a port
scanner is enough) walked a cascade of logs that all described the same
event: `gbuf2json()` logged "json_load_callback() FAILED" with a full
stack trace, `iev_create_from_gbuffer()` logged "gbuf2json() FAILED",
`ac_on_message()` logged "iev_create_from_gbuffer() FAILED", and none of
them named the peer. All three were `gobj_log_error` — errors are for our
own broken invariants, and a stranger sending junk is not one.
`C_IEVENT_SRV`/`C_IEVENT_CLI` now ask for the parse silently
(`verbose = 0`, which `iev_create_from_gbuffer()` newly honors for its own
log too) and emit a single `gobj_log_warning` under `MSGSET_PROTOCOL`,
carrying `peername`/`sockname` and a dump of the offending bytes capped at
`MAX_LOG_DUMP_SIZE` (256), mirroring `c_prot_tcp4h`. The dump reads from
the gbuffer head, so it survives the parser having consumed the data and
leaves the read pointer alone. Behaviour is unchanged: the connection is
still dropped. Callers passing a non-zero `verbose` keep the old logs.
- **feat(authz): a subscriber of `EV_AUTHZ_USER_LOGIN` can now refuse a
login.** `mt_authenticate()` published the event fire-and-forget and threw
away the result, so a subscriber unable to accept the user had no way to
say so — the login succeeded regardless, and the user was left
authenticated but unregistered downstream. It now checks
`gobj_publish_event()`'s return and answers `result: -1` ("Some subscriber
refusing user") when it comes back negative. Contract note for out-of-tree
gclasses: an action handling this event that returns a negative value now
denies the login; every in-tree subscriber (`c_controlcenter`, `c_agent`,
`c_mqtt_broker`) returns 0 and is unaffected. Two caveats worth knowing:
the checked value is the *sum* of the subscriber returns, and a subscriber
holding `__own_event__` short-circuits before the accumulation, so its
refusal is not seen. The auth failure paths also log `peername`/`sockname`
now.
- **fix(controlcenter): a login arriving before the service played crashed
against a NULL treedb.** `C_CONTROLCENTER` subscribes to the authz service
in `mt_start` but only opens `treedb_controlcenter` in `mt_play`, so every
login landing in that window called `gobj_get_node()` / `gobj_create_node()`
with a NULL gobj and logged "hgobj NULL or DESTROYED" with a stack trace —
loud at startup, when many agents reconnect at once. `ac_user_login` and
`ac_user_new` now check the treedb is open and refuse the login until the
service reaches `mt_play` (fail-closed, riding the authz contract above),
so a user never gets in without its controlcenter record.7.8.0¶
A feature release built around the C_TRANGER / C_NODE read surface — the
command set a GUI needs to browse a timeranger2/treedb without pulling whole
topics into the browser: server-side key filtering/sorting/paging
(list-keys), cursor pagination (open-iterator / get-page), realtime feeds
(open-rt / close-rt), the restored record-read commands, and
print-tranger’s bounded, drillable raw-JSON dump. Alongside it, a
timeranger2 audit pass that fixed several ways a failed append could be
reported as success.
No BREAKING changes: the one signature correction is a header/implementation mismatch (prototype names only, no ABI change), and the one event-payload change is additive.
Ships with @yuneta/gobj-js 7.8.0 and @yuneta/gobj-ui 4.0.0 (a
MAJOR — five BREAKING contract changes; see its CHANGELOG before upgrading a
v2 SPA).
- **feat(treedb): C_NODE gains a `print-tranger` command.** Mirrors
C_TRANGER's: it dumps the tranger the treedb lives on as bounded,
`kw_collapse()`-truncated JSON, and accepts `path=` to lazily drill into
one subtree (arrays indexed by numeric position, so the path round-trips
through `kw_find_path`). This is what the gui_treedb "Raw JSON" viewer
calls to inspect a treedb's raw tranger; a primitive path returns an
explicit error rather than a silent null.
- **feat(gobj-c): `kw_collapse()` now accepts a top-level array.** Drilling
`print-tranger path=<array>` used to fail because `kw_collapse()` required
a dict at the requested path; it now collapses a top-level array too (new
`collapse_array`, mirroring the array-value branch of `collapse()` —
object elements recursed, arrays/primitives copied shallow, element paths =
numeric index). Dict output is byte-for-byte unchanged, and only
`print-tranger` calls it. Covered by `tests/c/kw/test_kw1` (top-level-array
collapse shape + `kw_find_path` round-trip + primitive-rejected).
- **fix(timeranger2): the realtime feeds ignored `only_md`, always reading
and delivering the record body.** A feed opened with `only_md` wants a
record's metadata but not its content — the historical iterator honors it,
but both realtime paths (the master append fan-out to rt_mem lists, and
`publish_new_rt_disk_records` on a follower) always called
`read_record_content()` and handed the callback the full body. On rt_disk
that was an extra disk read per live record, and within a single `only_md`
list historical rows arrived md-only while live rows carried full content.
They now honor `only_md`: the only_md feeds get NULL `jn_record`, and the
rt_disk read is skipped when no audience of the key needs the body. No
consumer relied on the old behavior (c_tranger already synthesizes an
md-only record for a NULL `jn_record`; tr_queue and the mqtt broker use
`only_md` only for one-shot historical dumps). Covered by
`tests/c/timeranger2/test_rt_disk_multi_feed` (an rt_disk and an rt_mem
`only_md` feed asserted to receive metadata, never a body).
- **fix(treedb): `treedb_create_node()` could return a dangling pointer.**
When the primary already existed (`save_id` false) and every listed pkey2
had no `indexy` (a schema/index inconsistency), the node landed in no
index, yet the final `json_decref` dropped its only ref and the freed
pointer was returned. It now frees the node and returns NULL when it
reached no index.
- **fix(timeranger2): two fs_watcher edge cases.** (1) the inotify read
buffer leaked when `yev_create_read_event()` failed (which only happens
before it takes ownership of the gbuffer). (2) after a concurrent
watch-descriptor removal `get_path()` returns NULL, and every event branch
except IN_DELETE built a `"(null)/..."` path or handed the callback a NULL
directory; the stale event is now skipped once, up front.
- **fix(tr2migrate): a failed record append was counted as migrated.** The
migration callback ignored `tranger2_append_record()`'s return and bumped
its counters before appending, so a failed append still counted toward the
migrated totals. It now counts only after a successful append.
- **fix(mqtt): `tr2q_append()` enqueued a garbage entry when the append
failed** (the broker session queue). Like `trq_append2` it ignored
`tranger2_append_record()`'s return, building a `q2_msg_t` from an
uninitialized `md_record` (bogus `rowid`) that later readers would trust;
and because it had already `KW_EXTRACT`'d the gbuffer out of `kw`, the
failure path also leaked that gbuffer. It now frees the extracted gbuffer,
decrefs `kw`, and returns NULL on a failed append.
- **fix(timeranger2): `tranger2_append_record()` reported success after a
file-open failure.** Both write stages are gated by `if(fd >= 0)` with no
else, so when `get_topic_wr_fd()` failed for the content or the md2 file
the function fell through and returned 0 (success) — persisting an index
row with a bogus offset/size (content open failed), or content bytes with
no index row and `g_rowid` 0 fed to the realtime feeds (md2 open failed).
It now returns -1 at either failure, like the sibling lseek/write errors.
- **fix(timeranger2): use-after-free in `fs_watcher` `remove_watch()` under
`TRACE_FS`.** `path` (the `IN_DELETE_SELF` caller passes `get_path()`'s
borrowed string, aliasing an entry of `jn_tracked_paths`) was read by the
trace and by the `inotify_rm_watch` error log AFTER `json_object_del()`
freed the backing string. It is now snapshotted before the delete.
- **fix(tr_queue): `trq_append2()` enqueued a garbage entry when the append
failed.** It ignored `tranger2_append_record()`'s return, so on failure it
built a `q_msg_t` from an uninitialized `md_record` (bogus `rowid`/`__t__`)
that later readers would trust. It now decrefs `kw` and returns NULL on a
failed append, matching `tr_msg`/`tr_msg2db`.
- **fix(tr_msg2db): `msg2db_close_db()` could leak the shared descriptor
with concurrent msg2dbs.** It released `topic_cols_desc` with an
unconditional `JSON_DECREF` (which nulls the global), so closing one of two
open msg2dbs nulled the global while the other still held a ref — leaked on
the second close. It now uses the same refcount-guarded decref as
`treedb_close_db`.
- **fix(treedb): a failed treedb/msg2db open leaked the shared column
descriptor.** `treedb_open_db()` / `msg2db_open_db()` incref-or-create the
module-global `topic_cols_desc` before validating the schema, but the two
early error returns ("No topics found", "TreeDB ALREADY opened") skipped
the matching decref — and a failed open is never paired with a
`close_db()`, so the descriptor stayed alive to `gobj_end()` (a leak under
`CONFIG_DEBUG_TRACK_MEMORY`, failing ctest). Both paths now undo the
incref/create with the same refcount-guarded decref `close_db` uses (plain
`json_decref` while another open still holds it — the double-open case —
else `JSON_DECREF` to free and null the global). Covered by
`tests/c/tr_treedb_hook_hygiene` (a no-topics open + the end-of-test memory
check).
- **fix(timeranger2): a corrupt md2 file was detected, logged, then read
anyway.** `load_first_and_last_record_md()` logged "md2 file corrupted" on
a negative or misaligned `lseek(SEEK_END)` but did not return, falling into
`if(offset >= sizeof(md2_record_t))` — where the signed `offset` is promoted
to unsigned, so a negative (lseek-error) offset passed the guard and a
truncated (misaligned) file was read one record short, both feeding a bogus
"last record" into the key cache. It now closes the fd and returns -1 at the
corruption point, so the caller aborts instead of caching garbage.
- **fix(treedb): force-deleting a node with an array hook skipped every
other child and then aborted.** The down-link teardown in
`treedb_delete_node(force=1)` iterated the parent's hook array while
`_unlink_nodes()` removed each child from that SAME array in place, so the
index-based loop stepped over the shifted tail: with three children it
unlinked the 1st and 3rd, left the 2nd, and the re-check then found a
leftover down link and refused the delete — leaving a half-unlinked graph
persisted on disk (the cleared children were already saved) while
reporting failure. The teardown now snapshots the child refs before
unlinking. Covered by `tests/c/tr_treedb_hook_hygiene` (force-delete of a
config with three linked yunos).
- **fix(treedb): `parse_schema()` validated every schema against an EMPTY
descriptor.** It built a local column descriptor but passed the
module-global `topic_cols_desc` to `parse_schema_cols()` — and that global
is NULL until the first `treedb_open_db()`. `parse_schema()` is the
validate-before-open helper the gclasses call at `mt_start`, i.e. before
any treedb is open, so `json_array_foreach(NULL, ...)` ran zero checks and
ANY malformed schema passed unvalidated. It now validates against the
descriptor it builds.
- **fix(timeranger2): an md2 open failure poisoned the key's cache with a
~1.8e19 row count.** `load_first_and_last_record_md()` returns -1 on a
failed `open()`, but its type was `uint64_t`, so the caller's
`if(file_rows < 0)` never fired: -1 became `UINT64_MAX` and the cache cell
was built with that as its row count. A follower reloading a key whose
`.md2` was unlinked mid-append (`tranger2_delete_key` racing an append)
then ran `publish_new_rt_disk_records` with `to_rowid = UINT64_MAX`, a
near-unbounded read loop. The function is now `json_int_t`, the guard
fires, and both callers bail (or skip the file) on NULL.
- **fix(timeranger2): a deleted key left its watermark behind in every disk
feed.** `published` holds one mark per key, and nothing dropped it when
the key died: a keyless feed on a topic that cycles its keys (an hourly
bucket) kept a mark for every key that ever existed, for as long as it
lived. Worse, a key RE-CREATED in the same file inherited the corpse's
mark — and a mark above the reborn key's rowids is a ceiling, not a
watermark, so its records were served to nobody. The mark now dies with
the key, on both delete paths (the master's `tranger2_delete_key()` and
the follower's `FS_SUBDIR_DELETED` branch, which share
`fire_key_deleted_locally()`). Covered by
`tests/c/timeranger2/test_rt_disk_multi_feed` (a key deleted and
re-created in the file its mark counted in).
- **fix(c_tranger): `list-keys` answered an UNSORTED page as if it were
sorted.** When the sort could not allocate its row array it gave up and
returned, and the command went on to page the untouched list and answer
OK: `from`/`limit` then cut it at positions that mean nothing, and the
client — which had asked for `order` — had no way to tell. The sort now
says whether it sorted, and the command is refused with the reason.
- **perf(c_tranger): `list-keys` sorts with `qsort`, not an insertion
sort.** The ordered variant inserted each key by linear scan — O(n²)
with a jansson refcount round-trip per swap, inside the event loop;
`order=records` over a topic with enough keys to need `list-keys` at
all could stall the yuno for seconds. Ties on the record count now
break by key, so equal counts page deterministically (`qsort` is not
stable).
- **fix(timeranger2): the first record of every new md2 file reached NO
realtime disk feed** — a Live card left open across midnight silently
dropped one record a day, per key. A feed's `published` watermark counts
rows IN A FILE (`read_md()` seeks `(rowid-1)` records into `<file_id>.md2`,
and the batch bounds come from that file's cache cell), but the topic
ROTATES its file (`filename_mask`, by default one a day). The mark was
stored per key alone, so at the rotation yesterday's mark — say row 226651
— met a file whose rowids restart at 1 and became a ceiling instead of a
watermark: `from_disk_rowid` sat above `to_rowid` and the batch was served
to nobody. It self-healed on the next append (the mark is overwritten with
the new file's row count), which is exactly why it read as "a record went
missing" and not as a broken feed. The mark now carries the file it counts
in: a mark of another file is no mark, and the feed is re-seeded at the
new file's start.
With it, the rowid a disk feed is GIVEN is now the GLOBAL rowid of the key
(the file's base plus the position in it) — what the callback's contract
promises (*"global rowid of key"*) and what the master's rt_mem path
already delivered (`update_new_record_from_mem()` returns `g_rowid`). The
follower's disk path was handing out the file-relative one, so `g_rowid`
in the published `__md_tranger__` was wrong after the first rotation and a
consumer that dedupes by it (the Live cards' key counter) took the new
file's records for records it already had. Covered by
`tests/c/timeranger2/test_rt_disk_multi_feed` (a rotation phase: the first
append of the new file reaches each feed exactly once, with the key's
global rowid).
- **fix(timeranger2): a realtime DISK feed re-broadcast every other feed's
wake-up, so N feeds on a key meant N copies of every record for each of
them.** With a per-key Live card and a whole-topic Live card open on the
same key (gui_treedb), every row appeared DUPLICATED in both. The master
hard-links each new md2 into the directory of EVERY feed that wants the key
(`master_to_update_client_load_record_callback`), so a follower is woken
once per feed — but `publish_new_rt_disk_records()` then fanned the new
records out to EVERY feed of that key, turning N wake-ups into N x N
publishes. The wake-up now serves the feed whose `/disks/<rt_id>/`
directory fired, and nobody else: the `rt_id` was already parsed off the
path and thrown away. Each feed carries its own `published` watermark per
key, because the shared key cache advances with the FIRST wake-up and a
feed served by a later one would otherwise find "nothing new" and lose the
record; the in-process rt_mem lists of a follower keep being fed from that
shared cache, so they still see each record exactly once.
The watermark of every feed of the key is SEEDED at the batch start on
the batch's first wake-up, while that start is still known: a feed with
no watermark yet (one that just opened) cannot recover it from the
already-advanced cache, and unseeded it permanently lost the first
records after opening — they reached the sibling Live card and never
that one (on a brand-new key, whose batch starts at rowid 1, BOTH feeds
lost the first record; presence in `published` is what says "seeded",
since the legitimate seed there is 0). Covered by `tests/c/c_tranger`
(a keyed feed and a keyless feed over the same key, one publish each,
rt_mem on the master) and by `tests/c/timeranger2/test_rt_disk_multi_feed`
(the rt_disk fan-out itself: master + in-process follower, three feeds,
first-append-after-open and brand-new-key cases); verified against the
live backend by counting the WebSocket frames of a browser with both
cards open.
- **fix(timeranger2): a brand-new key made every rt_disk follower log an
`unlink() FAILED`.** On staging, `agregador_wz` logged one *"unlink()
FAILED, errno 2 (No such file or directory)"* per new key — an hourly
bucket meant an hourly ERROR in the log, and monitors alerting on nothing.
A follower is notified of new records by a hard link the master drops in
`<topic>/disks/<rt_id>/<key>/<file>.md2`; the follower unlinks it (consuming
the notification) and reads what is new. A brand-new key reaches it
**twice**, by construction: `fs_watcher`, on the `IN_CREATE` of the key
directory, adds the watch on it **first** and calls back **after**
(`fs_watcher.c`), so in the window between the two the master hard-links the
`.md2` inside — and the file then gets its OWN `IN_CREATE`
(`FS_FILE_CREATED_TYPE`) *and* is found by the directory scan
(`scan_disks_key_for_new_file`). Both paths lead to
`update_key_by_hard_link()`; whichever runs second finds the file already
gone. It is exactly the hazard the inotify(7) note quoted in that handler
warns about — the scan exists to cover it, but nothing deduplicated what
the watch would also report.
`ENOENT` is therefore not an error: it is the notification having been
consumed already. It is traced (`fs`) instead of logged as an error; any
other errno still is one. The read still runs — it is idempotent (it
publishes only the rows the cache does not have yet), so consuming twice
costs a re-read and nothing else, while skipping it would lose the records
if the other path had unlinked and then failed.
- **feat(c_tranger): `list-keys` filters, sorts and pages the keys IN THE
SERVER.** It answered every key of the topic, always: a client that wanted
the keys of one device was handed a hundred thousand of them and filtered
what it had been given — and a browser that only shows 15 at a time was
transferring, holding and sorting the whole index on its main thread.
New parameters: `rkey` (a PCRE2 regex the key must match), `order`
(`key`|`records`) + `desc`, and `from`/`limit`.
Backwards compatible by shape: with no `limit` the answer is the plain
list it has always been, so every existing client keeps working. Asking
for a PAGE gets the same envelope `get-page` uses —
`{total_rows, pages, data}` — so a client pages KEYS exactly as it pages
records, and `total_rows` counts the MATCHING set, not the page. Sorting
has to happen here for the same reason: a client holding 15 of 100.000
keys cannot order what it was not given.
PCRE2 and not POSIX `regcomp()`: every yuno already links `libpcre2-8`
(`tools/cmake/project.cmake`) and gobj-c already speaks it
(`json_replace_vars.c`), and this is the one place in the read path that
runs the same pattern against up to a hundred thousand subjects — the case
its JIT exists for. The pattern is compiled and JIT-compiled once, outside
the loop.
Cost is paid where it belongs: the record COUNT of a key is a cache lookup,
so it is taken for the whole matching set (that is what lets `order=records`
sort by it and `total_rows` be exact), while the key's TIME SPAN copies the
cache totals into a fresh dict and is therefore built only for the keys
that actually travel. A bad `rkey` and an unknown `order` are refused with
an error, never answered as "nothing matched".
- **feat(timeranger2): `rkey` governs a keyless list — the disk load AND the
realtime feed.** A list opened with no `key` is the whole topic; `rkey`
narrows it to the keys matching a regex. It was documented and half true:
the one-shot read (`open-list return_data=1`) filtered the keys, but a LIVE
list **refused** `rkey` outright (*"rkey is only supported with
return_data=1"*) — and rightly so, because `tranger2_open_list()` visited
every key on load and the realtime feeds (`rt_mem` / `rt_disk`) knew
nothing about it. A list filtered on load and unfiltered on append is worse
than no filter at all: the caller cannot tell it is being lied to.
`rkey` is now honoured in both halves, so a live list can finally be opened
on a SUBSET of a topic's keys. The three dispatch sites that decided who
gets an appended record (mem feed, disk feed, and the non-master path) went
through one `list_wants_key()`, so the two halves cannot drift apart again.
The compiled pattern lives IN the list (a pointer in its json, like its
`load_record_callback`) and dies with it: it is consulted once per appended
record, so compiling it per record would cost more than the load it saves.
PCRE2 + JIT, and `c_tranger`'s one-shot path was moved off POSIX
`regcomp()` onto it as well — one flavour of regex per parameter, or the
same `rkey` would mean two different things depending on which half of
`open-list` ran it.
A malformed `rkey` REFUSES the list (and the feed), rather than degrading
to "every key": that would push the caller exactly the records it asked not
to receive. The refusal is logged as a **warning**, not an error with a
stack trace: an `rkey` arrives from a remote peer (a SPA opening a live
list), so a bad pattern is peer input, not a broken internal invariant —
the decoder-severity policy. `gobj_log_set_last_message()` carries the
regex's compile error (pattern + offset) into the refusal the caller gets
back, so the SPA can show WHY the pattern was rejected.
(An earlier note in this section claimed `open-list`'s `rkey` was *silently
ignored*. That was wrong, and worth correcting: it was honoured in the
one-shot read and explicitly refused in the live one. The dead code is
timeranger2's commented-out `find_keys_in_disk`.)
- **fix(c_tranger): use-after-free closing a handle whose topic was closed
(SIGSEGV on shutdown).** A topic OWNS the iterators / rt_mem / rt_disk
handles opened on it: `tranger2_close_topic()` closes them all
(`tranger2_close_all_lists`) and frees the topic. C_TRANGER registered
what a remote client opened as a RAW POINTER, so once anything closed a
topic — the app closing a treedb at shutdown, a `delete-topic` at runtime
— its registry pointed at freed memory, and the next close dereferenced
it. In production this crashed a yuno with SIGSEGV on EVERY shutdown that
had a Rows card open: close-treedb frees the topics, then C_TRANGER's
`mt_destroy` walked its registry calling `tranger2_close_iterator()` on
iterators that no longer existed (jansson reading a dead hashtable). The
daemon then relaunched the yuno, so its binary was never idle and
`update-binary` failed with `copyfile() FAILED` — an un-updatable yuno was
the visible symptom. The registry now stores the topic name beside the
pointer and every use goes through a guard (new
`tranger2_topic_is_open()`): topic gone ⇒ the handle was freed with it ⇒
drop the entry, never touch it. Regression test: open an iterator + a feed,
close the topic under them, then command and tear down — it SIGSEGVs
without the guard.
- **fix(timeranger2): the cache totals of a key stored a timestamp with the
metadata flags baked into it.** In an on-disk `md2_record_t` the 16 high
bits of `__t__` carry the `user_flag` and those of `__tm__` the
`system_flag`. The disk-load path masks them off; the APPEND path did not,
and fed the raw record to `update_cache_cell()` — so on the master, a key's
cached `fr_t`/`to_t`/`fr_tm`/`to_tm` were poisoned by the flag bits (a `tm`
of 946684800 was cached as 17593132729216). This was not cosmetic:
`get_segments()` chooses which `.md2` files to read by comparing against
those very ranges, so **file selection by time was wrong for any record
carrying a non-zero user_flag or system_flag** (treedb tagged records,
queue pending flags). `update_cache_cell()` now reads the times through the
`get_time_t()` / `get_time_tm()` accessors. `test_topic_pkey_integer`'s foto
had the defect baked in — it held BOTH forms of the same value (polluted in
the in-memory cache, clean after a reload) and has been repointed.
- **feat(timeranger2): a filtered paging iterator now honors its
`match_cond` per RECORD (row index).** `get_segments()` can only reason
about whole files, and `tranger2_iterator_get_page()` ignored the
iterator's conditions entirely (it built its own `match_cond` with just
`from_rowid`/`to_rowid` and reported the FULL key row count), so the
time / rowid / user_flag conditions of `open-iterator` were only applied at
file granularity: a 4-record key filtered down to 2 still reported 4 and
its page returned all 4. An iterator that filters now builds its row
**index** when it opens (`tranger2_match_metadata` over each record's
32-byte metadata, no content read): `tranger2_iterator_size()`, `pages` and
the pages themselves count only matching records, and `get_page`'s
`from_rowid` is a position among THOSE rows. An unfiltered iterator builds
no index — its open stays O(1) on the key size and its positions remain the
global rowids. A dead row (`tranger2_delete_instance`) never enters the
index, so a filtered iterator's count no longer over-reports it.
- **feat(c_tranger): `list-keys` reports each key's time span, and `topics
expanded=1` its topic descs.** `list-keys` now returns `fr_t`/`to_t` and
`fr_tm`/`to_tm` alongside `records`, so a client can bound a time picker to
what the key actually holds without reading a record;
`topics expanded=1` returns a desc per topic (`topic_name`, `system_flag`,
`pkey`, `tkey`) — `system_flag` is the only thing that says whether the
topic's `t`/`tm` are seconds or milliseconds. Both are additive: the
default `topics` answer (an array of names) is unchanged. New public
`tranger2_topic_key_range()`.
- **fix(timeranger2): `get-page` reported "0 pages" for an out-of-range page.**
`pages` is a property of the KEY and the requested page size, not of the
particular answer: a page past the end has no data, but the key still has
its pages. Returning 0 told the client "there is nothing at all" and
collapsed its pager to a single page — a remote-paginated table then
refused to move ("Next page would be greater than maximum page of 1") and
its Last button landed on an empty table. Both empty branches
(out-of-range `from_rowid`, no first segment) now report the real page
count.
- **fix(c_tranger): realtime feeds leaked, and every leaked feed duplicated
records for EVERY subscriber.** `publish_rt_callback()` runs once per OPEN
FEED, and each run published `EV_TRANGER_RECORD_ADDED` to the whole
channel. Feeds were only ever released by an explicit `close-rt` (or at
`mt_stop`), so a remote client that died without one — browser reload,
closed tab, dropped websocket — left its feed alive forever. Every
surviving feed then re-published each append, so N leaked feeds meant N
copies of the same record delivered to all subscribers, not only to the
session that leaked them (observed in gui_treedb: the same `rowid`
repeated ~20 times in a Live card).
Two fixes: (1) the payload now carries the `rt_id` of the feed that
produced it, so a record can be routed to the subscriber that OPENED that
feed instead of broadcast (a foreign feed can no longer duplicate anyone's
rows); (2) `mt_subscription_deleted` closes the feeds **and iterators**
opened by a subscriber once its last subscription is gone — the owner is
stamped into the rt/iterator as `src_gobj` at `open-rt` / `open-iterator`.
Consumers of the event get one new field (`rt_id`) and are otherwise
unaffected; a consumer that does not filter on it keeps working.
An iterator opened by a session that never subscribes (Rows-only browsing)
is still only reclaimed at `mt_stop` — see `TODO.md`.
- **feat(c_tranger): `open-iterator` accepts metadata match conditions.**
Beyond `key` + `backward`, the command now forwards the record-metadata
conditions honored by `tranger2_match_metadata` into the iterator's
match_cond, so they pre-filter the page index and the reported
`total_rows` / pagination reflect the filtered set: `from_t`/`to_t`
(t, epoch seconds), `from_tm`/`to_tm` (tm, epoch ms),
`from_rowid`/`to_rowid` (1-based; negative = from end) and the user_flag
conditions (`user_flag`, `not_user_flag`, `user_flag_mask_set`,
`user_flag_mask_notset`). Each is optional — `0`/empty means unset. Only
keys actually supplied are added to match_cond. Record-FIELD filters
(e.g. `voltage > 200`) are deliberately NOT plumbed here: they are not
indexable at the metadata level, so they stay client-side in the SPA.
`open-rt` is unchanged — the realtime feed filters only by `key` (its
stored match_cond is not consulted on append), so no non-functional
params were added there. Enables gui_treedb's Rows-card request options
(yunos-js).
- **fix(timeranger2): fail loud on inotify `IN_Q_OVERFLOW`.** Under a burst
the kernel drops inotify events and emits a single `IN_Q_OVERFLOW`
(`wd == -1`); until now `fs_watcher` ignored it, so a dropped
`FS_FILE_CREATED`/`FS_SUBDIR_*` left an rt-disk follower silently out of
sync. `fs_watcher` now logs it `critical` with `LOG_OPT_ABORT`: the yuno
aborts and ydaemon relaunches it, and the clean reload re-establishes
every feed correctly — the proven recovery path, chosen over a
hard-to-test in-place resync. The deb/rpm packagers also raise the
inotify sysctl provisioning in `99-yuneta-core.conf`:
`fs.inotify.max_user_instances` 1024 → 4096,
`fs.inotify.max_user_watches` → 524288, and
`fs.inotify.max_queued_events` → 65536 (a defensive cushion above the
16384 kernel default), so the abort/relaunch stays rare.
- **feat(c_yuno): `info-inotify` command** — reports the system inotify
limits (`/proc/sys/fs/inotify/*`) and this yuno's own usage (instances +
watches, via `/proc/self/fd` + `fdinfo`), alongside `info-cpus` /
`info-ifs` / `info-os`. The `/proc/self/fd` probing lives in
`helpers.c` as `get_inotify_self_usage()`.
- **feat(c_tranger): realtime feed commands `open-rt` / `close-rt`, and
`EV_TRANGER_RECORD_ADDED` is now `EVF_PUBLIC_EVENT`.** `open-rt {rt_id,
topic_name, key}` opens a realtime-only feed on a topic key — NO history
load, no data retention — that streams each NEW append to subscribers as
`EV_TRANGER_RECORD_ADDED {topic_name, key, rowid, record}` (the record
carries a `__md_tranger__` with the same field names `get-page` emits, so
live and paged rows render identically). rt-by-memory on the master
(fires on `tranger2_append_record`), rt-by-disk on a reader (inotify).
`close-rt {rt_id}` closes it (via `tranger2_close_list`, which dispatches
by list_type). Feeds live in a per-gclass registry closed at `mt_destroy`
if a client never sends `close-rt`. Making the event `EVF_PUBLIC_EVENT`
is what lets a remote subscriber (a browser SPA) subscribe over the
ievent gate (`c_ievent_srv` requires the flag). The regression test
(`tests/c/c_tranger/`) subscribes a probe and asserts a fresh append
publishes exactly once, a cross-key append does not reach the feed, no
publish after `close-rt`, and a left-open feed is leak-free at destroy.
Enables gui_treedb's Live records card (yunos-js).
- **fix(timeranger2): three defects found while documenting the public API.**
(1) `tranger2_topic_key_size()` with an empty `key` passed `gobj` where a
`json_t *tranger` is expected, so the whole-topic fallback silently
returned 0; now it returns the real total. (2) `tranger2_close_all_lists()`
was declared `(…, rt_id, creator)` in the header but the implementation and
both callers use `(…, creator, rt_id)`; the header (and the `kernel/c`
README signature) are corrected to match — a caller trusting the old header
filtered by the wrong field (prototype names only, no ABI change).
(3) `tranger2_get_iterator_by_id()` / `_get_rt_mem_by_id()` /
`_get_rt_disk_by_id()` tested `empty_string(creator) && empty_string(creator)`
(the second operand was meant to be the stored `creator_`); an empty
query-creator now matches only creatorless entries (both-empty) instead of
any creator. The three functions also normalize a NULL query-creator to
`""` (like `tranger2_open_rt_mem()` does), so a NULL `creator` can no
longer reach `strcmp(NULL, …)` when the entry has a creator.
Also: the whole `timeranger2.h` public API was re-documented
(ownership, return semantics, NULL/error paths, master-only, disk-vs-memory)
and the doc.yuneta.io timeranger2 API page synced.
- **feat(c_tranger): restore the record-read commands `open-list`,
`get-list-data` and `close-list`** — v7-port stubs ("Pending to review")
rewired to the current timeranger2 iterator/list API, keeping the v6
contract: `open-list` accepts the full match_cond parameter set
(`key`, `rkey`, `from/to_rowid` incl. negative-from-end, `from/to_t(m)`,
`fields`, `only_md`, `backward`, user-flag masks). With `return_data=1`
it is a ONE-SHOT snapshot read — loads the matching records per key with
short-lived iterators and auto-closes (a remote client, e.g. a SPA, may
never send `close-list`); without it the list stays open (registry in
the gclass, closed at destroy), collecting history + realtime appends in
its `data` (readable via `get-list-data`) and publishing appends as
`EV_TRANGER_RECORD_ADDED`. `only_md` records are synthesized md-only
dicts (the current loader hands a NULL content record). Bool/int params
are read with `KW_WILD_NUMBER` (command-yuno/-agent forward strings).
`add-record` stays stubbed (write path). Consumed by gui_treedb's new
tranger records browser (yunos-js repo).
- **feat(c_tranger): cursor-pagination command surface — `list-keys`,
`open-iterator`, `get-page`, `close-iterator`.** Exposes the timeranger2
per-key iterator primitives (`tranger2_open_iterator` +
`tranger2_iterator_get_page`) as commands so a remote client can page
through a key's records with a real cursor instead of re-reading a
growing snapshot. `open-iterator` builds the key's row index only (no
upfront record load, no realtime feed — `load_record_callback` NULL) and
returns `{iterator_id, total_rows}`; `get-page from_rowid=<1-based>
limit=<n> [backward=1]` returns `{total_rows, pages, data}`, reading
records lazily; `close-iterator` closes and deregisters. Open iterators
live in a per-gclass registry (mirror of the open-lists one) and are
closed at `mt_destroy` if a client never sends `close-iterator` (a leaked
iterator would retain file handles). `list-keys` returns a topic's keys
with their record counts (`[{key, records}]`) — the input for a
two-level keys→records browser. All check `read`/`list` authz; ints/bools
read with `KW_WILD_NUMBER`; the parameter is named `iterator_id` (not
`id`) to dodge the command-yuno `id=` collision. New regression test
`tests/c/c_tranger/` drives the four commands through `gobj_command`
(forward paging, out-of-range, dup-open, close-then-404, missing key /
topic, and a deliberately-left-open iterator proving the destroy-time
cleanup is leak-free). To be consumed by gui_treedb's two-level tranger
records browser (yunos-js repo, Phase 2).7.7.2¶
C / SDK patch release — agent TTY console lifecycle: open-console re-attach
(the gui_agent Terminal shell now survives a browser refresh instead of leaking
a PTY per reload), max_consoles off-by-one, and clean shutdown with consoles
open. Pairs with the gui_agent stable per-tab console name + screen restore
(tracked in the yunos-js repo CHANGELOG); @yuneta/gobj-js is 7.7.2 on npm.
- **fix(agent/agent22): `open-console` re-attach — a browser refresh of the
gui_agent Terminal no longer accumulates PTYs until `max_consoles`.**
Re-opening an EXISTING console from the same channel (the
controlcenter↔agent route, which stays up across browser refreshes, so
the agent never sees the client disconnect) answered `-1 "Console
already open"` and each refresh had to fork a new console. Now it is a
re-attach: the stored route `__md_iev__` is refreshed so the tty stream
routes to the NEW requester (not the dead one), and `EV_TTY_OPEN` is
replayed to the requester with the same kw shape C_PTY publishes at
start (name/process/uuid/cwd/rows/cols) so the client leaves
"Connecting…" — the shell session survives the refresh. Also fixed the
`max_consoles` off-by-one (`>` → `>=`: the limit admitted max+1
consoles). Pairs with the gui_agent stable per-tab console name
(yunos-js); both `c_agent.c` and `c_agent22.c`.
- **fix(agent/agent22): close live consoles on shutdown.** With a console
open, an orderly exit (Ctrl-C) reached `gobj_end` with the C_PTY still
started — "Destroying a RUNNING gobj" + "hgobj destroying" + a running
`YEV_READ_TYPE` event destroyed hot. `mt_stop` now deletes every entry
of `list_consoles` (the volatil C_PTY services stop before the tree is
destroyed). Note the key-aliasing trap: `delete_console` KW_EXTRACTs the
console from `list_consoles`, freeing the dict key mid-call, so the
loop passes a copy of the name.7.7.1¶
C / SDK patch release — control-plane / auth-BFF hardening and correctness
fixes, plus an mbedTLS backend bump. JavaScript framework changes are tracked in
their own repositories (@yuneta/gobj-js, @yuneta/gobj-ui CHANGELOGs); the
gui_agent / gui_treedb yunos record their own UI changes under
yunos/js/*/README.md.
- **chore(ext-libs): bump mbedTLS 4.1.0 → 4.2.0 (v1.21).** The mbedTLS
TLS backend (runtime-selectable, statically linked into every yuno built
with `CONFIG_HAVE_MBEDTLS`) moves to the `v4.2.0` upstream tag;
`repos2clone.sh` re-pins `TAG_MBEDTLS` and `configure-libs.sh` bumps the
ext-libs `VERSION` 1.20 → 1.21. Rebuild ext-libs (`extrae.sh` +
`configure-libs.sh`) and relink the affected yunos.
- **harden(c_auth_bff):** cookie/token paths sized and fail-closed. One
`BFF_TOKEN_MAX` (8 KB) now covers the stored refresh_token, the cookie
extraction buffers and the Set-Cookie build (a Keycloak access_token with
many roles exceeds the old 4 KB extraction buffer — `/auth/token` would
forward it silently clipped and remote backends rejected the signature with
no evidence). `make_set_cookie` refuses to emit a truncated cookie (logs +
returns NULL; the login/refresh answer becomes `500 token_too_large`),
`extract_cookie` treats a value that doesn't fit as missing (clean 401
instead of a corrupted token), and `/auth/token` checks its `json_pack`.
- **fix(c_authz):** `create-user` / `update-user` no longer report success
when the treedb write failed — the `EV_ADD_USER` result is propagated, so a
failed autolink (bad `roles^ROLE^users` string) or tranger write error
answers `Can't create/update user`. Also: "User not exist" → "User does not
exist", and `update-user` shares `pm_create_user` (the table was a verbatim
copy).
- **fix(yuno_agent):** a `write-tty` naming a console that no longer exists
now answers the requester with a synthetic `EV_TTY_CLOSE` (same path as
`ac_tty_close`), so a client whose original close was lost in a link flap
can close its Terminal tab instead of typing into the void. `multiple_dir`
logs on snprintf failure/truncation (a truncated tags path silently placed
the yuno under a wrong repos dir). `ac_stats_yuno_answer` gains the same
self-name short-circuit as its command twin (no more "Event NOT DEFINED"
noise when the agent itself is the stats requester).
- **fix(prot):** restore the `gobj_has_bottom_attr()` guard on the
peername/sockname log reads in `c_prot_tcp4h` / `c_ievent_srv` — the bare
`gobj_read_str_attr()` logged "Attribute NOT FOUND" + stack trace in the
race windows where the bottom chain is unset (dropped by 390b8c679, which
overlooked that internal log).
- **fix(glogger):** the `TRACE_GBUFFERS` pretty-print also accepts
`[`-rooted JSON arrays (they fell through to the hex dump).
- **chore(trace samples):** drop the `"monitor"` / `"event_monitor"` lines
from the commented trace blocks in `ycommand` and the yuno skeletons — those
global levels don't exist (uncommenting yielded "global trace level NOT
FOUND"); the stale `gobj.h` name list is now synced with
`s_global_trace_level`.
- **feat(c_auth_bff):** new opt-in `POST /auth/token` endpoint returns the
access_token to JavaScript so a single SPA can forward it in a `C_IEVENT_CLI`
identity_card to Yuneta backends on **other** hosts (multi-backend browsing,
e.g. `gui_treedb`). A deliberate, scoped SEC-06 relaxation, off by default and
double-guarded: `expose_access_token` attr (default `false`; when off the
endpoint is an invisible `404`) **and** fail-closed Origin pinning (the token
is emitted only when the request `Origin` exactly matches `allowed_origin`,
else `403 origin_not_allowed`; enabling the flag without pinning an origin
yields nothing). Every existing BFF (wattyzer, estadodelaire, hidraulia) keeps
full SEC-06 — they never enable the flag. The server side already accepts an
identity-card JWT with priority over the cookie (`c_ievent_srv.c`) and
validates it against the issuer JWKS (`c_authz.c`); each remote backend must
have that JWKS provisioned. Docs: `YUNO_AUTH.md` §2.2.
- **fix(c_ievent_srv):** `ac_mt_command`'s "Service not found" error path
answered with `EV_MT_STATS_ANSWER` (copy-paste from the stats handler)
instead of `EV_MT_COMMAND_ANSWER`, so a remote command to an unknown service
got its error back under the wrong answer event.
- **fix(c_tcp):** a running client dropped via `EV_DROP` could stall in
`ST_STOPPED` (the reconnect `EV_TIMEOUT` was then ignored) and never
reconnect; `try_to_stop_yevents()` now finalizes to `ST_STOPPED` only when the
gobj is actually stopping. Generic to every C_TCP client dropped while alive.
- **fix(c_pty):** `EV_TTY_CLOSE`'s `json_pack` had a stray extra `}` → NULL kw
(no `name`/`uuid`/`slave_name`), so consumers keying off the console name
(e.g. gui_agent's Terminal) never saw a usable close.
- **fix(controlcenter):** `write-tty` now matches a node by UUID **or** hostname
(like `command-agent`) and, on a no-match, logs instead of `EV_DROP`-ping the
requester's shared control socket.
- **fix(controlcenter):** the `write-tty` no-match log is now a **warning**
(`MSGSET_PROTOCOL`), matching the agent side's demotion of the same benign
client race — an agent briefly disconnecting while a Terminal tab is focused
no longer emits one ERROR per keystroke into logcenter.
- **fix(glogger):** `gobj_trace_json`'s `TRACE_GBUFFERS` pretty-print path
parsed the gbuffer with `string2json(…, verbose=TRUE)`: a gbuffer starting
with `{` that is not one complete JSON value (partially-consumed buffer,
back-to-back messages, binary starting `0x7B`) injected a spurious
`gobj_log_error` + stack trace from a pure trace path. Now non-verbose — the
existing hex-dump fallback covers the parse failure.
- **fix(yuno_agent):** `write-tty` no longer drops the whole control link on a
benign per-write error ("console not found" → warning); it logs and drops just
that message.7.7.0¶
C / SDK release — capability marker. No C/SDK source change since 7.6.8;
this minor advertises the agent version boundary from which a controlcenter
can drive the command/stats control plane of a node’s managed yunos. JavaScript
framework packages (@yuneta/gobj-js, @yuneta/gobj-ui) are unchanged this
cycle; the gui_agent yuno records its own UI changes in
yunos/js/gui_agent/README.md.
- **Controlcenter-driven `command-yuno` / `stats-yuno` of managed yunos —
minimum agent version is now `7.7.0`.** The agent-side plumbing that
returns a controlcenter-cascaded answer to the original requester
(SPA → controlcenter → agent → yuno, back again) shipped in 7.6.8: the
agent hands the answer to its outbound `controlcenter` `C_IEVENT_CLI`,
which serializes the inner inter-event via its `EV_SEND_IEV` action
exactly as `C_CHANNEL` does server-side
(`kernel/c/root-linux/src/c_ievent_cli.c`, `yunos/c/yuno_agent/src/c_agent.c`).
7.7.0 **promotes this to an advertised compatibility boundary**: the agent
reports `YUNETA_VERSION` as its `__version__`
(`yunos/c/yuno_agent/src/main.c`: `APP_VERSION = YUNETA_VERSION`), so a
controlcenter can gate on `agent >= 7.7.0` before routing commands/stats
down to a node. An agent below 7.7.0 must not be assumed to answer these
cascaded control-plane calls.
- **chore(gui_agent): controlcenter web console rollup.** The console that
consumes this capability matured into the `yunos-js` submodule: a live
**Stats** panel per selected node, console command **history**,
command **shortkeys** (ycli parity) managed from Preferences, a
**copy-response** button, and TreeDB removed from the console. Tracked in
`yunos/js/gui_agent/README.md`.
- **chore(repo): `yunos/js` extracted into the `yunos-js` submodule**
(`github.com/artgins/yunos-js`, tracks `main`). It sits at its original
path, so the yunos' local `@yuneta/gobj-js` / `@yuneta/gobj-ui` `file:`
deps resolve unchanged. Edit in `yunos/js`, commit on `main` in the
standalone repo, then bump the submodule pointer here — the same flow as
`gobj-js` / `gobj-ui`.
- **chore(cli): yunetas CLI `0.12.0`** — pre-build ext-libs version guard,
published to PyPI.7.6.8¶
C / SDK release. JavaScript framework changes are tracked in their own
repositories (@yuneta/gobj-js, @yuneta/gobj-ui CHANGELOGs); the gui_agent
yuno records its own UI changes in yunos/js/gui_agent/README.md.
- **fix(agent): controlcenter-cascaded `command-yuno` / `stats-yuno`
answers now return to the requester.** A command routed down from a
controlcenter (SPA → controlcenter → agent → target yuno) reached the
yuno and produced its answer, but the agent dropped it on the way back
(`ac_command_yuno_answer` / `ac_stats_yuno_answer` logged
`child not found`), so the caller only saw the synchronous dispatch ack.
The reverse hop resolved the requester **only** among `__input_side__`
children; a command that arrived over the agent's outbound
`controlcenter` `C_IEVENT_CLI` link has its requester on the client side,
not there. The answer is now handed to that `C_IEVENT_CLI`, which gained
an `EV_SEND_IEV` action (`kernel/c/root-linux/src/c_ievent_cli.c`) that
unwraps and serializes the inner inter-event exactly as `C_CHANNEL` does
on the server side — one uniform path for both server- and
client-initiated links, covering command and stats answers alike
(`yunos/c/yuno_agent/src/c_agent.c`).
- **feat(authz): add `update-user`, split from `create-user`.**
`create-user` already rejected an existing user, so there was no way to
modify one: `create-user` now cleanly creates (reject if exists) and the
new `update-user` modifies (reject if missing); both funnel through
`EV_ADD_USER`. Two latent defects fixed while reviewing the new path
(`kernel/c/root-linux/src/c_authz.c`): a password-less update no longer
wipes stored credentials or silently re-enables a disabled account
(`disabled` is written only when supplied or the user is new), and the
existence-check node is no longer leaked. The `ROLE` format
(`roles^ROLE^users`) is now spelled out in the command help.
- **chore(ext-libs): bump jansson 2.15.0 → 2.15.1 (v1.20).** Patch release,
no API/ABI change. Caps recursion depth in `json_dump` / `json_equal` /
`json_deep_copy` (anti-DoS hardening on functions every `kw`
serialize/compare/copy runs through), rejects a negative string length in
the `json_pack` `s#` / `+#` formats, and adds the offending key/index to
`json_unpack` type-mismatch errors. `libjansson.a` links statically into
every yuno, so all must be rebuilt + relinked.
- **chore(ext-libs): bump liburing 2.14 → 2.15 (v1.19)**. Pin-only — no
API/ABI change and no removed/renamed symbols, so no yuneta consumer,
header or CMakeLists change rides along. Two of the 2.15 bug fixes land
on the exact APIs the event loop uses (`kernel/c/yev_loop/src/yev_loop.c`):
`io_uring_peek_cqe()` drops out-of-line round trips and a redundant
acquire ordering (loop drain path), and a stale-CQE-pointer fix on
wait-with-timeout errors (`io_uring_wait_cqe_timeout` in the wait path).
The new 2.15 helpers (`register_bpf_filter` / `register_query` /
`register_zcrx_ctrl`) are additive and unused here. `liburing.a` is
linked statically into every yuno, so yunos that link `yev_loop` must be
rebuilt + relinked (`extrae.sh` + `configure-libs.sh`) to pick it up.
- **chore(trace): richer unknown-event drop diagnostics + JSON-gbuffer
pretty-print.** On the empty-`iev_event` drop path both `C_IEVENT_SRV`
and `C_IEVENT_CLI` now dump the offending kw via `gobj_trace_json` — the
informative error is already logged upstream by
`gclass_find_public_event`, so the CLI's duplicate `gobj_log_error` was
removed. `gobj_trace_json` also pretty-prints a JSON gbuffer under
`TRACE_GBUFFERS` instead of a hex dump (and decrefs the parsed value, so
the trace path no longer leaks). (`kernel/c/gobj-c/src/glogger.c`,
`kernel/c/root-linux/src/c_ievent_{cli,srv}.c`)7.6.7¶
- **fix(security): close three buffer/parse defects found in a source
audit.** (1) `release-packages.yml` interpolated `github.event.release.
tag_name` / the `workflow_dispatch` input straight into a `run:` shell
body — an actor with write access could inject commands via a crafted tag;
the values now flow through `env:` and are referenced as quoted shell
variables. (2) `c_auth_bff`'s `make_set_cookie()` reused `snprintf`'s
return value without clamping, so an oversized token value (truncation)
drove the later `buf+n` / `sizeof(buf)-n` offsets past the stack buffer —
an out-of-bounds write in the auth path; `n` is now clamped to
`[0, sizeof(buf)-1]`. (3) `c_agent`'s `multiple_dir()` advanced
`p += ln; bflen -= ln;` on the `snprintf` return without a truncation
check, so a domain component that overflowed the buffer sent `p` past the
end and `bflen` (int) negative — widening to a huge `size_t` in the next
`snprintf`; it now breaks on truncation.
- **harden: route every `strtok` through `strtok_r`; add split-helper
tests.** `get_cpus()` (`c_yuno`), `split2()` (`helpers`) and
`json_unflatten_dict()` (`kwid`) relied on `strtok`'s hidden static state
— not a live bug (single-threaded, self-contained parses) but a
reentrancy footgun, notably for the public `split2()` (16 call sites). All
now use `strtok_r` with a local saveptr; no behaviour or signature change.
Added `tests/c/helpers/` (split2 + a reentrancy regression asserting
split2 no longer clobbers a caller's in-progress `strtok` parse) and a
`json_unflatten_dict` case in `tests/c/kw`. The vendored `linenoise.c/.h`
reference snapshot was refreshed (it is non-compiled — the console uses
`c_editline`) and `modules/c/console/README.md` now documents that.
- **fix(js): bump `@yuneta/gobj-js` submodule to 7.6.7 — restore the
`EV_ON_CLOSE`-on-deliberate-stop contract in `c_ievent_cli`.** 7.6.6
nulled the WebSocket `.onclose` handler in `mt_stop()`, so a deliberate
stop never delivered the async `EV_ON_CLOSE` to the FSM and never
published it to subscribers — diverging from the C kernel (where
`mt_stop` stops the bottom transport and `ac_on_close` publishes
`EV_ON_CLOSE` when a session was open). Consumers that drive their logout
UI teardown from that event (estadodelaire) stopped hiding the app. 7.6.7
keeps `.onclose` wired and instead guards the handler with
`gobj_is_destroying()`, so only the stop+destroy-in-the-same-turn case
(gui_agent) bails out silently while a stop that keeps the gobj alive
still gets its `EV_ON_CLOSE`. wattyzer/gui_agent (explicit teardown) are
unaffected. gobj-js-only patch: `YUNETA_VERSION` stays 7.6.6.
- **refactor(decoders): protocol decode errors caused by a malformed packet
from the peer are now warnings, not errors.** Across the protocol gclasses
(`c_prot_tcp4h`, `c_websocket`, `ghttp_parser`, `c_prot_mqtt`,
`c_prot_mqtt2`) a malformed/unexpected frame from the peer logs as
`gobj_log_warning` with its assigned category (`MSGSET_PROTOCOL`, or
`MSGSET_MQTT` for mqtt) plus a length-capped dump of the offending frame
(`MAX_LOG_DUMP_SIZE`, 256). MQTT uses a single capped dump at the
`frame_completed()` dispatch chokepoint, and the peer-malformed warnings
no longer carry `LOG_OPT_TRACE_STACK`. This also covers `mqtt_read_string`'s
"malformed utf8" path (both gclasses), previously mis-tagged
`MSGSET_INTERNAL` with a stack trace: a peer sending an invalid-UTF8 string
field is peer-malformed, now a `MSGSET_MQTT` warning with a capped dump of
the offending bytes. Reserved for our own faults
(`gobj_log_error`): broken internal invariants (e.g. `c_ievent_srv`'s
"gbuffer NULL", where the gbuffer must arrive in the `kw`), allocation
failures, outgoing-encode paths, and unsupported/unknown protocol fields
("NOT IMPLEMENTED"/"NOT FOUND") that mark our own TODO/map-gaps. The point:
error logs should mean "our fault", so routine peer misbehaviour stops
polluting error counts. Test `c_mqtt/malformed` updated to expect the
rejection as a warning.
- **feat(agent): report each binary's on-disk file time in
`*list-binaries` / `*list-binaries-instances`.** Both commands now add
`time` (epoch seconds) and `time_str` (local timestamp) next to `size`,
computed live by `stat()`ing the stored `binary` path (no treedb schema
change; covers binaries installed before the field existed; never mutates
the in-memory node — the listed records are fresh `node_collapsed_view`
dicts). The motivation is `sync_binaries`: `size` alone calls a rebuild
that kept the byte count identical (a one-char log edit, or a relink
against a changed static lib) "up-to-date", so it was never offered for
`update-binary`. Added `add_binary_file_time()` in `c_agent.c`.
- **fix(sync-binaries): detect a same-version rebuild by file time, not just
size.** `classify()` now flags a `REBUILD` when the local file is newer
than the agent's installed slot even when `Δsize` is 0: it prefers the
numeric `time` (file mtime) the agent reports next to `size`, and falls
back to the embedded build `date` (`__DATE__ " " __TIME__`, in
`--print-role` / `*list-binaries`) for an older agent. The candidate table
gains a `note` column spelling out a date-triggered rebuild ("newer
build") so a 0-`Δsize` `REBUILD` doesn't read as a no-op. `tools/README.md`
updated. The same newer-than-slot check now also applies in the snap-pinned
(`INSTALLED`) branch.
- **refactor(ytls): clarify the rejected-handshake log line (both backends).**
The default-on `gobj_log_warning` in `do_handshake` dropped the misleading
parenthetical hint from its `msg` — OpenSSL's
`(check ssl_min_version for legacy peers)` and mbedTLS's
`(mbedTLS floors at TLS1.2; use OpenSSL backend for legacy peers)` both now
read just `TLS handshake rejected`. The hint implied every rejection was a
protocol-floor issue, but most are internet background noise
(HTTP-on-TLS-port, port-scan garbage, open-proxy `CONNECT`); only
`unsupported protocol` / `version too low` are actual legacy peers. The
OpenSSL `tls_version` field was renamed to `negotiated_version` (since
`SSL_get_version()` returns the **server** object's version — equal to the
peer's offer only when the ClientHello was parsed far enough, otherwise the
server default, so a plaintext-HTTP probe is logged as `TLSv1.3`), and the
same `negotiated_version` field was **added** to the mbedTLS line via
`mbedtls_ssl_get_version()` (which honestly returns `"unknown"` pre-
negotiation) so both backends log a symmetric field set. Log-text/field-name
only; no behaviour change.
- **fix(ytls): raise the `ssl_verify_depth` default from 1 to 2.** OpenSSL
counts the trust anchor in the chain depth, so the minimal verification
path against any public CA is leaf(0) → intermediate(1) → root(2). With
the old default of 1, a verifying TLS **client** (`ssl_verify_mode`
`required`/`optional` with a CA) rejected every normal modern chain as
`X509_V_ERR_CERT_CHAIN_TOO_LONG` at depth 2 — observed on `auth_bff`
connecting to a Let's Encrypt-fronted Keycloak (`certificate verify
failed` after "certificate chain too long"). 2 is the de-facto floor;
cross-signed / extra-intermediate chains still need an explicit higher
`ssl_verify_depth`. Only the computed default changed in `openssl.c`;
`ytls.h` and `guide_tls.md` updated. The mbed-TLS backend has no depth
knob and is unaffected.
- **refactor(ytls): drop the handshake "forensic transcript".** Both TLS
backends captured every inbound handshake byte into a 16 KB per-socket
buffer and dumped it (hex) on handshake failure. The dump was useless — in
TLS 1.3 everything after ServerHello is ciphertext, and the cleartext
records are better read with the `C_TCP` `traffic` trace or a pcap — while
every connection paid for the allocation and a per-chunk memcpy. Removed
`HANDSHAKE_TRANSCRIPT_MAX`, the `handshake_transcript` field,
`capture_handshake_bytes()` and all decref sites from `openssl.c` /
`mbedtls.c`. The default-on `gobj_log_warning` recording the rejection
reason (error, peername, sockname, SNI, negotiated_version) is kept. The two
`test_handshake_dump_{openssl,mbedtls}` tests were repurposed as
`test_handshake_reject_*` (a bogus HTTP-on-TLS-port handshake is rejected
cleanly: `error=-1`, no crash) — both backends pass.
- **fix(agent-sync): classify `sync-binaries`/`sync-configs` against every
installed slot, not just the active primary.** `*list-binaries` /
`*list-configs` report only the primary; with a snap active the primary
can be an OLD version while the freshly built one is already installed as a
non-primary slot, so every role was mis-classed `BUMP` and then fired a
doomed `install-binary` / `create-config` (the agent rejects with "Node
already exists"). The full set is now read from `*list-binaries-instances`
/ `*list-configs-instances` and used as the authoritative "is this version
already installed?" check: a version installed but not primary is the new
`INSTALLED` status (skipped, with a hint to promote via
`yunetas upgrade-yunos`); `BUMP` now means "not installed and newer than
the primary". `tools/README.md` tables updated.
- **fix(agent-sync): skip the kill/restart cycle when the rebuilt version
isn't the one running.** The REBUILD path (`update-binary`) stopped and
restarted the role if ANY instance was live, but `update-binary` overwrites
only the slot whose version equals the uploaded binary's, and
text-file-busy bites only when a LIVE process is mapped to that exact file.
`deploy_update_with_restart` now reads `role_version` per instance from
`*list-yunos` and only kills/restarts when an instance running the version
being written is live (unknown `role_version` → treated as on-target,
killed on the safe side).7.6.6¶
- **build(js): extract `@yuneta/gobj-js` to its own repository as a git
submodule (symmetric with gobj-ui).** gobj-js was the last in-tree JS
framework package; it now lives at `github.com/artgins/gobj-js` (public)
and is embedded as a git submodule at `kernel/js/gobj-js`, the same model
as gobj-ui and `utils/python/tui_yunetas`. The new repo is a clean
snapshot (history not preserved), single line on `main`, tag `7.6.5`
tracking `YUNETA_VERSION`. **Clone with `--recurse-submodules`** (or
`git submodule update --init`). The submodule sits at the original path,
so local `file:` consumers (`wattyzer`, in-repo `yunos/js/gui_treedb`)
resolve unchanged and npm consumers (`estadodelaire`, `hidraulia`) are
unaffected. New publish flow: bump `package.json` in lockstep with
`YUNETA_VERSION` and `npm publish` **in the standalone repo**, then bump
the submodule pointer here. `.gitmodules` uses the HTTPS url (so
`--recurse-submodules` works for every cloner) for all three submodules;
the gobj-ui submodule's stale internal name (`lib-yui`) was also aligned
to `gobj-ui` and its url switched SSH → HTTPS in the same pass.
- **chore(ext-libs): bump nginx 1.30.2 → 1.31.2 (v1.17)**. Fixes three
CVEs: CVE-2026-42530 (use-after-free in `ngx_http_v3_module`),
CVE-2026-42055 (buffer overflow in the HTTP/2 paths of
`ngx_http_proxy_module` / `ngx_http_grpc_module`) and CVE-2026-48142
(buffer overread in `ngx_http_charset_module`). Pin-only — nginx is a
separate dynamically-linked binary (see `configure-libs.sh` v1.10), so
no yuneta consumer / header / CMake change rides along. NOTE: `1.31.x`
is the nginx *mainline* branch (odd minor), not the `1.30.x` stable
line we were on — chosen because the fixes landed there. openresty
(`1.29.2.5`) is a separate binary and is **not** covered by this bump;
track upstream openresty for a release that picks up these patches.
Each deployed project must rebuild its own nginx copy.
- **chore(ext-libs): bump openresty 1.29.2.5 → 1.31.1.1 (v1.18)**. Advances
the openresty-bundled nginx core from `1.29.2` to `1.31.1` (released
2026-05-29). Pin-only — openresty is a separate dynamically-linked binary
(so its bundled OpenSSL 3.5.6 is irrelevant to our build). ⚠️ **CVE
status:** nginx 1.31.1 does NOT cover the three CVEs fixed in nginx
`1.31.2` (CVE-2026-42530 / 42055 / 48142, 2026-06-17); openresty 1.31.1.1
was tagged *before* nginx 1.31.2, so the openresty binary — the one that
actually fronts the SPAs — remains exposed until upstream ships a release
based on ≥ nginx 1.31.2. The standalone nginx binary IS patched (v1.17).
Each deployed project must rebuild its own openresty copy.
- **observability(prot): attribute protocol parse errors to the source IP
(`peername`).** Server-side protocol gclasses logged malformed-input
errors without the remote peer's address — `peername` is set on the bottom
`C_TCP` (`SDF_VOLATIL`) and the upper layers never copied it into their own
logs, so a bad-frame / bad-header event was not attributable to a device
or attacker without cross-referencing the `C_TCP` `Connected` line by
timestamp. The canonical read pattern (already in `c_websocket.c` /
`c_prot_mqtt2.c`) is now applied in the cold error branch of each
remote-data parse-error log: `c_prot_tcp4h.c` (head-too-long,
protocol-error disconnect, protocol timeouts) and `ghttp_parser.c` (the
invalid-UTF-8 header-value store error; the main "non-HTTP data received"
violation already carried it). `c_prot_http_sr.c` / `c_channel.c` only
emit registration / internal "no bottom" logs (not remote-attributable)
and are left untouched; outbound clients (`c_prot_http_cl.c`) and the
`c_prot_mqtt2.c` gap-fill are deferred. No FSM/schema/API change — logs
gain a `peername` field only. See TODO.md "source-IP attribution".
- **security(glogger): escape invalid UTF-8 in log fields (logcenter parse
DoS).** `_ul_str_escape()` copied every byte `0x7f-0xff` verbatim (via the
`json_exceptions[]` table) without validating UTF-8. A corrupted device
payload logged verbatim (e.g. an FS00802_4G sensor leaking modem AT
commands + raw bytes into its MQTT JSON) leaked lone invalid bytes
(`0x8a`, `0xc2`, ...) into the log record; the record was then no longer
valid UTF-8, so the logcenter's `gbuf2json()` rejected and dropped it
("unable to decode byte 0x8a"), losing the log. Fixed at the root so no
field from any emitter can produce an unparseable record: `json_exceptions[]`
and the non-thread-safe static `exmap` are gone; a strict
`utf8_valid_seq_len()` validator (rejects overlong encodings, surrogates,
`> U+10FFFF`, never reads past the NUL terminator) now drives the escaper,
which copies valid UTF-8 sequences verbatim (logs stay readable) and
escapes invalid/control bytes as `\u00XX`. Worst-case output size is
unchanged (<=6 bytes/char), so the caller's buffer sizing is untouched.
Regression test `tests/c/glogger_utf8` registers a capture log handler and
re-parses the emitted record with `anystring2json` (what the logcenter
does), proving an invalid-UTF-8 payload now yields a valid JSON/UTF-8
record while legitimate UTF-8 (`café`, `€`) is preserved verbatim.
- **security(mqtt): reject zero-length payload frames that crashed the
broker (NULL-gbuf remote DoS).** A control packet whose MQTT "remaining
length" was 0 left `frame_completed()` with a NULL payload gbuffer, which
it then handed to a handler that dereferenced it
(`gbuffer_leftbytes(NULL)`) → SIGSEGV. A single malformed packet from a
remote client crashed the whole broker process (observed in
`handle__subscribe`). Fixed in both protocol gclasses (`C_PROT_MQTT` and
`C_PROT_MQTT2`) with two layers: at header validation, `frame_length == 0`
is now rejected for every command carrying a mandatory payload, before a
NULL gbuf can reach a handler (MQTT5 still permits a zero-length
DISCONNECT/AUTH, whose handlers already tolerate it; PINGREQ/PINGRESP keep
their must-be-zero check); and the read primitives
(`mqtt_read_uint16/uint32/bytes/byte/varint`) now treat a NULL gbuf as a
malformed packet before touching `gbuffer_leftbytes`. Regression test
`tests/c/c_mqtt/test_mqtt_malformed` injects a malformed in-session
SUBSCRIBE (`0x82 0x00`) at the transport: it crashes with the exact
production backtrace without the fix and passes with it.
- **build(js): the JS UI library was extracted to its own repository and
renamed `@yuneta/lib-yui` → `@yuneta/gobj-ui`.** It now lives at
`github.com/artgins/gobj-ui.js` and is embedded as a git submodule at
`kernel/js/gobj-ui` (clone with `--recurse-submodules`), the same model as
`utils/python/tui_yunetas`. The repo carries two maintained lines, each
consumed a different way:
- **`main`/v2** (npm dist-tag `latest`, tag `2.0.0`+, `src/` layout) —
active development: the declarative shell
(`C_YUI_SHELL/NAV/PAGER/WIZARD`) on top of the legacy stack. **The
yunetas submodule now tracks this line**, and `wattyzer` consumes that
checkout locally via a `file:` dependency (importing
`@yuneta/gobj-ui/src/*` by package specifier).
- **`v1`** (npm dist-tag `legacy`, tag `1.0.1`, `src/` layout) — the
frozen legacy GClass GUI stack. Consumed from the **npm registry** as
`@yuneta/gobj-ui@^1.0.1` by `estadodelaire`, `hidraulia` and the
in-repo `yunos/js/gui_treedb` (NOT a local `file:` — the local
submodule is v2 now).
The old `lib-yui` name collided with Yahoo's YUI on npm; only the package
identity changed — internal naming (`C_YUI_*`, `c_yui_*`, `yui_*`, `yi-*`)
is unchanged. Both lines use the `src/` layout (v2 was restructured to
match v1). Published to npm as `@yuneta/gobj-ui` (`latest`=2.0.0,
`legacy`=1.0.1); the abandoned `@yuneta/lib-yui` was unpublished.
- **build(js): `@yuneta/gobj-js` is versioned to `YUNETA_VERSION` and
published to npm.** `kernel/js/gobj-js/package.json` now tracks the SDK
version (currently `7.6.5`); bump it in lockstep and `npm publish`.
`estadodelaire`/`hidraulia` consume it from the registry
(`@yuneta/gobj-js@^7.6.5`); `wattyzer` and `yunos/js/gui_treedb` keep a
local `file:` dependency on `kernel/js/gobj-js`.7.6.5¶
- **security(libjwt): re-review against upstream v3.4.0 and backport the
reachable hardenings.** v3.4.0 is a large feature release (full JWE, the
`crit` header, `jti` callbacks, PEM→JWK public API), almost none of which
touches Yuneta's compiled subset. After filtering to the JWS verify/parse
path, three items were backported into the vendored tree:
- **`18133e4` (L17): reject duplicate JSON members on the token parse**
(`jwt-verify.c`). The inbound header/payload is now parsed with
`JSON_REJECT_DUPLICATES` (RFC 8725 §2.4), so a peer that selects a
different occurrence of a duplicated claim/header cannot be made to
disagree with us.
- **`d180cc7`: enforce strict base64url on decode** (`jwt.c`). The
decoder accepted the standard-base64 `+`/`/` and silently truncated
on an embedded `=`; it now rejects anything outside `[A-Za-z0-9_-]`.
Reachable on every token segment and JWK member decode.
- **`fe8840a`: enforce the RFC 7515 §4.1.11 `crit` (Critical) header**
(`jwt-verify.c` + both checker entry points). The parser previously
ignored `crit`; since this copy understands no extension headers, any
token carrying a well-formed `crit` is now rejected (checker side
only — the builder side is not ported, C_AUTHZ does not sign).
The batch's only CVE-class bug (`5fada81`, mbedTLS RSA short-signature
OOB read) is not present here: the vendored mbedTLS backend is a v4.0/PSA
rewrite using the length-aware `mbedtls_pk_verify_ext`, immune by
construction. Regression coverage added to `test_jwt_alg_confusion`
(`crit` rejection on both entry points; positive controls still verify).
Full classification in `kernel/c/libjwt/README.md`.7.6.4¶
- **fix(tr_msg2db): stop `msg2db_open_db` logging spurious schema errors
when reopening with `jn_schema=NULL`.** The persistent reopen path (no
schema dict passed — the schema is loaded from
`<db>.msg2db_schema.json`) read the name and `schema_version` straight off
the NULL `jn_schema`, so every open via `msg2db_list` (and any non-master
reopen) emitted three red errors — `kw must be list or dict` /
`path NOT FOUND` for `id`, and the same for `schema_version` — before the
function then correctly loaded the schema from file. Now mirrors
`treedb_open_db`: the name comes from the passed `msg2db_name_` when
`jn_schema` is NULL (and is read with a non-`KW_REQUIRED` flag otherwise),
and `schema_version` is guarded with `jn_schema? kw_get_int(...) : 0`. The
resolved name and version are unchanged for the one non-NULL-schema caller
(`c_mqtt_broker`, whose passed name already equals `jn_schema["id"]`);
only the noise is gone.
- **refactor(msg2db_list): modernize the CLI to the `treedb_list` style.**
The tool still carried its V6-era flags — most visibly `--path` / `-a`
as the only way to point it at a store. It now takes the store as a
**positional `PATH` argument** (the `-a` flag is gone) and shares the
`treedb_list` ergonomics:
- `resolve_msg2db_path()` deduces the tranger root, `--database` and
`--topic` from `PATH` (a tranger root, a `<db>.msg2db_schema.json`,
or a topic directory), auto-discovering the single schema when
`--database` is omitted and listing the candidates when it is
ambiguous or missing.
- new presentation flags `--mode form|table` / `-m` and
`--fields` / `-f` (field selection implies table mode), rendering
columns from the topic's `cols` `fillspace` just like `treedb_list`.
- new `--dry-run` / `-n` prints the resolved path / database / topic
plus the ids and filter JSON, then exits without listing.
- `PATH` is normalized (trailing slashes stripped, `./` prefixed for
bare relative names); the per-topic header is highlighted and the
recursive walk keeps its per-database record count.
`--follow` is intentionally NOT ported: `tr_msg2db` exposes no
change-callback equivalent to `treedb_set_callback`. The mega-header
`<yunetas.h>` include was replaced by the specific kernel headers.
- **chore(ytls): de-duplicate and de-noise the rejected-handshake logs.** A
single rejected connection (e.g. a non-TLS/HTTP client hitting the TLS
port) emitted overlapping lines across the ytls and transport layers.
Now:
- "TLS handshake rejected" is INFO (was WARNING) in both backends — a
sub-floor/legacy/non-TLS peer is routine, not actionable; it was
inflating "Global Warnings".
- that default-on line is self-contained: it now carries
`peername`/`sockname`, handed to ytls by the transport through a new
optional `ytls_set_peer_name()` (per-`sskt`, both backends), so the
offending peer is identifiable even with `connections` trace off.
(ytls does NOT reinterpret `user_data` as a gobj — unit tests pass a
non-gobj `user_data`; callers that skip the setter just log `""`.)
- the transport's `ytls_on_handshake_done_callback` no longer logs the
FAILS case (it duplicated the ytls line); it keeps "TLS Handshake OK".
- `set_trace` no longer logs the per-connection `trace:0` disable (pure
noise on every accept); it logs only when enabling. Both backends.
- **feat(libjwt): trace the claims JSON on a failed-claims verification.**
`__verify_config_post` now calls `gobj_trace_json` with `jwt->claims` when
`__verify_claims` reports one or more failed claims, so the offending
token's `iss`/`aud`/`exp`/`nbf`/… values are visible at the point of
rejection. Emits unconditionally on the failure path (LOG_DEBUG). libjwt
now back-references `gobj_trace_json`: real yunos already pull `glogger.o`,
but the standalone libjwt unit test pulls nothing from it, so its link line
repeats `libyunetas-gobj.a` after `JWT_LIBS` to resolve the reference.
- **fix(utils): TLS client utilities could not connect over `wss://` /
`https://`.** The verify-by-default change made `build_ssl_ctx` refuse any
TLS client whose `crypto` config lacks server-certificate validation, but
the CLI utilities were never ported: they passed an empty `crypto` to
their sockets, so every remote TLS connection (e.g. `ycommand` against the
controlcenter, including its OIDC `task-authenticate` to the issuer) failed
with *"TLS client refused: no server-certificate validation"*. All
utilities now pass `crypto: {ssl_use_system_ca: true}` to their C_TCP (and
to `C_TASK_AUTHENTICATE` where present): `ycommand`, `ystats`, `ybatch`,
`ytests`, `ycli`, `mqtt_tui`, `emu_device`. `ycommand` additionally gains
`--ssl-use-system-ca` (default on), `--ssl-trusted-certificate` (private
CA) and `--ssl-allow-insecure-client` (MITM bypass) for non-public-CA
endpoints. Plain `ws://` is unaffected (C_TCP ignores `crypto` without TLS).
- **feat(ytls): log `ssl_server_name` in TLS diagnostics; drop dead
fields.** Every post-init `gobj_log_*` in the OpenSSL and mbedTLS backends
now carries `ssl_server_name`, so handshake/verify/read/write errors show
which SNI/server name the context was for. Also removed the never-used
`rx_bf[16*1024]` field from `sskt_t` in both backends (~16 KB per live TLS
connection) and the unused `error` field from the mbedTLS `sskt_t`.
- **fix(tui_yunetas 0.10.1): quieter `upgrade-yunos` output.** The two
`find-new-yunos` steps no longer dump ycommand's raw stdout — the preview
prints once (formatted) and `create=1` shows a one-line `Created N new
yuno row(s).` summary. The post-`sync-binaries` install-binary reminder
now leads with `yunetas upgrade-yunos` (raw ycommand sequence kept as the
manual equivalent).
- **feat(tui_yunetas 0.10.0): agent-aware deploy.** `sync-configs` without
`--host` now matches each registered project's `yunos/batches/<host>/`
directories against the realm_ids the local agent manages
(`*list-realms`) and syncs every match — a node running several realms
deploys all the relevant ones in one pass, since a batches dir is named
after its realm_id (the deploy FQDN). `--host` still targets one dir; an
unreachable agent falls back to the legacy single-hostname guess; new
`--url`/`-u`. New `upgrade-yunos` command bundles the version-bump
promotion flow: optional rollback snapshot (idempotent by name,
`pre-upgrade-<YYYYMMDD>`, `--no-snap`) -> `find-new-yunos` preview +
confirm (`--yes`) -> `find-new-yunos create=1` -> `deactivate-snap`
(restart_nodes: SIGKILL + treedb reload, newest release wins).
`--dry-run` prints the agent commands without running them.
- **fix(tools): a resumed deploy is now idempotent instead of failing.**
When a prior run installed the binaries / configs and registered the new
yuno rows but never promoted them (`deactivate-snap` not reached),
re-running the deploy hit the agent's "... already exists" answers.
`sync_binaries.py` / `sync_configs.py` now report such an
`install-binary` / `create-config` as `ALREADY PRESENT` (idempotent) and
count it as ok, not a red `FAILED`. The matching `upgrade-yunos`
fall-through (don't abort when `find-new-yunos create=1` only hits
already-existing rows) ships in the tui_yunetas CLI 0.11.1. A genuine
(non-idempotent) error still fails closed.7.6.3¶
- **feat(treedb): immutable (non-deletable) topics and records.** A record
can be marked immutable (md2 system_flag bit `sf_immutable_record`,
surfaced as `__md_treedb__`immutable`) and a topic non-deletable
(`system_topic` in `topic_var.json`) — the protection is METADATA, not a
data column, so it needs no user-schema change and no `topic_version`
bump. `treedb_delete_node` / `treedb_delete_instance` /
`treedb_delete_topic` refuse it and `force` does NOT override; the record
bit is inherited across updates and survives reload. New
`treedb_set_node_immutable()` and a `system_topic` param on
`treedb_create_topic`; the `__system__` treedb structural topics and per-
treedb `__snaps__`/`__graphs__` are marked system. `c_authz` `mt_start`
runs a master-only idempotent ensure-loop that stamps the Authz seed
(`root` role / `yuneta` user) immutable on every start — deployed stores
protected on next restart, no schema change, no wipe. Out of scope on
purpose: `delete-treedb` / whole-store wipe. Test
`tests/c/tr_treedb_immutable`; design in
`kernel/c/timeranger2/DESIGN-immutable-topics-records.md`; docs in
`YUNO_TREEDB.md` §3.10 + `YUNO_AUTH.md` §4.2.
- **fix(ytls): portable system-CA trust (`ssl_use_system_ca`) for static
binaries, both backends.** A fully-static binary doesn't inherit the host
OPENSSLDIR / `SSL_CERT_FILE`, so OpenSSL's `set_default_verify_paths()`
loaded an EMPTY store and a valid public cert failed to verify. New
`ytls_get_system_ca_bundle()` probes the well-known CA bundle FILES across
distros (Debian/Ubuntu, RHEL/Rocky/Alma/Fedora, SUSE, Alpine — the
hashed-dir CApath is not portable); OpenSSL loads it via
`load_verify_locations`, and mbedTLS (no system store of its own) now
parses it too instead of refusing the client. `C_PROT_HTTP_CL` gains a
`crypto` attr (default verify-by-default) forwarded to its bottom C_TCP,
and emailsender's `c_smtp_session` the same — so HTTPS polls (e.g. ESIOS)
and SMTPS verify out of the box. This unbroke auth_bff's IdP TLS and
stopped a ~1.2 MB/s "TLS handshake FAILS" log flood (-> ~0.6 KB/s).
- **feat(c_tcp): opt-in exponential reconnect backoff.** New
`timeout_between_connections_max`: when > `timeout_between_connections`,
the reconnect delay backs off from base up to the cap, resetting to base
once a connection is established (for a TLS client, only on a successful
handshake). A peer that keeps failing — e.g. an IdP whose cert won't
verify — no longer hammers at the base cadence. `auth_bff` uses it (100 ms
first retry, 30 s cap).
- **fix(logcenter): `search` / `tail` no longer crash on a truncated log,
and read fast.** `extrae_json` brace-counting `abort()`ed the whole daemon
when a truncated UDP log entry left a `{` with no `}` (it grew past the
max block) — taking logcenter down on a read-only command. Records are now
split on the `<PRIORITY>: ` line prefix `rotatory_write()` already writes
(robust to truncation and multi-line JSON; a `\0` on-disk terminator was
rejected as it would make the log binary for grep/less/vim). Reads use
64KB blocks + `memchr` instead of `fgetc()` per byte, and `tail` seeks to
the last window: on a 518MB log, tail 72s -> 2s, search 70s+/crash -> 2-7s.
- **refactor(tui_yunetas 0.9.1): project registry moved to
`~/.yuneta/projects.json`.** The external-project registry was written
inside the source tree (`$YUNETAS_BASE/.projects.json`, gitignored). But
which projects to build alongside the SDK is runtime/usage state, not a
property of any checkout — it does not belong in the tree at all. It now
lives in the user's home (`~/.yuneta/projects.json`), independent of
`YUNETAS_BASE`. A one-time soft migration moves an existing legacy file
on the next CLI run; no manual step. The `.gitignore` / `.hgignore`
entries for the old path are kept as a safety net while pipx CLIs on
other nodes still write the legacy location.
- **fix(treedb): refused `treedb_delete_instance()` no longer drops a
borrowed node ref.** The snapshot-tag guard's refusal path decref'd the
node even though callers (`mt_delete_node`, tests) pass the index's
borrowed pointer — a refused per-instance delete of a snap-tagged node
left the index slot one ref short (latent use-after-free / double-free).
The refusal path now leaves the node untouched; the ref is consumed only
on success, where `delete_secondary_node()` extracts it from the index,
same convention as `treedb_delete_node()`.
- **fix(performance): `perf_yev_ping_pong2` no longer reports a first-run
`tranger2_startup` error.** The startup phase expected NO logs, but on a
node where `~/tests_yuneta/` had never been created (e.g. a fresh VM)
`tranger2_startup` emits the one-time INFO "Creating
`__timeranger2__.json`" — flagged as unexpected by the strict
expected-results FIFO. Simply adding the log to the expected list would
break the opposite case (database already created by a previous test or
run). Fix follows the established pattern
(`test_tr_treedb_update_instance.c`): wipe the database with `rmrdir`
before `tranger2_startup` so the creation INFO is always emitted, and
expect it. Verified with back-to-back runs (fresh and leftover store).
- **chore(packages): drop `stress_*` lab binaries from the .deb/.rpm
payload.** The CI builds the whole tree and the packagers copied
`outputs/` wholesale, so the stress load-generators (`stress_auth_bff`,
`stress_listen`, …) shipped on every production node. They are now
stripped at staging time. `perf_*` benchmarks stay on purpose: fully
static, they are handy to measure a target machine right after install
(validated on the 7.6.2 Ubuntu VM).
- **fix(packages): pipx CLIs install for the operator, not for root.**
`install-yuneta-dev-deps.sh` (deb/rpm) runs as root, so `pipx install
kconfiglib yunetas` landed in `/root/.local/bin` — invisible to the
`yuneta` operator account (verified on a clean Ubuntu VM: `yunetas:
command not found` after a full install). The script now installs the
pipx apps for the `yuneta` user when it exists (falling back to
`$SUDO_USER`, then root), via `runuser -l` so pipx resolves the right
`$HOME`. The staged `profile.d/yuneta.sh` already has
`/home/yuneta/.local/bin` on PATH, so the CLIs work on next login with
no `pipx ensurepath` step.7.6.2¶
- **feat(tui_yunetas 0.9.0): external projects integrated into the
`yunetas` CLI.** New `register-project` / `unregister-project` /
`list-projects` commands keep a machine-local registry in
`$YUNETAS_BASE/.projects.json` (gitignored); `init` / `build` / `clean`
now also process each registered project's `yunos/` after the SDK
(select with positional project names, or `--sdk-only` to skip them).
New `sync-binaries` / `sync-configs` subcommands wrap
`tools/agent/sync_*.py`, forwarding arguments; `sync-configs` walks the
registered projects' `yunos/batches/<host>/` directories (`--host`
selector with hostname auto-match), closing the discovery gap that
forced a manual `cd` into each batches dir.
- **fix(env): `yunetas-env.sh` exported a stale artefacts layout.**
`YUNETAS_OUTPUTS` / `YUNETAS_YUNOS` pointed at the PARENT directory of
the repo (`$(dirname $YUNETAS_BASE)/outputs`), a layout nothing else
uses: `project.cmake`, the CLI and the `.deb`/`.rpm` payload all agree
on `$YUNETAS_BASE/outputs[_ext]`. Both variables now follow that rule,
a new `YUNETAS_OUTPUTS_EXT` is exported, and `deactivate_yunetas`
unsets all of them. The legacy `$HOME/yunetaprojects` branch was
dropped from the profile script staged by the `.deb`/`.rpm` packagers,
and the docs (`CLAUDE.md`, `installation.md`) were aligned.
- **feat(packages): one outputs/ path on every node.** The `.deb`/`.rpm`
used to stage the SDK payload directly under `/yuneta/development/`
(`outputs/`, `outputs_ext/`, `tools/`, `.config`), so runtime-only nodes
had a different `YUNETAS_BASE` (and outputs path) than source checkouts.
Both packagers now stage the same payload as a sparse SDK under
`/yuneta/development/yunetas/` — the SAME base path as a full source
tree — so `outputs/` is `/yuneta/development/yunetas/outputs` everywhere
and the two-branch layout conditional in the staged `profile.d/yuneta.sh`
collapses to a single unconditional block. The `yunetas` CLI handles
these runtime-only trees (no `YUNETA_VERSION`): `init <project>` /
`build <project>` work against the shipped headers, and a plain `init`
refuses to wipe the shipped `outputs/`. Legacy `/yuneta/development`
fallbacks remain in the resolution chains for nodes installed with older
packages; on upgrade dpkg/rpm relocate the payload automatically, but
already-configured project `build/` dirs cache the old paths — re-run
`cmake` (or `yunetas init <project>`) after upgrading.
- **fix(ycli,ycommand): assemble local-config paths via `build_path`.**
The `save_local_json/string/base64` helpers built
`$HOME/.yuneta/configs/<name>` with `snprintf` into a NAME_MAX buffer
while the sanitized name was also NAME_MAX, so GCC emitted
`-Wformat-truncation`. `build_path()` (the hard-rule path helper) does
the assembly and logs LOG_CRIT on real overflow instead of silently
truncating.
- Note: the 7.6.1 tag already shipped two undocumented renames — the
legacy `cli` gobj/service is now `ycli`, and its global config key
`Cli.shortkeys` is now `ycli.shortkeys`.7.6.1¶
- **fix(authz,root-linux): a root superuser reaches any service of the
node.** The 7.6.0 per-message `dst_service` gate (`is_service_authorized`
in `c_ievent_srv.c`) authorized only the channel's `authorized_services`
— the *keys* of `services_roles`. But the local trusted `yuneta` user
authenticated through the `yuneta_by_local_ip` shortcut in `c_authz.c`,
which hardcoded an EMPTY role set (`{"agent":[]}`, the old
`// TODO not need role?`), so its real `root` role (`service="*"`,
`realm_id="*"`) never reached the channel. Result: the local control
plane (`ycli` warming its command cache with `list-gobj-commands` to
`dst_service="__yuno__"`, `ycommand`, …) was REJECTED at `__yuno__` and
any sibling service — root could not reach the yuno root. Now the local
`yuneta` goes through the SAME `get_user_roles()` filter as any user (no
hardcode); `get_user_roles()` flags the channel `superuser` when the user
holds an effective `service="*"` role (computed from the wildcard, not the
literal role name), propagated in the authenticate response and stored as
the `is_superuser` channel attr. `is_service_authorized()` returns TRUE for
a superuser: any realm/service/permission by definition, so it is not a
cross-service escalation. Scoped roles (`developer`, `sysop`, …) stay
limited to their granted services — the 7.6.0 cross-service protection is
intact for them. What a command may DO is still governed by the
default-off per-command authz, orthogonal to this routing gate.
- **fix(root-linux): the ievent server never leaves a channel zombie when it
refuses a message.** `ac_on_message` rejected an unrouted `dst_service`
(unauthorized or not found) with a bare `return -1`, which both skipped
any answer AND left the socket read un-rearmed (`c_tcp` only re-arms on a
`0` return from the `EV_RX_DATA` publish chain): the channel stayed
connected but deaf and the peer waited forever. New `reject_unrouted_iev()`
never returns `-1` silently: `command` / `stats` (which have a natural
answer channel) get a negative `EV_MT_*_ANSWER` with the reason and the
read re-arms (`return 0`); `subscribe` / `unsubscribe` / `inject` (no
answer) `drop()` the channel for a clean disconnect. Applied to both the
unauthorized-service and the service-not-found paths.
- **build(cmake): link the kernel and external static libraries by full
path so consumers auto-relink.** `tools/cmake/project.cmake` listed the
`.a` files as bare names resolved via `link_directories()` `-L`, which
CMake treats as plain `-l` flags with NO file dependency: after editing a
kernel source and rebuilding its `.a`, dependent yunos were NOT relinked
(`make` reported `Built target` with the stale binary; the workaround was
to delete the binary first). Each archive is now given by full path
(`${LIB_DEST_DIR}/...` for yuneta's own, `${EXT_LIB_DIR}/...` for
`outputs_ext/lib` third-party), so CMake tracks it as a link dependency
and `make` / `yunetas build` relinks automatically when a lib changes.
System libs (`pthread`, `dl`) stay bare. Verified: a full clean rebuild is
green and touching `libyunetas-core-linux.a` relinks the agent with a
plain `make`.7.6.0¶
- **security(root-linux): authorize per-message dst_service against the
authenticated service set on the ievent server.** `ac_on_message`
(subscribe / unsubscribe / inject) and `ac_mt_stats` resolved the
`dst_service` / `service` from the attacker-controlled routing stack and
dispatched against any registered service. A peer authenticated for
service A could subscribe to events of, inject into, or read/reset the
stats of another service B by naming it. Both paths now check the resolved
service against the set this channel is authorized to reach, captured at
identity-card time from the `services_roles` returned by
`gobj_authenticate()` (the keys: the primary `dst_service` plus any
`required_services` the user holds real treedb roles in — see `append_role`
in `c_authz.c`; the no-treedb path yields just the primary service). This
implements the long-standing `available_services` TODO: a single
authentication legitimately grants several services (the multi-service GUI
frontends authenticate against `db_history_wz` and reach
`treedb_wattyzer` / `treedb_authzs` / …), while a service outside the
granted set is refused. The authorized set derives from real roles, not
the client-supplied `required_services`, so it cannot be spoofed.
Validated end-to-end: the wattyzer SPA against a patched `db_history_wz`
logs in, opens its ievent channel, and loads multi-service data with zero
gate denials. (`ac_mt_command` cross-service reach stays gated by the
default-off per-command authz — threat-model T7, an accepted posture.)
- **security(root-linux): resolve the `C_PTY` `process` attr against a
trusted-dir allowlist, never `$PATH`.** The remote-settable `process`
value (set by the authz-gated `open-console` command) reached `execvp()`,
which consults the inherited `$PATH` — a planted PATH entry could hijack a
bare name. New `resolve_process_path()` accepts an absolute path only if
executable, resolves a bare name against a fixed list of system dirs
(`/bin`, `/usr/bin`, `/sbin`, `/usr/sbin`, `/usr/local/bin`), rejects
relative-with-slash, and fails closed (empty argv[0] → no exec).
`execvp` → `execv`.
- **security(ycommand/ycli): sanitize peer-supplied config record names
before they become local filenames.** A malicious peer's command-answer
record `name`/`id` (`view-config` / `read-json` / `read-file` /
`edit-config`) flowed unsanitized into the `"<editor> <path>"` string that
`pty_sync_spawn()` hands to `/bin/sh -c` — RCE on the operator host (e.g.
`"x; rm -rf ~ #"`). New `sanitize_config_name()` folds everything outside
`[A-Za-z0-9._-]` to `_`, forbids a leading dot, and collapses path
separators to a single inert basename, applied in all `save_local_*`
builders of both tools.
- **harden(ytls/mbedtls): opted-in insecure client is never silent —
observability parity with openssl.** An accepted
`ssl_allow_insecure_client=true` client now logs the same *"TLS client
WITHOUT server-certificate validation (MITM surface)"* warning as the
openssl backend, and runs the handshake under `VERIFY_OPTIONAL` instead of
`NONE` so mbedTLS still computes the verify result and the tolerated
failure is surfaced at handshake end (openssl records it natively even
under `VERIFY_NONE`; mbedTLS skips verification entirely under `NONE`).
The accept decision is unchanged (`OPTIONAL` never aborts; `CA_CHAIN_REQUIRED`
fires only under `REQUIRED`, and the missing-hostname hard error is
`REQUIRED`-only too). The *"did NOT verify"* warning guard now keys off
the effective authmode instead of `has_ca_cert` (under `NONE`,
`verify_result` holds `BADCERT_SKIP_VERIFY` and must not false-fire).
Fixes the 11 TLS ctest failures under an mbedTLS-only `.config` — the
test expectations had encoded openssl-only emissions. Verified 112/112
with each backend.
- **security(gobj-c): reject `gbuffer_create()` `data_size == SIZE_MAX`.**
`GBMEM_MALLOC(data_size+1)` wrapped to `malloc(0)` — a non-NULL ~0-byte
buffer that slips past the `__max_block__` guard while `gbuf->data_size`
stays `SIZE_MAX`, defeating every later bounds check (`gbuffer_freebytes`
/ append `memmove`). Now rejected at creation. Regression test
`test_gbuffer_guards.c::test_wrap_guard`.
- **security(gobj-c): fix OOB heap over-read in `kwid.c` `collapse()`.** The
in-tree `gbmem_strndup(str, size)` is a raw `memmove(s, str, size)` (not a
real `strndup`), so building the path with
`gbmem_strndup(path, strlen(path)+strlen(key)+2)` read past `path` by
`strlen(key)+1` bytes and left the buffer unterminated before the
`strcat`s. Replaced with `GBMEM_MALLOC` + `strcpy` at the exact size.
- **security(libjwt): pin the exact JWT algorithm, not just the key family.**
On the common pinned path (`config->alg == config->key->alg`, both set) the
alg-vs-alg chain in `__verify_config_post` never compared the token alg, and
the `kty` backstop is only family-granular (`jwt_alg_required_kty` maps every
RS/PS alg to RSA) — so e.g. an **RS512 token verified against an RS256-pinned
key**. Added an exact-alg check against whichever alg is pinned. Purely
additive: no legitimate token is newly rejected. Complements the
GHSA-q843-6q5f-w55g cross-family fix already in 7.x.
- **security(yev_loop): fix use-after-free when a callback destroys its own
event.** A callback calling `yev_destroy_event()` on its own event freed it
synchronously, then `callback_cqe`'s re-arm block and dispatch tail
dereferenced freed memory (last in-flight CQE). New `in_dispatch` flag
defers the free to the dispatch tail; the re-arm blocks now also test
`!destroy_requested` so a dying event is never re-armed.
- **security(root-linux): guard NULL header value in `ghttp_parser.c`
`on_header_value()`.** A previous chunk's `json_string()` failing on invalid
UTF-8 leaves no value under `cur_key`; the next continuation chunk then hit
`strlen(NULL)` (attacker bytes + TCP segmentation). Guarded before `strlen`,
restart the accumulator from the current chunk, and log the store failure
instead of silently truncating.
- **security(root-linux): fix re-entrant-free UAF and TLS-error teardown in
`c_tcp.c` `set_secure_connected()`.** The post-handshake `ytls_flush()` can
re-enter and destroy the connection (an `EV_RX_DATA` subscriber) or report a
TLS error. The return was ignored, so `start_pending_writes()` then ran on a
freed gobj. Now: `-2222` (re-entrant free) bails without touching gobj/priv;
`-1111` (TLS error, gobj alive) calls `try_to_stop_yevents()` and returns —
matching the decrypt-path discipline. Requires the ytls change below.
- **security(ytls): propagate `flush_clear_data()` errors out of openssl
`flush()`.** It swallowed the negative return (including the `-2222`
re-entrant-free sentinel); now returns it so `c_tcp` can act on it.
- **security(emailsender): reject CR/LF/control chars in `attachment` and
`inline_file_id`.** Both reach MIME headers raw via `append_attachment_part()`
(`Content-Type name=` / `Content-Disposition filename=` / `Content-ID`), and
`EV_SEND_EMAIL` is public — so they were an SMTP/MIME header-injection vector.
Added to the single-line control-char rejection set alongside the envelope
and display fields.
- **security(gobj-c): pre-auth NULL deref in `gbuffer_deserialize()`.** A
malformed base64 `data` field makes `gbuffer_base64_to_binary()` return
NULL, fed straight into the unguarded `gbuffer_setmark()` inline — a daemon
crash reachable pre-auth via the ievent server (`ac_on_message` →
`kw_deserialize`). Added the NULL check after decode plus a central guard
on the `gbuffer_setmark`/`getmark` inlines. Same change: `kwid_find_record_in_list()`
returned `0` (a valid index) on not-found instead of `-1` (silent
wrong-record match in the list comparator), and the flatten/unflatten
helpers in `kwid.c` moved off raw libc `malloc`/`strdup`/`free` onto the
mandated `gbmem_*`. Regression tests in `test_kw1.c`.
- **security(libjwt): make the JWT verify contract fail closed.**
`jwt_checker_verify2()` handed back the parsed claims regardless of outcome
— the verdict lived only in `jwt_checker_error()`, so a caller trusting the
non-NULL return would accept a forged / expired / alg-confused / unsigned
token. It now returns NULL on any verification failure (the return value
carries the verdict) and `jwt_verify_complete()` aborts early on
`__verify_config_post` failure. Also fixed a wrong-free on the
`jwk_process_one()` OOM path (freed the borrowed `jwk`, not the owned
`item`). Regression: `test_jwt_alg_confusion.c::test_verify2_fail_closed`.
- **security(timeranger2): validate pkey/id against path traversal.** A
string primary key or treedb/msg2db node id becomes a `keys/<key>/`
directory component, so an attacker-influenced value containing `/` or
beginning with `.` could escape the topic's `keys/` dir on append
(mkrdir/newfile), `tranger2_delete_key` (rmrdir), the disk mirror, and
`tranger2_delete_instance`. Rejected at every sink in `tranger2_append_record`
/ `tranger2_delete_key` / `tranger2_delete_instance` / `treedb_create_node`
/ `msg2db_append_message`. Regression:
`tests/c/timeranger2/test_pkey_path_traversal.c`.
- **security(yev_loop): connect() the static DNS resolver socket.** The
`CONFIG_FULLY_STATIC` resolver read UDP replies from any source, so
authenticity rested only on the 16-bit transaction id — an off-path
attacker could forge an A/AAAA answer and redirect a yuno's outbound
connection. `dns_query()` now `connect()`s the UDP socket to the chosen
nameserver (both IPv4/IPv6 branches) so the kernel drops datagrams from any
other source. Regression:
`tests/c/yev_loop/static_resolv/test_static_resolv_spoof.c`.
- **security(ytls): verify-by-default for TLS clients (BREAKING).** A TLS
*client* that would run `VERIFY_NONE` (no CA / effective authmode NONE) is
now refused at ctx/state build time in both backends instead of merely
logging a warning — closing a live MITM hole (the auth_bff → Keycloak
outbound client ran unverified). Opt back in per gate with the new
`ssl_allow_insecure_client=true` (default false) for self-signed / PSK /
IoT bring-up. The mbedTLS gate keys off the effective authmode for openssl
parity. The `C_AUTH_BFF` `crypto` and `c_authz` `kc_crypto` defaults
flipped to a verifying posture (`ssl_use_system_ca` + `ssl_verify_mode=required`).
**Rollout:** any TLS-client deployment relying on silent `VERIFY_NONE` must
add a CA (or `ssl_allow_insecure_client=true`) before it will connect.
- **fix(root-linux): check `ytls_init()` NULL in `c_tcp.c` connect path.**
Follow-up to verify-by-default: the refusal is a soft failure (`ytls_init`
returns NULL, yuno stays up), but `ac_connect` never checked it and the
established connection then called `ytls_new_secure_filter(NULL, ...)` —
SEGFAULT (caught by `perf_c_tcps` test4/test5). Now the connect aborts
cleanly via `try_to_stop_yevents()`. Also added the missed
`ssl_allow_insecure_client=true` to the `perf_c_tcps` test4/test5 client
crypto blocks (the `tests/c/c_tcps*` sweep skipped `performance/`).
- **harden(yuno_agent/yuno_agent22): pin the controlcenter client to the
canonical agent certificate.** The outbound controlcenter client
(`tcps://<arch>.<owner>.<output_url>`, active when `node_owner != "none"`)
ran with no CA and would be refused under verify-by-default. It now pins
`/yuneta/agent/certs/yuneta_agent.crt` — the self-signed canonical cert
the controlcenter actually serves (fingerprint-verified), already present
on every node (the agent's own wss server uses the same file) — with
`ssl_server_name=yuneta_agent.yuneta.io` (the cert has no SAN, hostname
check falls back to CN). Supporting change in `c_tcp.c`: a config-supplied
`ssl_server_name` now wins over the url-derived host, enabling pinning
where the pinned cert's name differs from the dialed host. **Caveats:**
the pin authenticates "a yuneta install" (the key ships on every node),
not the controlcenter specifically — still a real upgrade over
no-verification; and on a node missing the cert file the first
controlcenter dial exits the agent (`ssl_trusted_certificate` load is
fatal on first init) — verify the file exists before deploying with a
non-none owner.7.5.12¶
- **fix(packages): default agent `node_owner` to `"none"` — no controlcenter
on a fresh node.** The bundled `yuneta_agent.json.sample` /
`yuneta_agent22.json.sample` shipped `"node_owner": "owner"`, a placeholder.
The agent starts the controlcenter client whenever `node_owner != "none"`
(`c_agent.c` `mt_start`), so a freshly installed standalone node kept
dialing `tcps://<arch>.owner.yunetacontrol.com:1994` and logging
`getaddrinfo() FAILED` every ~17 s. The default is now `"none"`, the
design's built-in off-switch, so a fresh box is quiet. Operators **with** a
controlcenter still opt in with `YUNETA_OWNER=mycompany` at install time
(the postinst sed now swaps `"none"` → their owner). Note: blanking the
config to `{}` does **not** silence it — the framework default for
`node_owner` is `""`, which is also `!= "none"`; `"none"` must be explicit.
Affects fresh installs only (the `.json` files are conffiles, never
overwritten on upgrade).
- **refactor(packages): minimal, sanitized agent config templates.** The two
`*.json.sample` templates are trimmed to the smallest valid standalone
baseline (the operator parameterizes per project afterwards): replaced the
deprecated `authz.authz_yuno_role` key with `authz.authz_service` (the
former is `SDF_DEPRECATED` in `c_authz.c`), and dropped the confusing
`__realm_id__` override (`"/yuneta_agent.trdb"`) — the agent's treedb store
dir is now inherited from the compiled-in `main.c` default
(`/yuneta/store/agent/yuneta_agent.trdb`), exactly as a real parameterized
node does. No `authz.jwks` / `authz.initial_load` in the template: the
`root` role + `yuneta` user come from `main.c` via the config merge, so a
fresh agent still bootstraps local `ycommand` on port 1991.7.5.11¶
- **security(ext-libs): bump vendored OpenSSL 3.6.2 → 3.6.3.** Security patch
release (`configure-libs.sh` v1.16, `TAG_OPENSSL=openssl-3.6.3`). Fixes one
**High** CVE — CVE-2026-45447, heap use-after-free in `PKCS7_verify()` —
plus a batch of CMS / QUIC / ASN.1 / AES CVEs (CVE-2026-34180..34183,
35188, 42764..42770, 45445/45446, 7383, 9076). No API change (3.6 series).
OpenSSL is linked **statically into every yuno**, so every yuno must be
rebuilt + relinked to pick it up; the release CI builds ext-libs fresh, so
the published `.deb`/`.rpm` get it automatically. Stayed on 3.6 (not 4.0),
same LTS rationale as before.
- **feat(install): no prompt — `install.sh` runs straight through.** The
installer no longer asks `Install the developer toolchain? [Y/n]` mid-run;
it installs everything in one pass without stops. Use `--runtime-only` to
skip the toolchain on a pure deployment box. (Served from `main`, so it
ships on push.)7.5.10¶
- **fix(deb): drop obsolete `libpcre3-dev` from the dev-deps helper.** PCRE1
(`libpcre3-dev`) was removed from current Ubuntu (26.04) — it is "referred
to but has no installation candidate", so the helper printed
`[!] Failed: libpcre3-dev`. Yuneta does not need it (it builds its own
static PCRE2, and the bundled nginx is static), so it is removed;
`libpcre2-dev` stays. The `apt-cache show` guard didn't catch it (the
transitional record still resolves), so the helper now just installs and
reports the real failures.
- **fix(deb): honest dev-deps end summary.** The Debian helper ended with a
vague `Dev environment setup attempt complete` and only printed failures
inline; it now collects them and reports `all N packages installed` or the
exact list that did NOT install, matching the `.rpm` helper.7.5.9¶
- **fix(deb): never auto-reboot in `postinst`.** The Debian `postinst` forced
a reboot at the end of install (auto-yes when non-interactive), which under
`curl | sh` rebooted the box mid-flight — killing the SSH session before
`install.sh` could install the developer toolchain. The kernel tuning is
already applied live (`sysctl --system`), so a reboot is not required: it
now only leaves the `reboot-required` hint and recommends a reboot, never
forcing one. Matches the `.rpm` `%post` policy.
- **feat(install): `install.sh` installs certbot on both distros.** After the
package, the installer now runs the bundled certbot helper (snap on Debian,
EPEL `dnf` on RHEL) so TLS for the bundled web server is set up in the same
run, regardless of `--runtime-only` (it is a runtime/ops tool).7.5.8¶
- **fix(rpm): dev-deps helper used a dnf5-only flag that RHEL 9 rejects.**
The 7.5.7 helper ran `dnf install --skip-unavailable`, but
`--skip-unavailable` only exists in dnf5 (Fedora); RHEL 9 / Rocky 9 ship
dnf4, which errors `unrecognized arguments: --skip-unavailable` and aborts
the whole transaction — so the toolchain (git, clang, gcc, wget, …)
installed nothing again. Now uses `--setopt=strict=0`, the dnf4-native way
to skip unavailable packages (and valid on dnf5 too).
- **fix(rpm): create the `nogroup` group so the bundled nginx starts.** nginx
falls back to its compiled-default group `nogroup`, which exists on Debian
but not on RHEL; without it nginx aborted at startup
(`getgrnam("nogroup") failed`). `%post` now creates it (RHEL-only, before
the service starts).
- **fix(rpm): honest web-server start in the init script.** `start_web()` ran
`nginx || true; log_end_msg 0`, printing `OK` even when nginx failed to
start. It now captures the real exit code and reports it, like the agent
start. (Matches the no-silent-failure rule.)
- **feat(install): `install.sh` is now a single cross-distro installer that
sets up everything in one run.** It detects the distro (`apt` vs `dnf`),
and on RHEL/Rocky/Alma enables **EPEL + CRB** first; pulls the matching
package (`.deb` / `.rpm`) from the latest Release and installs it; then
installs the **full developer toolchain** (git, mercurial, clang, gcc,
cmake, ninja, wget, pipx, …) by delegating to the bundled, resilient
`/yuneta/bin/install-yuneta-dev-deps.sh` — so a fresh box is build-ready
from one command, with no second script to remember. Asks first when a
terminal is attached (reads `/dev/tty`, so it works under `curl | sh`);
installs by default when non-interactive. `--runtime-only` skips the
toolchain for pure deployment boxes. Served from `main`, so it reaches
users on push (it installs the latest published Release packages). Was
Debian-only and runtime-only before.7.5.7¶
- **fix(rpm): dev-deps helper no longer installs nothing when one package is
unavailable.** `install-yuneta-dev-deps.sh` ran `dnf -y install "${PKGS[@]}"`
as a single atomic transaction, so one unfindable package (typically a
CRB-only `-devel`/`-static` when CRB was never enabled) aborted the WHOLE
set — leaving `clang`, `wget`, `gcc`, `cmake` … none installed — and the
`|| echo continuing` + final `[✓] complete` hid it. It now uses
`dnf --skip-unavailable` (installs every available package, skips only the
missing) and reports per-package with `rpm -q` which ones did NOT install,
pointing at EPEL/CRB when a `-devel`/`-static` is absent, instead of a
green lie. Matches the `.deb` helper's per-package resilience. (`wget` and
`clang` are dev-deps installed by this helper, not by the base `.rpm`.)7.5.6¶
- **fix(rpm): honest agent start in `%post`; don't source the `set -u`-unsafe
RHEL init functions.** Two RHEL-only packaging bugs in the 7.5.5 `.rpm`
(the `.deb` was never affected). (1) The generated `/etc/init.d/yuneta_agent`
runs under `set -u`; on RHEL `/lib/lsb/init-functions` is absent so it fell
to sourcing `/etc/init.d/functions`, which references unset vars
(`SYSTEMCTL_SKIP_REDIRECT`…) and aborted the whole init script with
"unbound variable" **before the agent was ever launched** — the binary was
fine, the service "failed". It now defines the only two functions it uses
(`log_daemon_msg`/`log_end_msg`) itself and only sources Debian's
`set -u`-clean LSB file when present. (2) `%post` started the agent with
`service start || true`, hiding a failed start behind RPM's always-"Complete"
transaction. It now re-reads the effective `kernel.io_uring_disabled` after
`sysctl --system`, only starts when io_uring is usable, captures the real
result, and prints a loud "AGENT IS NOT RUNNING" warning with diagnosis
hints (`systemctl status` / `journalctl` / `getenforce` for SELinux) instead
of a green install over a dead agent.7.5.5¶
- **refactor(packaging): split Debian packaging into `packages/deb/`.**
With the new `packages/rpm/`, the Debian scripts moved from the root of
`packages/` into a sibling `packages/deb/` (`AMD64`/`ARM32`/`ARMhf`/
`RISCV64` wrappers + `make-yuneta-agent-deb.sh` + its `README.md`), so the
two packagers are now symmetric: `packages/deb/` and `packages/rpm/`. The
shared agent config samples stay at `packages/templates/` (referenced by
both via an absolute `$YUNETAS_BASE/packages/templates/` path);
`packages/README.md` is now a short index. The deb arch wrappers read
`../../YUNETA_VERSION` + `../../RELEASE` (one level deeper); the release CI
and the `.gitignore` `authorized_keys`/`webserver` rules follow the move.
No change to the produced `.deb` or its contents.
- **feat(build): RHEL/Rocky/Alma build support (was Debian-only).**
Yuneta now builds and runs on the RHEL family; verified end-to-end on
Rocky Linux 9.7 (full static build + 110/110 ctest). New
`install-dependencies.sh` auto-detects the distro from `/etc/os-release`
and installs the right packages with `apt` (Debian) or `dnf` (RHEL,
enabling EPEL + CRB). RHEL-specific build fixes that ride along:
`configure-libs.sh` v1.15 forces `-DCMAKE_INSTALL_LIBDIR=lib` on the
CMake libs (mbedtls/pcre2/jansson/argp), which on RHEL default to
`lib64` and so were missed by the kernel's `outputs_ext/lib` link path
(no-op on Debian); `set_compiler.sh` gained a `dnf reinstall` branch;
and the postgres module includes `<libpq-fe.h>` (not
`<postgresql/libpq-fe.h>`) with the dir resolved via `pg_config` in
CMake, since the libpq header sits in `/usr/include` on RHEL vs
`/usr/include/postgresql` on Debian. RHEL also needs `glibc-static`/
`libstdc++-static`/`libxcrypt-static` (CRB) for the default static link.
- **feat(runtime): document the io_uring requirement on RHEL.** Yuneta's
`yev_loop` is io_uring-based, and RHEL 9 / Rocky 9 / Alma 9 ship
`kernel.io_uring_disabled=2` (fully disabled), so every yuno aborts at
startup until it is re-enabled (`kernel.io_uring_disabled=0`). Called
out in `installation.md`; `install-dependencies.sh` warns when it
detects the disabled state. No code change — a deployment prerequisite.
- **feat(packaging): RPM packaging for the Yuneta Agent (`packages/rpm/`).**
Counterpart of the Debian `packages/`: stages the same `/yuneta` payload
and builds an `.rpm` with `rpmbuild` (`make-yuneta-agent-rpm.sh` +
`x86_64`/`aarch64`/`riscv64` wrappers). RHEL-specific envelope: `.spec`
instead of `control`, `%post`/`%preun`/`%postun` instead of
maintainer scripts, `useradd`/`wheel`/`chkconfig`/langpacks/EPEL-certbot,
and the shipped `kernel.io_uring_disabled=0`. Built + inspected on Rocky
9.7 (`rpm -qlp`/`--scripts`/`rpmlint`); not installed.
- **ci(release): `release-deb.yml` -> `release-packages.yml`, now also
publishes the x86_64 `.rpm`.** The release job builds the `.rpm` next to
the AMD64 `.deb` and uploads both as release assets. It runs on the same
Ubuntu runner: the default build is fully static, so the binaries also
run on RHEL/Rocky — no EL9 container needed.
- **fix(tests): make `ytls/test_cert_info` portable to OpenSSL >= 3.5.**
The expired-cert helper used `openssl x509 -req -days -1`, which
OpenSSL >= 3.5 rejects ("end date before start date"). It now falls back
to `-not_before/-not_after` with fixed past dates (OpenSSL >= 3.2) when
the negative-span form fails, keeping older OpenSSL (e.g. Debian 12's
3.0.x) working. Not RHEL-specific — any host with a modern OpenSSL.
- **fix(emu_device): don't truncate the final frame on standalone exit.**
`finish_replay()` called `exit(0)` right after queueing the last frames,
before io_uring completed the write — the final frame could be lost. In
standalone CLI mode it now waits for the C_TCP `EV_TX_READY` drain signal
(tx queue empty + in-flight write completed) and exits from `ac_tx_ready`;
empty replays still exit immediately. Verified end-to-end: a 3-frame replay
delivered all bytes including the last frame, clean exit. Agent-managed
mode is unchanged (it never exited).
- **chore(auth_bff): remove the deprecated `idp_url` + `realm` pair.** The
legacy Keycloak path-scheme fallback (build
`<idp_url>/realms/<realm>/protocol/openid-connect/{token,logout}`) was
`SDF_DEPRECATED` since the 2026-04-30 OIDC migration and a release has
shipped with the warning. Both `SDATA` attrs and the resolution branch in
`c_auth_bff` `mt_create` are gone; configure `issuer` (discovery) or the
explicit `token_endpoint` + `end_session_endpoint` instead. No yunetas or
private deployment still set the legacy pair (all on `issuer`). Docs
updated (`YUNO_AUTH.md` §2.5, `guide_oauth2_pkce_bff.md`). The now-subjectless
`tests/c/c_auth_bff/test17_legacy_idp_url` is removed; the suite is 18/18.
- **docs(auth): ROPC in `c_task_authenticate` deferred by design.** The CLI
grant stays `grant_type=password` (works on Keycloak, the only deployed
IdP). Documented the constraint and the real migration path (device-flow
for interactive + client-credentials for headless CI, not loopback PKCE —
the six CLI callers have no browser) in the `c_task_authenticate.c` header,
`YUNO_AUTH.md` §3.4, and `TODO.md`. No behavior change.v7.5.4 -- 08/Jun/2026¶
- **feat(mqtt/security): subscribe-side ACL enforcement.** Completes the
publish/subscribe ACL started in 7.5.3. The per-topic SUBACK reason is
built in the broker's `ac_mqtt_subscribe`, so the check lives there
(alongside the existing `deny_subscribes` gate), calling
`mqtt_acl_check(…, "read")` per requested filter — a denied filter is not
added and gets a `MQTT_RC_NOT_AUTHORIZED` (v5) / `0x80` (v3.x) SUBACK
reason, logged. Unchanged when `enable_acl` is off or a group has no
`subscribe_acl` patterns.
- **fix(security/authz): per-command authz gate redesign.** A local agent
pilot showed the 7.5.3 gate (`enable_command_authz`) was undeployable — it
denied a yuno's own internal startup commands (e.g. `open-treedb`) and the
yuno exited. Two bugs fixed in gobj-c: (A) a specific-authz lookup on a
concrete gobj now falls back to the **global** authz table
(`authzs_list`), so `__execute_command__` resolves on any gobj (was
"authz not found" → deny-all, root included); (B) the gate fires **only
for external commands** (those whose kw carries the authenticated
`__username__` injected by `c_ievent_srv`), so internal `gobj_command()`
calls are never gated (`command_parser`). Re-piloted: the agent now boots
clean with the gate on and `ycommand` (root) works. `enable_command_authz`
remains **default-off**.
- **feat(security/authz): seed root role model** in the C_AUTHZ yunos that
lacked one (`controlcenter`, `mqtt_broker`, `emailsender`) via
`Authz.initial_load` (role `root` + user `yuneta`, mirroring the agent),
a prerequisite for enabling the command-authz gate there.v7.5.3 -- 08/Jun/2026¶
- **feat(security): per-command authorization re-armed (gated opt-in).** The
`SDF_AUTHZ_X` check at the command-dispatch boundary in `command_parser.c`
(commented out for years) now runs again, but only when the yuno sets the
new `enable_command_authz` attr (`c_yuno`, `SDF_RD`, default `"0"`), so the
default posture is unchanged and non-breaking. Self-issued commands
(`src == gobj`) bypass the check; a denial returns `-403` and is logged
(`MSGSET_AUTH`). Turning the gate on requires a running `C_AUTHZ` role
model (the global checker is fail-closed). Test:
`tests/c/command_authz/`.
- **feat(mqtt): publish-side ACL (model A, group-based, default off).** New
`enable_acl` attr on `C_MQTT_BROKER` plus `publish_acl`/`subscribe_acl`
array columns on the `client_groups` topic (schema_version 25→26,
topic_version 3→4; additive). `C_PROT_MQTT2` queries the broker over a
direct `EV_MQTT_ACL_CHECK` event on PUBLISH; allow when ACL off or no
patterns authored, deny unknown clients, deny logged. Subscribe-side
wiring is staged (schema + helper ready) but not yet enforced. Test:
`tests/c/c_mqtt/acl`. Also fixed a latent fkey bug: the helper now passes
`{fkey_only_id:1}` so `client_groups` resolves as plain ids (otherwise the
ACL silently allowed all).
- **fix(emu_device): frame emission path.** `window`/`interval` were coerced
to 0 (CLI values passed as `json_string` into `DTP_INTEGER` attrs;
`cmd_write_*` used `kw_get_str`+`atoi`), so the replay sent nothing. Now
`json_integer(atoi(...))` on the CLI side and `kw_get_int(KW_WILD_NUMBER)`
in the commands; also freed the replay resources on `mt_play` error paths
and log a skipped record with no `frame64`.
- **refactor(emu_device): moved from `yunos/c/` to `utils/c/`.** It is a
standalone CLI utility yuno (a device-gate emulator run by hand for
testing), not a deployable service — now installed to `/yuneta/bin` like
the other `utils/c` tools.
- **docs:** `YUNO_AUTH.md` rewritten to describe the gated authz (was
documented as "commented out", including the auth-flow diagram);
`mqtt_broker.md` gains an Authorization (ACL) section.v7.5.2 -- 06/Jun/2026¶
- **security: hardening batch (memory-safety + injection across the stack).**
- **gobj-c:** NULL-guard in `gbuffer_deserialize`; bounded recursion in
`kw_find_path` and the kwid comparators (hostile-JSON stack exhaustion).
- **root-linux:** NUL-terminate accumulated URL / header field / header
value in `ghttp_parser` (over-read fed to `json_string`); guard
`/auth/logout` against an uninitialized refresh-token read.
- **ytls:** re-entrant use-after-free in `encrypt_data` WANT path + a
double `gbuffer_get` (stream corruption).
- **yev_loop:** bound DNS response parsing + unpredictable transaction id;
defer event free until in-flight io_uring CQEs drain (UAF).
- **timeranger2:** validate on-disk md2 `__offset__`/`__size__` before
read AND before the `delete_instance` payload wipe (cross-record
overwrite); guard `read_md` return; bound the inotify parse loop.
- **libjwt:** backport GHSA-q843-6q5f-w55g algorithm-confusion JWT forgery
fix + the full cfd8902 hardening; add an in-tree regression test
(`tests/c/libjwt/`, RSA/EC/EdDSA/`alg:none`).
- **modbus:** reject MBAP length < 3 (heap overflow). **dba_postgres:**
escape SQL identifiers/literals in insert + create-table (SQLi).
- **emailsender:** reject CR/LF in header/envelope fields (SMTP/MIME
injection). **mqtt:** reject property-length underflow in C_PROT_MQTT2.v7.5.1 -- 03/Jun/2026¶
- **fix(treedb): multi-version parent reverse-hook hygiene.** Two
in-memory hook quirks around versioned (pkey2) parents are fixed at the
treedb layer:
- **Unlink targeted only the primary parent version.** A child's fkey
ref carries just `parent_topic^parent_id^hook` (no version), so
`treedb_clean_node` unlinked from the PRIMARY instance, leaving a
stale entry on the non-primary version the child was actually hooked
on. It now locates the parent-version instance that really holds the
child (`find_parent_version_holding_child` over the pkey2 index) and
unlinks that one; the primary-instance behaviour is the fallback for
hook+fkey combos the read-only probe can't match.
- **Duplicate hook entries.** Repeated create/link of the same child id
left it more than once in the parent hook (`yunos:["5000","5000"]`),
inflating "Using in N". `_link_nodes` and the loader
`link_child_to_parent` now dedup by child id before appending
(idempotent link), and the child-side fkey array is deduped too.
A skipped duplicate is **warned** (not silently swallowed): a re-link, or
a duplicate fkey self-healed on load, is surfaced via `gobj_log_warning`.
Closes the TODO follow-up; both quirks also self-heal on reload.v7.5.0 -- 03/Jun/2026¶
- **fix(agent): version-aware, stale-safe `delete-config`/`delete-binary`
usage guard.** The "Using in N yunos" guard read the raw `yunos` hook
count, which is config-id-level (shared across versions) and can carry
stale/duplicate refs — so an UNUSED config/binary version could not be
pruned while another version was in use, and lingering refs blocked
deletes. New `count_yunos_using()` validates every hooked yuno id via
`gobj_get_node` (a deleted yuno → NULL → skipped) and, when a version is
given, counts only yunos pinned to THAT version (config↔`name_version`,
binary↔`role_version`). So an unused/superseded version prunes cleanly,
only the in-use version blocks, and `force=1` overrides. Combined with the
durable per-instance delete below, `delete-config version=`/`delete-binary
version=` now remove a single version durably. (Underlying treedb
multi-version reverse-hook hygiene remains a minor follow-up — see TODO.md.)
- **feat(treedb): durable per-instance (pkey2) node delete.**
`treedb_delete_instance` now tombstones EVERY md2 row belonging to a
`(key, pkey2_value)` via `tranger2_delete_instance()` (enumerated with a
transient disk list) and drops the in-memory pkey2 slot — previously it
only dropped the in-memory slot, so the instance resurrected on the next
reopen (a treedb instance spans several rows: create + each link/update
re-appends one with the same id/pkey2; tombstoning just the latest let an
earlier row reload). The primary index is untouched (callers route only
non-primary instances; the loader re-elects the highest surviving rowid).
Whole-key delete (`treedb_delete_node`/`tranger2_delete_key`) is unchanged.
Exposed through the agent: `delete-yuno`/`delete-config`/`delete-binary`
now accept the pkey2 (`yuno_release`/`version`) to prune a single
non-primary release/version, listing via `gobj_list_instances` and
reading the running-guard from the primary (instance records carry a
stale `yuno_running`). New regression test covers a multi-row instance +
close/reopen. (Remaining follow-up: stale reverse-hooks on linked parents
— see TODO.md.)
- **fix(agent): `find-new-yunos` inherits node placement across a version
bump.** A version-bump deploy (`install-binary` + `find-new-yunos
create=1` + `deactivate-snap`) created the fresh `yunos` row at schema
defaults, dropping the operator-set `start_priority`/`sched_priority`/
`cpu_core` and forcing a re-run of `tools/agent/set_start_priorities.py`.
`cmd_find_new_yunos` now copies those three fields from the prior primary
row into the emitted `create-yuno` command (and `pm_create_yuno` accepts
them), so launch tiers and CPU placement survive the bump. Same-version
REBUILD hot-patches were already unaffected (they keep the existing row);
genuinely-new yunos still get defaults plus the `util`-tag seed.v7.4.8 -- 03/Jun/2026¶
- **feat(agent): per-yuno `start_priority` launch tiers.** The agent's
`yunos` topic gains `start_priority` (band 0..9, default 5). `run-yuno`
launches ascending (utilities first), `kill-yuno`/`pause-yuno` descending
(utilities last, so logcenter captures everyone's shutdown), stable within
a tier. The node-wide relaunch (`run_enabled_yunos`, used by
`restart_nodes`/`deactivate-snap` and at startup) honours the same order;
the force-SIGKILL pass stays unordered (no graceful drain to sequence).
`create-yuno` seeds `start_priority=1` for `util`-tagged yunos (the set
`run_util_yunos` already starts first) — no app role names in the agent.
Schema `topic_version` 19→20 + `schema_version` 22→23; the bump only
refreshes the col schema files, record data is untouched (no store wipe).
Assign app tiers per node with `tools/agent/set_start_priorities.py`.
- **feat(agent): node CPU placement (`sched_priority`, `cpu_core`) from the
agent treedb.** Both are new `yunos` columns the agent injects into the
launched yuno's config as its `sched_priority`/`cpu_core` attrs, so OS
scheduling/affinity is a node-local decision instead of being baked into
the config that travels across nodes. Defaults only: the user config file
is merged after the agent's and still wins. `cpu_core=0` (default) = no
boost, unchanged behaviour.
- **refactor(c_yuno): scheduling attr `priority` renamed to `sched_priority`.**
The `sched_setscheduler` attr (default 20, applied only when `cpu_core>0`)
collided with the per-service start order (0..9) and the agent's
`start_priority`; renamed so the name states what it does. `SDF_PERSIST`
fallback is a no-op in practice (consulted only at `cpu_core>0`, which no
shipped yuno sets); no migration shim. The per-service `priority` is
unchanged. See `TODO.md`.
- **refactor(agent): `set-ordered-kill` renamed to `set-graceful-kill`.** The
command never ordered anything — it only sets `signal2kill=SIGQUIT` (the
yuno catches it and shuts itself down cleanly). Renamed to the honest axis
(graceful SIGQUIT vs quick SIGKILL, `set-quick-kill`), which also frees
"ordered" for the real `start_priority` ordering above. No alias kept.
- **feat(tools): `set_start_priorities.py` — assign `start_priority` by role.**
One-shot operator tool mapping each managed yuno's role to a launch tier
(defaults: utilities=1, `gate_*`=4, `db_*`=7; unmatched left as-is) and
writing the differences via `update-node` (record base64'd into
`content64`; the inline `record={...}` form is not coerced by the CLI).
`--rule PATTERN=PRIO` adds/overrides (matched before the built-ins),
`--dry-run`/`--all`/`--show-all`. Same OAuth2-once + `-j` plumbing as the
other agent scripts, so it can drive a remote wss:// agent.
- **feat(tools): `sync_binaries.py`/`sync_configs.py` order restarts by
`start_priority`.** Both bounced yunos alphabetically; now they read the
per-yuno `start_priority` the agent exposes via `*list-yunos` and restart
ascending, so infrastructure comes back before its dependents. Both
degrade to the previous order when the agent has no `start_priority` yet.
The quit/decline message also reads `Cancelled - no changes made.` instead
of `Aborted.` (which looked like a crash).
- **feat(c_yuno): `print-role` command — runtime equivalent of `--print-role`.**
Every yuno (the agent included) now answers a `print-role` command that
returns its basic identity: `role`, `name`, `alias`, **`version`** (the
yuno's own APP_VERSION) and **`yuneta_version`** (the framework version),
plus description/tags/required_services/public_services/service_descriptor.
Until now that info was only printable offline via the binary's
`--print-role` flag; there was no way to read a *running* yuno's version.
Lives in C_YUNO's command table, so it is inherited by all yunos. Address
the yuno gobj with the `-S __yuno__` flag: `ycommand -S __yuno__ -c
'print-role'` (inline `service=...` is a command parameter, not routing).
- **feat(tools): `sync_binaries.py` automates the same-version REBUILD
hot-patch.** A `REBUILD` (`update-binary`) overwrites the slot the running
yuno executes from, so it failed with `text-file-busy` and the script only
printed a "kill-yuno first" reminder. Now, once both confirmation gates are
cleared, it runs the documented per-role cycle itself, scoped by
`yuno_role` (never node-wide): `kill-yuno` (only if running; orderly
SIGQUIT, so the gbmem audit runs) → poll `*list-yunos` until the process
exits → `update-binary` → `run-yuno play=0` (if it was running) →
`play-yuno` (if it was playing). Prior run/play state is read from
`*list-yunos` and restored per role, and a role with several instances
across realms is handled in one shot. New `--no-restart` flag keeps the old
print-only behaviour. The version-bump path (`find-new-yunos` +
`deactivate-snap`, a node-wide bounce) stays a reminder.
- **feat(tools): `sync_configs.py` gains an opt-in `--restart`.** Installing
a config does NOT need a kill — unlike `update-binary` (which hits
`text-file-busy` while the yuno runs), a config push always succeeds on a
running yuno; it just does not take effect until that yuno next (re)starts.
So by default the script still only pushes and prints the affected yuno ids
(from the agent record's `yunos` field) as a `kill-yuno` + `run-yuno`
reminder — restarting is a separate, optional step. Pass `--restart` to
also bounce the using yunos right away, scoped by yuno `id` (never
node-wide): `kill-yuno` (only if running; orderly SIGQUIT) → poll
`*list-yunos` until it exits → `run-yuno play=0` → `play-yuno` (if it was
playing), preserving prior run/play state. A stopped yuno is left stopped;
NEW configs (no agent record) print a reminder.v7.4.7 -- 02/Jun/2026¶
- **feat(c_authz): `create-user` password is now optional.** KC/IdP-
authenticated users have no local password (`credentials` null) — auth is
by JWT. The command no longer rejects an empty password; it only hashes
credentials when one is given, otherwise creates the user password-less,
the same way `register-idp-user` and the `initial_load` users do.
- **fix(c_authz): stop resetting a user's "Created Time" on update.** In
`ac_create_user` the `new_user` flag was inverted (`user?TRUE:FALSE` is
TRUE when the node already exists), so updating an existing user wrote
`time=now` into the record, overwriting its creation timestamp. New users
were unaffected (treedb auto-stamps the `time`-flagged column on create).
Corrected to `user?FALSE:TRUE`.
- **fix(yuno_agent): silence spurious "Event NOT DEFINED in state" on every
login.** C_AGENT subscribes to all of its `authz` service's output events
but had no FSM entry for `EV_AUTHZ_USER_LOGIN`/`LOGOUT`/`NEW` (consumed by
controlcenter from its own local authz), so each login/logout logged an
error. Added accept-and-ignore handlers.
- **fix(ycommand): accept `EV_ON_OPEN_ERROR` in `ST_DISCONNECTED`.** A failed
connect / identity-card NAK publishes `EV_ON_OPEN_ERROR` after the close;
the FSM lacked it and logged "Event NOT DEFINED in state". Now handled like
`EV_ON_ID_NAK` (`ac_on_close`).
- **feat(tools): sync_binaries.py / sync_configs.py — OAuth2 passthrough for
remote agents.** Both scripts now log in ONCE (Keycloak password grant via
stdlib, or a `--jwt` passed verbatim) and thread the token through `-j` to
every `ycommand` call, so they can drive a remote `wss://` agent without
SSH. New flags: `-I/--issuer` (OIDC discovery), `-T/--token-endpoint`,
`-Z/--client-id`, `--client-secret`, `-x/--user-id`, `-X/--user-passw`,
`-j/--jwt`. The no-arg local path is unchanged (no auth). `$$()` already
resolves client-side, so the LOCAL build is what gets uploaded.
- **fix(emailsender): handle `EV_ON_OPEN` in `ST_WAIT_RESPONSE`
(reconnect-on-demand).** When the SMTP link had idle-closed, the head
message is dispatched anyway and `c_smtp_session` reconnects to deliver
it; on reaching `ST_IDLE` it publishes `EV_ON_OPEN` *before* beginning the
stashed message — but `c_emailsender` is already in `ST_WAIT_RESPONSE` (it
changes state before `EV_SEND_MESSAGE`), so every reconnect-to-deliver
cycle logged a spurious *"Event NOT DEFINED in state"* even though the
mail was delivered. `ST_WAIT_RESPONSE` now accepts `EV_ON_OPEN` via
`ac_on_open_waiting`, which only marks the link ready and does NOT
re-dequeue (a message is already in flight). Also corrected the misleading
`ac_disconnected` comment in `c_smtp_session`.
- **feat(c_authz): IdP (Keycloak) user provisioning.** New commands
`register-idp-user` (create the user in Keycloak via the admin REST API +
the local `treedb_authzs` user with the chosen role, then email a
set-password invite), plus `set-kc-config` / `view-kc-config` to configure
the admin connection. The connection params are neutral persistent attrs
(`kc_*`, `SDF_PERSIST`) set at runtime — no endpoints/secrets in code or
committed config; the secret is masked by `view-kc-config`. Lives in
C_AUTHZ so every auth-enabled yuno inherits it. The outbound work is an
async multi-job `C_TASK` over a lazily-created `C_PROT_HTTP_CL`.
- **feat(c_prot_http_cl): accept any JSON value as the request body.**
`data` is now read with `kw_get_dict_value` instead of `kw_get_dict`, so a
JSON array body (e.g. Keycloak `execute-actions-email`) is sent verbatim
via `json_dumps`; the x-www-form-urlencoded path is unchanged (it only
iterates objects).
- **fix(packages): cert-sync no longer reloads TLS on every tick.**
`copy-certs.sh` re-copied the certs each run (mtime bump → spurious
`reload-certs` broadcast every 15 min): GNU `install -C` never skips a
symlink source (letsencrypt `live/*.pem`) and re-copies on root/yuneta
owner mismatch. Now resolves the symlink with `readlink -f` and sets the
owner via `install -o/-g`, dropping the trailing `chown`.v7.4.6 -- 01/Jun/2026¶
- **feat(tools): `tools/agent/sync_binaries.py` — reconcile built yunos
against the agent and push updates.** Drives from the agent's installed
binaries (`ycommand -c '*list-binaries'`), looks each one up in
`outputs/yunos` (`--print-role`), and classifies it
BUMP/DOWNGRADE/REBUILD/UP-TO-DATE/NO-BUILD. After confirmation it runs
`install-binary` / `update-binary id=<role> content64=$$(<role>)` for the
chosen roles; it does not automate the node-wide lifecycle steps
(`kill-yuno`, `find-new-yunos` + `deactivate-snap`) but prints them as
reminders. Lives under `tools/` (shipped in the install `.deb`, usable on
a bare node), and is documented at doc.yuneta.io under the new **Tools**
section.
- **feat(tools): `tools/agent/sync_configs.py` — reconcile a directory's
configs against the agent and push updates.** Config-side sibling of
`sync_binaries.py`. Because configs are not centralized like binaries
(they live under each yuno's `batches/<host>/`), it drives from the
current directory: each `*.json` config's id is its filename (minus
`.json`) and its version is the `__version__` field inside the file
(`_*.json` batch helpers and files without `__version__` are skipped). It
looks each up via `ycommand -c '*list-configs'` and classifies it
NEW/BUMP/UPDATE/UP-TO-DATE/DOWNGRADE/agent-only. After confirmation it runs
`create-config` / `update-config id='<id>' content64=$$(<path>)`; a
DOWNGRADE (local older than the agent) is reported but never pushed. It
prints the affected yuno ids as a `kill-yuno` + `run-yuno` reminder rather
than automating the restart. Documented at doc.yuneta.io under **Tools**.
- **feat(yuno_agent): `install-config` alias for `create-config`.** Added by
analogy with `install-binary`, so config installs read symmetrically with
binary installs (`c_agent.c`). Corrected the config-command docs that this
exposed as stale: `YUNO_LIFECYCLE.md` claimed there was no `install-config`
and that `update-config` "creates and updates" (it only overwrites an
existing `(id, version)`); the onboarding recipes in `YUNO_LIFECYCLE.md` /
`SCAFFOLDING.md` / `YUNO_AUTH.md` used `update-config … version=<v>
zcontent=$$()` — the real form is `create-config … content64=$$()` (version
read from the file's `__version__`).
- **chore(packages): yuno binaries + `tools/agent` on PATH per layout.** In
`make-yuneta-agent-deb.sh` the hardcoded `outputs/yunos` PATH entries moved
into the profile snippet's layout-detection branch (full source tree vs
deployed `.deb` node), and `tools/agent` was added there too, so
`sync_binaries.py` / `sync_configs.py` are runnable by name on a node.
- **chore(tools): retire `tools/docs-migration/`.** The myst migration is
done and the Quarto pilot was abandoned, so the two one-off helpers
(`myst_to_quarto.py`, `strip_toctrees.py`) were removed.
`verify_api_coverage.py` is a repo-dev verifier (not a node tool), so it
moved to `scripts/`; its five stale "extra" reports were resolved by
correct header→landing mapping (no docs removed). `scripts/` is repo-only;
`tools/` ships in the `.deb`.v7.4.5 -- 30/May/2026¶
- **fix(yuno_agent): `delete-yuno`'s snap-tag guard read the wrong metadata
key — it was dead.** `cmd_delete_yuno` read `__md_treedb__`__tag__`, but
the metadata key is `tag` (set in `tr_treedb.c`; the kernel guard
`treedb_delete_node` reads `__md_treedb__`tag`). So the agent-level guard
always saw 0 and never fired — only the kernel `treedb_delete_node`
backstopped the actual delete (with a cryptic message), and the bogus
`KW_REQUIRED` on a missing key risked log noise. Fixed to read
`__md_treedb__`tag` with flag 0 (default 0 = untagged), matching the
kernel guard and the `delete-binary` guard, and clarified the message to
*"tagged by snap N (rollback)"*. Found while adding the `delete-binary`
guard. Verified the key is populated: `snap-content name=<tag>
topic_name=yunos` lists the tagged yuno records.
- **fix(yuno_agent): `delete-binary` refuses to purge a binary a snap
references (clear reason).** A snap pins the binaries it captured:
`shoot-snap` stamps its id on each topic's current-primary record
(md2 `user_flag`, surfaced as `__md_treedb__.tag`), `binaries` included,
and `activate-snap` rolls back to exactly those records — so the binary
file must survive or `run-yuno` fails with *"primary binary not found"*.
The kernel `treedb_delete_node` already refuses a tagged node unless
`force`, and `cmd_delete_binary` breaks before the `rmrdir` when the node
delete fails — so the file was never actually lost. But unlike the
sibling `delete-yuno`, `delete-binary` gave no reason (just a cryptic
kernel log + a generic failure). Added the explicit agent-level guard
(mirroring `delete-yuno`, reading `__md_treedb__.tag`): a snap-tagged
binary is refused with *"referenced by snap N (rollback)"* and `force=1`
overrides — which breaks that snap's rollback, as documented. Verified
the mechanism live: `snap-content name=<tag> topic_name=binaries` lists
the exact binary records the snap pinned.
- **fix(yuno_agent): `list-binaries` shows the binary in use, not every
instance.** `cmd_list_binaries` had been switched (df0e50e70) from
`gobj_list_nodes` to `gobj_list_instances`, which made it return one row
per `(role, version)` — identical to `list-binaries-instances`, and after
a same-version `update-binary` (an append) even two rows for the same
`(role, version)`. The reason given at the time ("the new instance is
invisible in the primary index until `deactivate-snap`") was the pkey2
staleness bug since fixed in `dbf532ec9`. Reverted `list-binaries` to
`gobj_list_nodes("binaries", …)`: ONE node per role — the primary, i.e.
the binary actually in use. `list-binaries-instances` keeps the full
`(role, version)` enumeration. Verified live: `list-binaries` returns 15
rows (one per role, the in-use version) while `list-binaries-instances`
returns 30 (every record). The primary correctly tracks the in-use
version — an `update-binary` updates it in place; an `install-binary` of
a new version only changes it once `deactivate-snap` promotes+reloads it
(correct: the new binary is not in use until then).
- **feat(ycommand): `history` / `!history` work non-interactively, and
fix the local-command hang.** Two problems with the history command
outside `-i`: (1) the line editor (`C_EDITLINE`, `priv->gobj_editline`)
is only created in interactive mode, and both `list_history()` (the bare
`history` intercept) and `cmd_local_history()` (the `!history`
local-table entry) read only that live editor — so `ycommand history` /
`ycommand -c history` printed nothing even though the history is
persisted to `~/.yuneta/history2.txt`. Both now fall back to that file
when there is no live editor. (2) A trailing **local** command in
non-interactive mode hung: the shutdown timeout is scheduled from
`ac_command_answer`, which only fires for **remote** commands (a local
`history` produces no `EV_MT_COMMAND_ANSWER`), so the queue drained and
ycommand waited forever for an answer that never came. Added
`schedule_exit_if_done()` at the tail of `run_next_pending()` — when the
queue is empty, the session is non-interactive, no async command is in
flight, and we are not in long-lived stdin-pipe mode, it schedules the
same shutdown timeout. Interactive sessions and pipe mode (which waits
for EOF) are unaffected.
- **feat(snap-content): friendlier snap inspection.** The agent's
`snap-content` (served by `C_NODE` in `c_node.c`) required the numeric
`snap_id` AND an exact `topic_name`, so you could not ask "where does
this snap point?" without already knowing the topic names. Two additive,
backward-compatible changes: (1) the snap is now selectable by
`snap_id`, `id` (alias), or `name` (resolved against `__snaps__`); the
legacy `snap_id=` keeps working. (2) `topic_name` is now optional — when
omitted, the command returns the **overview** of every topic the snap
tags and how many records each (a cheap count-only walk via a new
`snap_count_cb`, not a full load), e.g. `snap-content name=pre-744` →
`realms:3, yunos:16, binaries:15, configurations:16, public_services:2`.
Pass `topic_name=<topic>` to drill into one topic's foto as before. The
`id`/`name` params were added to the param schema in both `c_node.c` and
the agent's `c_agent.c`.v7.4.4 -- 30/May/2026¶
- **fix(c_websocket): stop synthesizing `EV_ON_OPEN_ERROR` at the
transport layer.** `EV_ON_OPEN_ERROR` is a high-level event owned by
the session layer (`c_ievent_cli`, which emits it with the remote-yuno
identity). The commit that introduced it ("EV_ON_OPEN_ERROR — close
before open") also added an emission in `c_websocket` `ac_disconnected`
for the "transport closed before the WS upgrade completed" case. That
emission is mislayered and has no consumer: no FSM declares
`EV_ON_OPEN_ERROR` as an input action, so `c_websocket` publishing it to
its parent (`C_CHANNEL`, sitting in `ST_CLOSED` because it never opened)
was rejected by `gobj_send_event` with "Event NOT DEFINED in state". On
slower nodes the `run-yuno` reconnect window widens and the race fired
once per affected yuno (1:1 with the close-before-upgrade warning).
`ac_disconnected` now publishes only `EV_ON_CLOSE` when a real session
existed, otherwise returns silently (pre-918be48b9 behavior), and
`EV_ON_OPEN_ERROR` was dropped from `c_websocket` `event_types`. The
high-level emission in `c_ievent_cli` is unchanged.
- **fix(c_websocket): raise default `timeout_handshake` 5s → 30s.**
During a mass yuno launch (`kill-yuno` + `run-yuno`) every yuno's
`agent_client` (`C_IEVENT_CLI` → … → `C_WEBSOCKET`) reconnects to the
single-threaded agent at once; the agent's event loop is stalled doing
launch work (loading binaries, fork/exec, treedb) and could not complete
each WS upgrade handshake within the old 5s window for the yunos at the
back of the queue → "Timeout waiting websocket handshake" in synchronized
bursts (one per launched yuno). The timeout firing was counterproductive:
it `ws_close` + `EV_DROP`s and reconnects after
`timeout_between_connections`, adding load to the very herd that caused
it. The new 30s default sits comfortably above the observed agent
loop-stall during mass launch; the attr is per-instance configurable
(`SDF_PERSIST`) so a public-facing WS server that wants faster dead-peer
detection can still tighten its own. The remaining root-cause work
(jitter on `timeout_between_connections` in `c_tcp` to break the
synchronized reconnect herd) is not addressed here.
- **feat(yuno_agent): single command response for `run-yuno`, plus a
`play` knob.** Scripts driving the agent need exactly ONE answer per
command to stay in sync; `kill-yuno`/`pause-yuno`/`play-yuno` already
do, but `run-yuno` emitted ~2N answers over N yunos. Two independent
causes were fixed: (1) `cmd_run_yuno` created one `C_COUNTER` with
`max_count=1` INSIDE the per-yuno loop (one answer each); it now
aggregates the `EV_ON_OPEN` filters into a single counter with
`max_count=total` AFTER the loop, mirroring kill/pause/play — one
`"N yunos found to run"` answer. (2) The implicit auto-play: on connect
`ac_on_open` reconciles `must_play` by calling `play-yuno`, one async
answer per yuno. A new `run-yuno play=0` parameter (default `1`,
backward-compatible) launches the process(es) WITHOUT auto-play, so a
script does `run-yuno play=0` (1 answer) then `play-yuno` (1 answer,
aggregated over already-running yunos). The suppression is per-launch and
kept **in-memory in the agent**: `run-yuno play=0` records each
`launch_id` in `priv->no_play_launches`, and `ac_on_open` consumes it by
matching the connecting yuno's `identity_card`launch_id`, deleting the
marker on first connect. It is NOT a treedb column and does NOT mutate the
persistent `must_play`; a watcher crash relaunch reuses the same
`launch_id` but the marker is already gone, so autonomous `must_play`
recovery is untouched.
- **fix(tr_treedb): refresh the pkey2 secondary index on a runtime
`treedb_save_node()`.** The secondary `pkey2` index kept objects
SEPARATE from the primary `id` index, populated only while loading
from disk (`load_pkey2_callback` gated on `sf_loading_from_disk`).
At runtime `treedb_update_node()` mutated the primary node in place
and `treedb_save_node()` only appended a tranger row — neither
touched the secondary index, so `treedb_get_instance()` /
`treedb_list_instances()` returned the OLD content after an update.
Surfaced as the agent's `list-binaries` showing the previous binary
right after a successful `update-binary` (the agent returns the
in-memory pkey2 index; the new record was already on disk). Now
`treedb_save_node()` re-points every pkey2 slot of the node at the
node object itself. No-op for topics without pkey2s. New regression
test `tests/c/tr_treedb_update_instance` (create → reload from disk →
update → assert via get_instance/list_instances); it fails against
the pre-fix code.
- **fix(emailsender): correctness + resilience of the SMTP send path.**
(1) Duplicate `MAIL FROM` → "503 MAIL already given": on AUTH-OK the
session published EV_ON_OPEN (whose subscriber already begins the
queued send, being idle) and then began it again; snapshot the
pending flag before publishing. (2) Permanent 5xx rejections now
dead-letter immediately (reply code forwarded via EV_ON_CLOSE /
EV_ON_MESSAGE; `code>=500` = permanent, 4xx/timeout/drop = transient
retry). (3) Binary (non-UTF-8) bodies persisted base64 under
`body_base64` instead of being silently dropped by `json_stringn`.
(4) Reconnection is owned by `c_smtp_session`, not the sender: after a
`timeout_inactivity` idle-close, `c_emailsender` just dispatches the head
message (it no longer carries any reconnect timer/backoff), and
`c_smtp_session` — which must redo the handshake — reconnects its bottom
`C_TCP` on `EV_SEND_MESSAGE` in `ST_DISCONNECTED` and sends on reaching
`ST_IDLE`. A handshake failure (AUTH 535, EHLO/banner 5xx) is transient
for the in-flight message (the server never saw it); only a 5xx in the
message's own MAIL/RCPT/DATA transaction dead-letters it. (5) Fixed a
shutdown SIGSEGV at its root: `tira_dela_cola()` now returns early when
the yuno is not playing (`gobj_pause()` clears the playing flag before
`mt_pause()`/`close_queues()`), so a deferred EV_ON_CLOSE delivered
during shutdown no longer touches the closed queue — no defensive NULL
check needed.
- **fix(c_tcp): retry with backoff after a failed reconnect in the
inactivity model.** `set_disconnected()` always cleared the timer in the
`timeout_inactivity` model — correct for a deliberate idle-close, but it
also stalled a failed on-demand reconnect (no retry, and no
EV_DISCONNECTED for a never-connected socket). Now only an idle-close
(`ac_timeout_inactivity` sets `idle_closed`) skips the retry; a connect
failure / dropped link schedules the `timeout_between_connections`
backoff and retries via `EV_TIMEOUT -> ac_connect`, like the classic
model. This is the layer that owns reconnection/backoff (the emailsender
rework relies on it).
- **fix(c_tcp): keep the pending tx queue across a FAILED reconnect in the
inactivity model.** `set_disconnected()` flushed `dl_tx` on every
disconnect, so bytes queued while disconnected (to be sent on the
on-demand reconnect) were lost if the connect failed before succeeding —
the message vanished silently and was never delivered. Now the queue is
kept when the connection was NEVER established (`inform_disconnection`
still FALSE) in the `timeout_inactivity` model on a running gobj, and
`start_pending_writes()` flushes it once a retry connects. An established
connection still flushes (its byte stream is broken); `mt_stop()` still
flushes unconditionally (no leak on stop). New regression test
`tests/c/c_tcp_inactivity` test4 (queue while the server is down → fail
retries → server up → echo confirms delivery); it fails against the
pre-fix code (no echo, FIFO timeout).
- **refactor(c_tcp): `timeout_inactivity` / `timeout_between_connections`
/ `rx_buffer_size` are deployment config (`SDF_RD`), not runtime
knobs** — dropped `SDF_WR` (and the misleading `SDF_PERSIST` on the
two timeouts); widened `priv->timeout_inactivity` to `json_int_t`.
- **fix(yev_loop): retry the static resolver's UDP `recv()` on `EINTR`,
and log `gai_strerror(ret)` not `strerror(errno)`.** A signal
interrupting the blocking DNS `recv()` made `yuneta_getaddrinfo()`
fail spuriously (the logged "Interrupted system call" was a stale
residual errno; getaddrinfo-family return an `EAI_*` code). Both
`getaddrinfo() FAILED` sites now log the real `gai_*` cause.
- **fix(ytls): send SNI in OpenSSL client handshakes.** The
OpenSSL backend never set the TLS `server_name` extension —
the code was a `// TODO SSL_set_tlsext_host_name` stub — so
client ClientHellos went out without SNI. Virtual-hosted TLS
endpoints behind a CDN/WAF (e.g. an Imperva Incapsula front
end) reject SNI-less handshakes with HTTP 403. The mbedTLS
backend already set SNI and `c_tcp` already supplies
`ssl_server_name`; only the OpenSSL path dropped it. Now
stores `ssl_server_name` in `init()` and calls
`SSL_set_tlsext_host_name()` per-connection in
`new_secure_filter()` for client sockets. Server side is
unaffected (no servername callback registered, so incoming
SNI is ignored). Verified end-to-end: 403 → 200 against an
Imperva-fronted HTTPS API.
- **fix(timeranger2): fire key-delete callbacks for `rt_by_disk`
followers.** `fire_key_deleted_locally()` skipped every entry
with an `fs_event_client` (i.e. every `rt_by_disk` follower),
assuming the `FS_SUBDIR_DELETED` inotify branch fired their
`key_deleted_callback`. But that branch's only firing
mechanism *was* `fire_key_deleted_locally()`, which skipped
them — so a `rt_by_disk` follower's key-delete callback never
fired on a `tranger2_delete_key()`. The follower's in-memory
state only reconciled on restart (LOADING reload from
`keys/`); live deletes were silently dropped for every
fs-watcher follower framework-wide. Split the fan-out by
transport via a new `fs_followers` flag: the master in-process
path fires only non-watcher subscribers (rt_mem
lists/iterators); the `FS_SUBDIR_DELETED` inotify branch fires
only the `rt_disk` followers (the inotify event IS their
signal). Each subscriber now fires exactly once; also removes
a latent same-process double-fire of non-watcher subscribers.
Verified: a master `tranger2_delete_key` now drops the key
from a separate-process follower's in-memory cache live, no
restart.
- **fix(yuno_agent): promote highest `yuno_release` to primary on
`restart_nodes`.** The treedb primary for a yuno-id is the
highest-ROWID record, not the highest `yuno_release`:
lifecycle writes (kill/run/snap) append records for whatever
release is *active*, so after `install-binary` +
`find-new-yunos` an older release could stay primary and
`deactivate-snap` relaunched it instead of the new version
(the long-standing "force volatil" TODO; a `shoot-snap`
between `find-new-yunos` and `deactivate-snap` reliably
triggered it). New `promote_highest_release_yunos()` runs in
`restart_nodes()` BEFORE the treedb reload: for each id whose
highest non-disabled `yuno_release` is newer than the current
primary, it re-appends that release so it becomes the highest
rowid; the reload then makes it primary and
`run_enabled_yunos()` launches it. An append does NOT move the
in-memory primary index — only the reload rebuilds it — so the
promote must precede the `gobj_stop/start`. Version order via
the existing `get_n_v()`. (`volatil` itself was already
honored — `mt_update_node` routes volatil updates to
`set_volatil_values`, in-memory only — so the culprit was the
non-volatil lifecycle/snap writes, not the run-update.)
Verified: a multi-yuno realm upgraded to a new release with a
single `deactivate-snap`.
- **fix(yev_loop): retry transient ENOMEM in
`io_uring_queue_init_params`**. A synchronised restart of
many yunos (e.g. an agent `deactivate-snap` on a node with
10+ yunos) used to drop 1-3 SIGABRT cores per yuno in
`/var/crash`, even though every yuno eventually came up
after the `ydaemon` watcher relaunched it. Root cause was
`yev_loop_create` aborting via `LOG_OPT_ABORT` on a
transient `-ENOMEM` from io_uring init: rings consume
pinned kernel memory (RLIMIT_MEMLOCK / vm.max_user_locks)
and a simultaneous restart of N yunos saturates that
budget for a few ms while the previous rings' pages are
released. Forensic evidence: 12 cores at 13:23 today with
identical bt bottoming at `yev_loop.c:184`, all `err=-12`
with `entries=32768`. Fix wraps the init call in a
5-iteration exponential-backoff retry (100/200/400/800/1600
ms ≈ 3 s) for `ENOMEM`/`EAGAIN` only; non-transient errors
(EINVAL, ENOSYS, EPERM…) fall through to the original
abort path unchanged. Each retry logs a warning so an
operator can see the pressure event without it being
silent. Local stress test (3 consecutive `deactivate-snap`
cycles = 48 yuno restarts) generated 0 cores.
- **fix(ycommand): keep stdin-pipe queue draining when a
command returns -1**. The long-lived stdin-pipe mode added
in 7.4.3 inherited the ybatch convention from `-c` / `-i` /
file-fed batches: a `-1` result with no leading `-` on the
command drops the rest of the queue. That convention is
hostile to stdin-pipe deploys — the operator has already
piped every line in, and one common non-fatal `-1` (e.g.
`install-binary` returning "Binary already exists" for a
slot that's already filled) silently swallows the rest.
Surfaced on the 7.4.3 wattyzer deploy: binaries got
registered, then `find-new-yunos` + `deactivate-snap`
vanished, leaving yunos on the old release until the
trailing commands were re-run by hand. Fix: in
`stdin_pipe_mode` every command behaves as `ignore-fail`
(no queue clear on error). Explicit `-` prefix path stays.
Repro matrix (4 install-binary in pipe, slot pre-filled):
before fix 3/4 responses, after fix 4/4. The earlier
"WS frame interleaving" hypothesis logged in TODO.md was
ruled out (kept as a post-mortem trail).
- **fix(emailsender): retry queued emails instead of
dead-lettering on the first failure, and persist the body**.
Any send failure (SMTP server down, wrong URL, rejected
AUTH, or simply the SMTP child still connecting when the
dequeue timer fired) used to move the email straight to the
`emails_failed` dead-letter queue and unload it — `max_retries`
was declared but never used and nothing drains the failed
queue, so one transient hiccup shelved the message forever.
Now the message stays at the head of `emails_queue` and is
only dispatched while the SMTP session is connected and
authenticated (a momentary outage just waits and retries on
reconnect); transient failures are retried up to `max_retries`
total attempts before being dead-lettered. The body is now
persisted as a string in the queue — it was carried as a
transient gbuffer pointer that the dequeued kw's auto-decref
freed after the first attempt, so retries (and any yuno
restart) lost the body. Also split RCPT recipients on `;` as
well as `,` (an Outlook-style list or a stray trailing `;`,
e.g. logcenter's summary `to`, was rejected by the server as
`501 Invalid TO`). Deployed and validated live on
`emailsender^artgins`.
- **feat(emu_device): implement the frame-emission path on
timeranger2**. The device-gate emulator was a scaffold: its
replay was written against the removed timeranger v1 API and
its `__output_side__` had no TCP connex (the v6 Connex/Tcp0
globals were dead). Ported to v7 — the output side is built
in code (`C_IOGATE > C_CHANNEL > C_PROT_RAW > C_TCP` to `url`,
like `sgateway`); `mt_play` loads matching `frame64` records
via `tranger2_open_list`, and on connect it sends the
`leading` frame then `window` frames every `interval` ms.
Compile-verified only; end-to-end runtime validation (needs a
`frame64` topic + a TCP sink) is tracked in `TODO.md`.v7.4.3 -- 27/May/2026¶
- **feat(emailsender)!: drop libcurl, native SMTP over ytls**.
`emailsender` was the only yuno that linked libcurl, which
dragged OpenSSL/libssh2/c-ares/libidn2/libpsl/libnghttp2/3/
zlib/brotli into its runtime graph. The dev-host glibc kept
bumping while production stayed on older versions (e.g. 2.36
on `app.wattyzer.com`), so emailsender was the one yuno where
every upgrade needed a build environment matched to the
target — and it was deliberately skipped from the 7.4.1
deploy bundle for that reason. Three new building blocks
land in this release: (1) `C_SMTP_SESSION` (yunos/c/
emailsender/src/c_smtp_session.{c,h}) — a CHILD-pattern
protocol gclass that owns a `C_TCP` bottom and walks the
RFC 5321 submission FSM (banner → EHLO → AUTH PLAIN →
MAIL FROM → RCPT TO → DATA → \r\n.\r\n → QUIT) with
multi-recipient support, RFC 5321 §4.5.2 dot-stuffing and
a best-effort QUIT on `mt_stop`. Uses `istream_read_until_
delimiter("\r\n", 2, EV_RX_LINE)` so raw bytes from C_TCP
(`EV_RX_DATA`) and parsed lines (`EV_RX_LINE`) flow through
distinct actions and never feed back into the istream.
(2) `mime_encoder.{c,h}` — pure helpers, no gclass; builds a
complete RFC 5322 message with optional single attachment
(`multipart/mixed`) or inline-image attachment with
Content-ID (`multipart/related`), base64-wrapped at 76
chars per RFC 2045 §6.8, and RFC 2047 base64 encoded-words
for non-ASCII Subject / From display-name. (3) Cutover in
`c_emailsender.c`: `priv->curl` → `priv->smtp` as a
pure_child created with `{url, username, password,
helo_name=gethostname()}`; the synchronous `gobj_send_event
(priv->curl, EV_CURL_COMMAND, …)` + immediate
`process_curl_response` flow becomes async — `ac_smtp_
command` MIME-encodes, sends `EV_SEND_MESSAGE` to the smtp
child and stays in `ST_WAIT_RESPONSE`; the response arrives
later via the new `ac_on_message` handler. Kernel-side
precursor: `_yev_protocol_fill_hints()` in `kernel/c/
yev_loop/src/yev_loop.c` learns the `smtps` schema (port
465, marked `secure=TRUE`) alongside the existing
`mqtts`/`wss` — strictly additive. Together this removes
`libcurl4-openssl-dev` from `docs/doc.yuneta.io/
installation.md`, drops `find_package(CURL REQUIRED)` and
`${CURL_LIBRARIES}` from the yuno's CMakeLists, deletes
`c_curl.{c,h}` outright, and brings the emailsender binary
down to `ldd` reporting only `libgcc_s` + `libc` (vs the
previous ~12 shared libs). `emailsender^artgins` is already
pointed at `smtps://ssl0.ovh.net:465` in all three realms +
the staging batch, so the cutover is a no-config-change
deploy. **Breaking change for callers building EV_SEND_EMAIL
kw manually**: the libcurl-era attrs `strict_tls` and
`auto_inline_images` are ignored (TLS is now decided by the
URL schema; auto-inline-image HTML rewriting was a libcurl
`curl_mime_*` feature not reimplemented). The `cmd_send_
email` command schema is unchanged; existing callers (the
whole estadodelaire batch + every realm config) keep
working without edits.
- **feat(ycommand): long-lived stdin-pipe session keeps OAuth2
auth open across many commands**. Until now `ycommand` had
three input shapes: `-c CMD` (one command, exit), `-i`
(interactive editline over a raw TTY), and an undocumented
synchronous pipe path inside `ac_on_open` that did
`while(fgets(line, stdin))` — fine for pre-buffered batches
but it blocked the yev_loop between lines, so any
programmatic driver that wanted to send a command, read its
response, then send another would hang. `-i` was also
unusable for non-TTY drivers (e.g. an AI coding agent running
`ycommand` via Bash) because `tty_keyboard_init`
unconditionally calls `enableRawMode`, which fails with
"NOT a TTY" on a piped fd. Net effect: every remote command
paid the full OAuth2 ROPC round-trip (~200-400 ms against
Keycloak) even when the caller had ten queued up. New
behavior, no new flag: when stdin is not a TTY and neither
`-c` nor `-i` was supplied, ycommand sets up an io_uring read
event on `dup(STDIN_FILENO)` (yev_loop refuses fd<=0, so we
dup) and drives lines through the existing
`split_commands_into_queue` + `run_next_pending` machinery.
The process stays alive between commands, EOF triggers an
orderly shutdown via the existing `set_timeout(timer,
wait*1000)` → `ac_timeout` → `exit()` path, and a
`priv->cmd_in_flight` gate serialises async dispatches so a
stdin line arriving mid-flight enqueues instead of racing.
Auth happens once; the rest of the session is free. Tested
against local `ws://127.0.0.1:1991` (4-line batches and
delayed sequences with 3 s gaps between lines) and against
`wss://app.wattyzer.com:1993` + OAuth2 (three commands with
2 s gaps, single ROPC). Backwards-compatible: the pre-existing
`-c` and `-i` paths are untouched, the pipe-mode trigger is a
strict superset of the previous synchronous fgets behaviour,
and `echo cmd | ycommand` keeps producing the same output as
before (just via an event-driven reader). `c_ycommand.c`
grew +220/-25; no changes elsewhere.
- **docs(philosophy): add "The Typed-Graph Model" chapter**.
New page under `philosophy/` slotted between
[Design Principles](doc.yuneta.io/philosophy/design_principles.md)
and [Domain Model](doc.yuneta.io/philosophy/domain_model.md),
articulating the conceptual claim the framework rests on: data
and behavior are two views of the same typed graph
(`topic`↔`gclass`, `node`↔`gobj`, `hook`/`fkey`↔
subscription/`bottom_gobj`, with `sdata_desc_t` describing
schemas on both planes). Sections cover: the unit is the
typed binding, not just the node; the two-plane primitive
table; what kinds of organisation the model can express
(hierarchies, matrix, workflows, communication topologies,
versioned-over-time); what does not fit cleanly (schemaless
iteration, OLAP, eventually-consistent distributed state,
truly opaque payloads); the implicit axiom; the payoffs at
scale; and the empirical justification from 15 + years of
v2/v6 production. Cross-links added in `philosophy/`
neighbours and reverse-links from `yunos/c/yuno_agent/`'s
`ENTRY_POINT.md` (See also), `GOBJ.md` (Conceptual frame
callout: behavior plane) and `YUNO_TREEDB.md` (Conceptual
frame callout: information plane) so a reader landing in the
technical chapters can step up one level on demand. Build is
warning-free; the new chapter appears in the TOC under
Philosophy.
- **fix(install-binary): surface the real cause in error
response**. `cmd_install_binary` built its failure comment as
`json_sprintf("Cannot create binary: %s",
gobj_log_last_message())`, which produced *"Cannot create
binary: "* (empty cause + trailing space) whenever the
underlying `treedb_create_node` returned NULL because the
`(id, pkey2=version)` combination already existed — that path
logs *"Node already exists"* via `gobj_log_warning`, and
`gobj_log_warning` does not populate `last_message` (only
`LOG_ERR` and above do). After the per-command reset in
`command_parser` (7.4.1, `b1abd7f69`), the buffer is `""` by
then. Two-layer fix, same shape as the snap commands in 7.4.1:
`treedb_create_node` now calls `gobj_log_set_last_message()`
alongside the warning so the cause (*"Node already exists in
'<topic>': id='<id>'"*) reaches every caller that pipes
`gobj_log_last_message()` into the response (≈13 callers in
`c_node.c` benefit alongside `cmd_install_binary`);
`cmd_install_binary` reads `last_msg` once and falls back to
`"(see log)"` if it's empty, so the response is always
informative regardless of whether layer-1 was reached. Drive-by:
removed the stale `// TODO check tranger2_write_user_flag`
marker above `treedb_shoot_snap` — the function was completed
in 7.4.0/7.4.1 (`4c89e4b2c` + `46f8f0434`) and the audit
confirmed no remaining wiring gap.v7.4.1 -- 27/May/2026¶
- **fix(command_parser): stop misleading stale strerror in
command responses**. Many `cmd_*` in `c_node.c` build their
failure comment as `json_string(gobj_log_last_message())`,
but `gobj_log_set_last_message()` is only called by
`gobj_log_*()` with priority `<= LOG_ERR`. If the failure
path logs at `LOG_INFO` (or doesn't log at all),
`last_message` keeps whatever it had from the previous
`LOG_ERR` — frequently `strerror(errno)` of an earlier TCP
disconnect ("Connection reset by peer"). The response was
being delivered correctly with that strerror as the comment;
ycommand rendered it verbatim and it looked indistinguishable
from a real network error, sending operators down a wild
diagnostic chase. Two-part fix at the kernel command
boundary in `kernel/c/gobj-c/src/command_parser.c`: (1)
`command_parser()` resets `last_message` to `""` at entry,
so every command dispatch starts with a clean slate;
(2) `build_command_response()` substitutes `"(see log)"`
when the response is a failure (`result != 0`) and the
comment is an empty string, so callers that use the bare
`json_string(gobj_log_last_message())` idiom produce a
useful placeholder instead of `ERROR -1: `. Success
responses keep their empty comment (cmd_topics etc. return
data without a comment by design). Documented the new
semantics in
`docs/doc.yuneta.io/api/logging/log.md`
(`gobj_log_last_message` + `gobj_log_set_last_message`).
- **fix(agent): `list-binaries` enumerates every
`(role, version)` instance**. `cmd_list_binaries` called
`gobj_list_nodes("binaries", ...)`, which only returns the
in-memory primary per id (role). After an `install-binary`
that added a second version under the same role, the new
instance was invisible until a `deactivate-snap` rebuilt
the primary index — and even then only the most-recent
version survived. The doc claimed *"returns all rows"* but
the call had never matched that promise. Switched to
`gobj_list_instances("binaries", "", ...)`: the topic has
`pkey2=version`, so the instances iterator returns one row
per `(role, version)` and multi-version installs are
visible from the moment `install-binary` appends the
record. Validated with two coexisting `emailsender`
versions (7.4.1 + 7.4.2): `list-binaries` now returns four
rows instead of three. `list-binaries-instances` stays as
the explicit "instances" alias. `YUNO_LIFECYCLE.md` table
updated to match.
- **fix(snap): implement `snap-content` + recover error
responses to ycommand**. (a) `cmd_snap_content` in
`c_node.c` was a literal `"TODO"` stub. Now walks the
requested topic with `tranger2_open_list` filtered by
`user_flag=snap_id` and a `load_record_callback` that
drains matching records into a `json_array` carried via
the rt's `extra` (merged into the rt object by
`json_object_update_missing_new`, so the callback reads
`list->snap_data`, not `list->extra->snap_data`).
Validates `topic_name` and `snap_id` (1..65534, matching
the `uint16_t` md2 `user_flag` range), returns the schema
+ the array and a `(count)` comment. (b) Error responses
from `shoot-snap` / `activate-snap` / `deactivate-snap`
were arriving at ycommand as `"Connection reset by peer"`
instead of the real cause. Root cause:
`json_string(gobj_log_last_message())` produced an empty
JSON string whenever the kernel function logged with
`gobj_log_info` (which doesn't populate `last_message`),
and the buffer still held the strerror of a prior
disconnect. Two-layer fix: `treedb_shoot_snap` /
`treedb_activate_snap` explicitly call
`gobj_log_set_last_message()` in their "already exists"
and "not found" paths; `cmd_shoot_snap` /
`cmd_activate_snap` / `cmd_deactivate_snap` build the
error comment with
`json_sprintf("Cannot ... '%s': %s", name, empty_string(last)?"(see log)":last)`
so the comment is never empty even if some future caller
forgets the layer-1 update. Note: there are ~13 other
`cmd_*` in `c_node.c` with the same
`json_string(gobj_log_last_message())` pattern — same trap
— covered by the systemic `command_parser` reset documented
in the entry above.
- **fix(snap): preserve previous snap's tag + harden
`run_yuno` launcher**. (a) `treedb_shoot_snap` stamped
`user_flag` IN PLACE on every primary record. A second
`shoot-snap` over a record that hadn't changed since the
first snap therefore overwrote the older snap's tag,
making that record unreachable from
`activate-snap(older)`. Fix: if the primary record already
carries a tag from a different snap, append a CLONE via
`tranger2_append_record(user_flag=new)` so the older
record keeps its tag. Untagged primaries (and same-snap
re-stamps) still take the in-place path — no extra
storage. The in-memory `__md_treedb__` is intentionally
left untouched on the clone branch so subsequent
`treedb_save_node()` appends keep using the original base
`rowid`. Validated end-to-end: shoot A → 13 records carry
`uflag=1`; shoot B (no intermediate change) → still 13
records carry `uflag=1` + 13 clones carry `uflag=2`;
`activate-snap A` and `activate-snap B` both restore all
yunos. (b) `build_yuno_running_script` in `c_agent.c` took
an uninitialised `char bfbinary[]` from the caller's stack
and could early-return `0` (silently) on a missing realm
or binary. All three callers ignored the return value,
then ran or serialised whatever garbage was on the stack —
hence the corrupted `/yuneta/bin/<role>^<name>.sh`
launchers (~21-27 bytes of stack noise) that
`activate-snap` produced when the binary record didn't come
back. Fix: zero-init `bfbinary` at function entry, log the
two early-return sites (no silent errors), and have every
caller check the return value and respond with an error
instead of piping uninitialised memory into
`run_process2()` or a JSON reply. Treedb `treedb_shoot_snap`
doc page updated to describe the new clone-vs-stamp
behaviour.
- **fix(timeranger2): silence two `-W` warnings without
losing errors**. (a) `treedb_shoot_snap` in
`kernel/c/timeranger2/src/tr_treedb.c` now returns the
accumulated `ret` so `tranger2_write_user_flag` failures
across topics surface to callers instead of being silently
dropped. (b) `mirror_key_delete_to_disks` in
`timeranger2.c` uses `build_path()` to assemble the
`disks/<rt_id>/<key>` path; `build_path` already syslogs
`LOG_CRIT` on overflow, so no silent skip on truncation
(closes a `-Wformat-truncation` warning without
introducing a silent early-out).
- **docs(api/treedb): document the real snap semantics**.
The `treedb_shoot_snap` / `treedb_activate_snap` API pages
described the surface only — name, parameters, return — and
missed the parts that determine whether snaps actually work
for the caller: `treedb_shoot_snap` tags the live `.md2`
record's `user_flag` in place via
`tranger2_write_user_flag` (so rowid order is preserved
and re-shoots overwrite prior tags on the same record);
`treedb_activate_snap("__clear__")` is the deactivate
path; the active/inactive toggle only flips a flag, and
the new primary visibility materialises on the **next**
`treedb_open_db()` — not on the call itself. Updated
`docs/doc.yuneta.io/api/timeranger2/treedb.md` §§
`treedb_shoot_snap` + `treedb_activate_snap` with the
reload semantics + the in-place tag mechanic + the
16-bit snap-id ceiling. Companion to the `treedb_shoot_snap`
completion shipped in the same release.
- **feat(tr_treedb): complete `treedb_shoot_snap` so
`activate-snap <name>` rolls back primaries correctly**.
The TODO at the heart of `treedb_shoot_snap` was a dead
branch: it walked the primary index of every topic but
the actual `tranger2_write_user_flag` call was commented
out, so snaps only ever created an entry in `__snaps__`
and never tagged any record. The agent's `activate-snap
<name>` path queried `snap_tag` correctly on reload (see
`treedb_open_db` line 1299 — the user_flag filter is
enforced), but with no record carrying the tag the load
returned an empty primary index, and the rollback silently
did nothing — the test suite never caught it because there
was none. Wired the tag write in-place via
`tranger2_write_user_flag(tranger, topic_name, key, t,
i_rowid, user_flag)` so the existing record gets stamped
without inflating the `.md2` (using `treedb_save_node`
would create a new instance at the highest rowid and
then steal "latest" after `deactivate-snap`, masking the
newer records the user actually wants live). Also tightened
the snap-id range check from 32-bit (`0xFFFFFFFF`) to 16-bit
(`0xFFFF`) since `user_flag` is `uint16_t` end-to-end. New
regression: `tests/c/tr_treedb_snap` walks 9 phases
modelling the agent's upgrade lifecycle (seed v1 → add v2
→ shoot snap_v1 → deactivate → reload picks v2 → add v3 →
deactivate → reload picks v3 → shoot snap_v3 → activate
snap_v1 → reload rolls back to v1 → activate snap_v3 →
reload to v3 → deactivate → stays at v3) on two topics
(`binaries` + `yunos`) keyed exactly like the agent's
`binaries` / `configurations` / `yunos`. `tr_treedb` and
`tr_treedb_delete_instance` rerun green against the patched
library — no regressions.
- **docs(yuno_agent): document the version-bump upgrade flow**.
`kill-yuno` + `run-yuno` does not pick up a new release —
`cmd_run_yuno` walks the `yunos` topic primary index, which
keeps pointing at the older `pkey2` (`yuno_release`) after
`find-new-yunos create=1` appends the new row. The fix is
always `install-binary` → `find-new-yunos create=1` →
`deactivate-snap`. `deactivate-snap` with no args (and no
active snap) is the only supported way to trigger
`restart_nodes()` (`c_agent.c:8816`), which SIGKILLs every
running yuno, `gobj_stop/start`s the treedb resource so the
primary index is rebuilt from disk with the newest `pkey2`
first, then runs every must-play yuno. Equivalent to
`yshutdown` + `restart-yuneta` at the agent-process level,
but without restarting the daemon. Added as `YUNO_LIFECYCLE.md`
§6.5 + §6.6 (rollback via `shoot-snap` / `activate-snap`) and
refreshed `CLAUDE.md` §3 with the same-version vs version-bump
split. Verified live with `gate_pvpc 1.3.1.0 → 1.3.1.1`
on the local agent.v7.4.0 -- 26/May/2026¶
- **chore(lib-yui)!: declarative shell stack removed —
`@yuneta/lib-yui` jumps to 8.0.0**. The new declarative
shell (`C_YUI_SHELL`, `C_YUI_NAV`, `C_YUI_PAGER`,
`C_YUI_WIZARD`, `shell_modals` and every `shell_*_helpers`
module + the full Playwright e2e suite and the `test-app/`
vite project) was already being maintained from the
wattyzer-vendored copy at `wattyzer/gui/src/lib-yui`. The
kernel copy was just dead weight in the bundle consumed
by estadodelaire (legacy apps that only use
`C_YUI_MAIN` + `WINDOW` + `TABS` + routing) and the two
copies had drifted enough to make every fold-back a
conflict. Verified by grep that neither legacy consumer
imports any shell symbol before deleting. `dist/lib-yui.es.js`
is now 3.4 MB / gzip 706 KB (~25% lighter). Migration: if
you need the declarative shell, the canonical copy is
`wattyzer/gui/src/lib-yui/` (private repo) as of
2026-05-15. The yuno-skeleton `js_gui` scaffold that
referenced `register_c_yui_shell` was dropped here too;
the replacement scaffold lives in
`wattyzer/templates/js_gui/`.
- **chore(ext-libs): three security bumps (v1.12 → v1.13 →
v1.14)**. v1.12: nginx 1.28.3 → 1.30.1 (CVE-2026-42945),
openresty → 1.29.2.4, openssl 3.6.1 → 3.6.2. v1.13: nginx
→ 1.30.2 (CVE-2026-9256, buffer overflow in
`ngx_http_rewrite_module`). v1.14: openresty → 1.29.2.5
(backports the CVE-2026-9256 patch into the
openresty-bundled nginx + a `proxy_protocol v2`
over-read fix). All three are pin-only — nginx and
openresty are separate dynamically-linked binaries (see
`configure-libs.sh` v1.10), so no yuneta consumer /
header / CMake change rides along. OpenSSL deliberately
held on the 3.6 LTS series; the 4.0 jump (non-LTS,
EOL 2027-05, drops engines / legacy init) is tracked
separately.
- **refactor(tranger2): rename `tranger2_delete_record` →
`tranger2_delete_key`**. Locks the vocabulary
timeranger2 was using loosely: *record* = a primary key
(whole `keys/<key>/` directory, deleted via
`tranger2_delete_key`); *instance* = one row of that
key's `.md2` index, addressed by
`(key, __t__, rowid)`. The legacy name is kept as a
source-level alias
(`#define tranger2_delete_record tranger2_delete_key`),
so external callers keep compiling unchanged. In-tree
caller (`treedb_delete_node`) updated; README, the
timeranger2 API page (with the old MyST anchor preserved
so external links to `(tranger2_delete_record)=` keep
resolving) and the appendix index follow the rename. A
subsequent commit dropped "soft" from the delete-instance
vocabulary: granularity, not reversibility — both
deletes are irrecoverable.
- **feat(tranger2): `tranger2_delete_instance()` — per-row
tombstone**. Mutates one row of the `.md2` index in place via
`sf_deleted_instance = 0x0400` (reinstated in `system_flag2_t`
on the inherited side of the mask, so `rt_by_disk` followers
see the same tombstone as the master). Optional `zero_payload`
overwrites the matching `__size__` bytes at `__offset__` in the
data `.json` for sensitive-data wipes. Three read sites honour
the bit and skip dead rows: `tranger2_open_iterator` history
loop, `tranger2_iterator_get_page`, and
`publish_new_rt_disk_records`. Treedb is downstream and
inherits the skip with no `tr_treedb` change. Master-only.
Second delete of the same row is a silent no-op. `rowid`s do
NOT renumber; `iterator_size` / `total_rows` keep counting
slots, not live rows. `tranger2_read_record_content` and
`tranger2_read_user_flag` still serve dead rows when the caller
addresses them directly (audit / wipe-verification tooling).
Coverage: `tests/c/timeranger2/test_delete_instance.c` (5
sub-cases).
- **feat(tranger2): `tranger2_delete_key()` propagates to
subscribers**. Pre-2026-05-26 the function `rmrdir`'d
`keys/<key>/` and cleared the in-memory rollup cache, but
never notified subscribers — `rt_mem` listeners kept stale
references and `rt_by_disk` followers in other processes kept
their cached view alive. Now: (1) `topic/disks/<rt_id>/<key>/`
subdirectories are removed BEFORE the live `keys/<key>/`, so
followers catch the deletion on the standard inotify channel
(`FS_SUBDIR_DELETED_TYPE`, the v7 TODO branch that had been
logging "NOT processed" since inception is now wired); (2)
in-process subscribers receive a registered
`tranger2_key_deleted_callback_t` via the new
`tranger2_set_rt_key_deleted_callback()` setter. Additive
typedef and setter — no breaking signature change to
`open_rt_mem` / `open_rt_disk` / `open_iterator`. Coverage:
`tests/c/timeranger2/test_delete_key_propagation.c` (5
sub-cases). Wattyzer's "tombstone-then-delete" workaround in
`db_history_wz` becomes redundant after one production cycle
of coexistence — cleanup planned in the wattyzer repo, not
here.
- **fix(tr_treedb): repair `treedb_delete_instance`
(pkey2-index cleanup only)**. The function was a dead
copy-paste of `treedb_delete_node` with the real work
fenced behind `if(0) { ... tranger2_delete_instance(...) }`
and an `else` returning -1 with *"Cannot delete node"*. The
dead branch was also wrong — it would have wiped the
underlying `.md2` row while the primary index and the
other `pkey2_*` indexes still referenced it (state
corruption). The only in-tree caller
(`c_node.c::ac_delete_node`) was getting -1 on every call
without acting on the return, so the breakage was silent.
Rewritten to do what the function name promises: drop the
in-memory entry for THIS pkey2 via `delete_secondary_node`,
fire `EV_TREEDB_NODE_DELETED`, preserve the JSON_INCREF /
DECREF pattern. The whole-node wipe stays the job of
`treedb_delete_node` → `tranger2_delete_key`. Contract
spelled out in the `.c` and `.h` docstrings.
- **fix(ytls/openssl): ship the full certificate chain**.
`build_ssl_ctx()` was loading the server certificate via
`SSL_CTX_use_certificate_file()`, which only parses the first
cert of a PEM bundle. With a Let's Encrypt fullchain.pem on
disk, that meant the listener served only the leaf — browsers
hid the issue via AIA-fetch / cached intermediates, but
strict-TLS clients (e.g. Node's native `fetch`, used by the
Playwright QA driver against the public URL) failed chain
verification. Switched to
`SSL_CTX_use_certificate_chain_file()` (chain-aware, PEM-only
— no `SSL_FILETYPE_PEM` arg). The mbedTLS backend
(`mbedtls_x509_crt_parse_file`) was always chain-aware so it
was not affected. Every yuno that exposes a TLS server with
the OpenSSL backend needs a relink + redeploy to pick up the
fix; for binaries shared by several live yunos (e.g.
`auth_bff` 1802+1804) the atomic `mv old old.bak; cp new old`
pattern avoids the `ETXTBSY` that breaks `update-binary`.
- **chore(ytls, c_authz): drop OpenSSL legacy init/cleanup
calls** (OpenSSL 4.0 prep). Removed four
deprecated-since-1.1.0 calls that were no-ops on the 3.x
series and disappear in 4.0: `SSL_library_init()` +
`OpenSSL_add_all_algorithms()` (with the redundant
`__initialized__` guard) in `ytls/openssl.c` init,
`EVP_cleanup()` in cleanup, and
`OpenSSL_add_all_digests()` (with its
`CONFIG_HAVE_OPENSSL` wrapper) in `c_authz.c`. OpenSSL
≥ 1.1.0 auto-initialises on first use and cleans up via
`atexit`. The `OPENSSL_API_COMPAT 30100` define already
gates the rest of the 1.1.x compat surface; the yuneta
source is now 4.0-clean (the jump itself stays deferred).
- **fix(c_prot_tcp4h, c_prot_mqtt): guard state reset
against in-publish disconnect cascade**. Under io_uring,
publishing `EV_ON_MESSAGE` is synchronous and can trigger
a full disconnect cascade upstream (authz NAK in
`C_IEVENT_CLI` → `EV_DROP` → `C_TCP` ac_drop →
`try_to_stop_yevents` → `set_disconnected` publishes
`EV_DISCONNECTED` → the protocol gclass moves to
`ST_DISCONNECTED`). The caller of `frame_completed`
then unconditionally reset the FSM back to
`ST_WAIT_FRAME_HEADER` / `ST_CONNECTED`, leaving the
protocol "connected" without an underlying TCP. Symptom:
*"Event NOT DEFINED in state: EV_CONNECTED in
C_PROT_TCP4H@ST_WAIT_FRAME_HEADER"* alternating with
broken-pipe / local-dropping cycles. Guard
(`state != ST_DISCONNECTED`) added — same form already
present in `c_websocket.c` and `c_prot_mqtt2.c`. The
twin guard initially added to `c_prot_modbus_m` was
reverted in a follow-up: the modbus-master flow does
not expose the same cascade (see `GOBJ.md §8.13`).
Verified live on `app.wattyzer.com` across the
controlcenter dial-out loop.
- **fix(gobj): `gobj_read_attrs` honours `mt_reading` via a new
`item2json` helper**. The bulk reader (behind `view-attrs`,
introspection and `db_save_persistent_attrs`) was the only
attribute path bypassing `mt_reading`, so `SDF_RSTATS` counters
kept in `priv->X` read as zero through `view-attrs` even though
`stats-yuno` (typed readers) saw the live value. The helper
dispatches by `DTP_*` and falls back to the stored value when
`mt_reading` is absent or returns `!v.found`. `gobj_read_attr`
(single, borrowed-ref) is intentionally left alone: its
"Return is NOT yours!" contract is incompatible with allocating
a fresh `json_t` from a typed override.
- **fix(gobj-js): mirror C — `gobj_read_attrs` honours
`mt_reading`**. Same shape as the C kernel patch, simplified
(dynamic types, no `DTP_*` switch): `undefined` falls back to
the stored value. Defensive — no JS gclass implements
`mt_reading` today, so this only closes the symmetry.
- **fix(c_tcp): set `v.found = 1` for `cur_tx_queue` in
`mt_reading`**. Pre-existing one-liner; the branch updated
`v.v.i` but forgot the discriminant, so the value was always
shadowed by the stored zero. Surfaces now that
`gobj_read_attrs` consults `mt_reading`.
- **fix(lib-yui): normalize `navigator.language` before
`Intl.DateTimeFormat`**. Playwright Firefox without locale
config (and some embedded webviews) report
`navigator.language` as the literal string `"undefined"`;
passing that to `Intl.DateTimeFormat` throws `RangeError`
and breaks SPA bootstrap. Guard with a string +
`"undefined"` check, fall back to `undefined` so Intl
resolves to the system locale. Fold-back of wattyzer
`1bf08aa`.
- **feat(yuno_agent): `stats-yuno` defaults `service` to the
matched `yuno_role`**. `cmd_stats_yuno` previously passed
an empty `service` when the operator didn't spell it out,
so the remote fell back to `priv->gobj_service` (the top
`C_YUNO` instance) and returned only the few attrs declared
on `c_yuno` — typically all zero, missing the real
`SDF_RSTATS` counters that live on the citizen service
(e.g. `C_AUTOMATIONS_WZ`'s `alarms_seen` /
`tracks_seen` / `fires_seen` / `runs_done`). Convention is
`gobj_create_default_service(yuno_role, GCLASS, ...)` so
service name == yuno role; the operator now gets the real
counters with the natural `stats-yuno yuno_role=X`
invocation. Pass `service=__yuno__` to explicitly query the
top `C_YUNO` attrs (previous default).
- **fix(tr_treedb): include the required-field name in error
logs**. The three "Field required" `gobj_log_error` sites in
`check_desc_field` / `normalize_node_field_value` /
`convert_node2tranger` logged the same opaque message; now
the field name is interpolated into `msg` so the offending
column is visible at a glance without expanding the
structured payload.
- **chore(yuno_agent): increase agent log file size**. Bumps
the agent's own log rotation threshold in `yuno_agent` and
`yuno_agent22` (two-line `main.c` change). Avoids
tighter-than-needed rollovers under the trace volume the
onboarding doc work surfaced.
- **note**: validated by relinking every yuno (15 binaries) and
running the full `ctest` suite (93/93 passed, 436 s). A
project-wide `yunetas build` after a kernel-side change does
pick up the relink correctly — the `rm <yuno_bin> &&
make install` workaround is only needed when rebuilding a
single yuno's `build/` directory in isolation.v7.3.4 -- 16/May/2026¶
- **chore(release): corrective republish of `@yuneta/lib-yui`**.
`@yuneta/lib-yui@7.3.3` was published to npm from a branch that
had the 7.3.3 release prep (gobj-js dependency pin, CHANGELOG,
version) but **not** the G6 v5 canvas panning fix (PR #115):
the published 7.3.3 tarball is missing
`src/g6_drag_canvas_touch.js` and the `autoResize:false` /
`ensure_drag_canvas_patch` changes, so installing lib-yui from
npm still had the broken touch/desktop panning. npm versions
are immutable, so 7.3.4 republishes `@yuneta/lib-yui` from
`main` with the complete set (#115 + #116). No source changes
versus what `main` already contained at 7.3.3 — this is a
packaging correction only.
- **chore(release): `@yuneta/gobj-js` 7.3.3 → 7.3.4 (lockstep)**.
`@yuneta/gobj-js@7.3.3` on npm was already correct (it carries
the `createElement2` nullish-`data-i18n` guard). It is bumped to
7.3.4 with no functional change purely to keep the two JS
packages in lockstep and avoid version-skew confusion; the
`@yuneta/lib-yui` peer range moves to `^7.3.4`.
- **note**: deprecate the bad artifact —
`npm deprecate "@yuneta/lib-yui@7.3.3" "incomplete: missing the
G6 panning fix; use >=7.3.4"`.v7.3.3 -- 16/May/2026¶
- **fix(lib-yui): G6 v5 canvas panning (touch broken, desktop
desynced)** (PR #115). G6 v5.1.0 `drag-canvas` derived the pan
delta from `event.movement`, which `@antv/g` fills from native
`PointerEvent.movementX/Y`: left at 0 for touch pointers on most
mobile browsers (canvas barely panned) and skewed by OS pointer
acceleration / `devicePixelRatio` on desktop (graph lagged the
cursor). New `g6_drag_canvas_touch.js` subclasses `DragCanvas`,
reuses its clamp/cursor logic, and pans by the `event.viewport`
delta — reliable on mouse and touch, correct for any canvas
scale/zoom. Registered once over the built-in `'drag-canvas'`
id so every graph (gobj-tree, json-graph, treedb editor) is
fixed without per-consumer changes. `c_yui_gobj_tree_js.js`
and `c_yui_json_graph.js` also stop fighting G6 `autoResize`
(window-only in v5) and size the canvas to its content box via
a self-contained `ResizeObserver`.
- **fix(gobj-js): `createElement2` no longer poisons `data-i18n`
with `undefined`**. A nullish `i18n` attribute (e.g. a field
whose `header` is undefined) rendered the literal
`data-i18n="undefined"` and suppressed translation, leaving
form labels blank. The attribute is now skipped when the value
is `null`/`undefined`; an explicit empty string is still
honoured.
- **fix(lib-yui): pin `@yuneta/gobj-js` dependency**. The
peer dependency was `"*"`, and a stale lockfile had frozen it
to the ancient `@yuneta/gobj-js@0.3.0` from npm — so lib-yui
(and downstream apps) silently built against 0.3.0 and updates
had no effect. Peer range is now `^7.3.3` and a
`file:../gobj-js` devDependency makes local builds use the
in-tree source. Downstream apps (wattyzer, estadodelaire)
must likewise repin and reinstall so they stop
resolving 0.3.0.
- **feat(lib-yui): shell-mountable developer panel**.
`build_dev_panel` plus a new `C_YUI_GOBJ_TREE_JS` hierarchical
gobj-tree viewer; the dev panel is now a real window box (was a
floating transparent overlay), with silent optional
`__yui_main__` lookup and theme-aware styling.
- **feat(lib-yui): TreeDB graph node redesign**. Unified node
design system: doc-style HTML node cards (theme-aware), soft
topic palette, rectangular leaves, size tiers, ports back to
the topic colour, no circles, click-detail popover (no hover
tooltips), redesigned edges/ports, dark-theme contrast fix.
- **feat(lib-yui): graph toolbar / context-menu**. All
toolbar/context-menu icons unified on one FA7 sprite;
theme-aware G6 context menu (readable in dark); G6 popup
transitions disabled (immediate, no glide); "reset zoom" home
icon restored; bolder/longer "create node" plus.
- **feat(lib-yui): shell / nav**. View-owned dynamic 3rd-level
runtime subroute; unknown route falls back to the default
route; secondary-nav zone collapses from config; single-row
toolbar on touch; `toolbar type:"connection"` + `context_action`
+ modal `on_close`; TreeDB topics persist the selected topic
across reloads with self-contained tab navigation; TreeDB
table row-count footer and email/tel/url subtypes; edit/delete
modal mount fallback.
- **fix(lib-yui): self-containment / responsiveness**.
`C_G6_NODES_TREE` self-contained `ResizeObserver` and toolbar
reconfig guarded until the graph is rendered;
`C_YUI_TREEDB_GRAPH` emits `EV_OPERATION_MODE_CHANGED`; g6
views detect theme from `<html data-theme>`; `C_YUI_UPLOT`
responsive width.
- **feat(treedb): system schema v5 → v6**. Restore the
`cols.topics` fkey, refresh `cols` topic, show all system
topics; keep the original ArtGins start year as a range
(2024-2026).
- **feat(ycommand): `--editor`/`-e` and stdout dump**. Dump a
file to stdout when stdout is not a TTY (was: always vim);
`ac_read_file` signals exit on success, not just on error;
`pty_sync_spawn` drains the master pty after child exit
(was: truncated `cat`).
- **feat(c_mqiogate): `broadcast` method** to fan out events to
every child (tidy `lastdigits` to mirror it).
- **chore(packages / ci)**. Ship sanitized agent JSON templates
from the repo (drop the `/yuneta/agent/` dependency); move
`RELEASE` to the repo root; `release-deb` generates a default
`.config` via `alldefconfig` and drops the pgdg apt source.
- **chore(ext-libs): TODO bump nginx 1.28.3 → 1.30.1**
(CVE-2026-42945) for the next ext-libs refresh.
- **misc(C)**: remove `sf_deleted_record` flag; fix stale
`md2_record_t` "Size: 96 bytes" comment; log the key with bad
metadata; revert the `C_TCP_S channel_filter` two-TLS-listener
change.
- **feat(tr2list): `--dry-run` and `--follow` modes**.
`--dry-run` / `-n` prints the resolved search parameters and
the `match_cond` JSON (times already resolved by `approxidate`),
plus a human-readable rendering of any `from-t`/`to-t`/`from-tm`/
`to-tm` set — respects `--print-local-time` and flags millisecond
input. `--follow` / `-F` opens an `rt_disk` list and runs
`yev_loop_run` until SIGINT (tail-f style; single topic, so it
errors out when combined with `--recursive`). `--help` now
documents the full `approxidate` grammar accepted by TIME
options (units, specials, absolute forms) via argp's `\v`
separator.
- **feat(treedb_list): `--dry-run` and `--follow` modes**.
`--dry-run` / `-n` runs `resolve_treedb_path` and prints the
deduced path / database / topic alongside the filter and
options JSON (also flags resolution failure and falls back to
the raw user input). `--follow` / `-F` keeps listening for
node CREATED / UPDATED / DELETED / LINKED / UNLINKED events
after the initial listing — uses the existing rt_disk path
that treedb already opens internally when `master=false`, plus
a `treedb_set_callback` that honours `--topic` and `--ids`.
Errors out on `--follow --recursive` and on `--follow` with
`--print-tranger` / `--print-treedb`.
- **fix(helpers/approxidate): accept short unit suffixes**.
`1s`, `1m`, `1h`, `1d`, `1w`, `1M`, `1y` (plus `1sec`, `1mi`,
`1min`, `1mo`, `1hr`, `1wk`, `1yr`) now resolve to the
expected relative duration instead of silently falling through
to the numeric date parser as day-of-month / year. Lowercase
`m` keeps minute, uppercase `M` is month (case-sensitive to
disambiguate, mirroring `sleep` / `find -mmin` conventions).
`mon` is intentionally left as Monday so weekday parsing
keeps its existing behaviour. Benefits every yuneta tool
that consumes `approxidate` (tr2list, treedb_list, ybatch,
tr2search, tr2keys, tr2migrate, ...).
- **refactor(tr2list): simpler `--dry-run` time block**.
Each time line now ends with `(<show_date_relative>)` —
`2 hours ago`, `3 days ago`, etc. — instead of echoing the
raw input and warning about parser footguns. The warning
block in `--help` is dropped and replaced by a `short`
group listing the new 1-3 char unit forms accepted by
`approxidate`.v7.3.2 -- 09/May/2026¶
- **feat(release): publish runtime `.deb` on GitHub Releases +
one-liner `install.sh`**. First CI workflow in the repo
(`.github/workflows/release-deb.yml`) builds the AMD64 `.deb`
on `release.published` (or `workflow_dispatch` against an
existing tag) via `packages/AMD64.sh` and uploads it as a
release asset. Pairs with a new `install.sh` at the repo
root: a POSIX one-shot installer that detects host arch
(`amd64` / `armhf` / `riscv64`), queries the GitHub Releases
API for the latest (or pinned) tag, downloads the matching
`yuneta-agent-*-<arch>.deb`, and installs it via
`dpkg + apt-get -f`:
curl -fsSL https://raw.githubusercontent.com/artgins/yunetas/main/install.sh | sudo sh
Pin a version with `sudo sh -s -- 7.3.2`. ARMhf / ARM32 /
RISCV64 wait for cross-compile or matching runners. Past
releases (7.2.0 .. 7.3.1) have no `.deb` assets — the
workflow operates forward.
- **refactor(packages): extract `RELEASE` to a shared
`packages/RELEASE` file** (reset to `1`). The four arch
wrappers had `RELEASE` hardcoded with divergent counters
(3× "9", 1× "6"); now all four read from one file the same
way they read `YUNETA_VERSION`. Yunetas isn't widely
distributed yet, so the renumbering is harmless.
- **docs(installation): rewrite as 7-step "guía burros" path**.
`installation.md` restructured: prerequisites + 7 numbered
steps from "create the `yuneta` user" through "build and
test", with verbose detail (apt explanations, miniconda
bootstrap, full `menuconfig` options) tucked into dropdowns.
Adds a top-of-page **Quick install** section with the
`install.sh` one-liner and clarifies that the PyPI `yunetas`
package (0.x) is the management CLI, **not** the framework
runtime (7.x). Step 5 documents the env vars `yunetas-env.sh`
exports — `YUNETAS_BASE`, `YUNETAS_OUTPUTS`, `YUNETAS_YUNOS`
— plus the `PATH` prepends and the layout contract
(`outputs/` and project repos as siblings of the `yunetas`
repo). Adds an explicit "re-source per shell" warning, a
silent footgun in cron / SSH / CI sessions where
`ybatch` / `ycommand` vanish from `PATH`.
- **fix(gobj-js): `DTP_STRING` attr coerces null / undefined to
`""`**. `json2item` used `JSON.stringify()` as the catch-all
coercion for non-string values; for `null` that produced the
literal 4-char string `"null"`, which leaked into IEvent
payloads. Specifically: `c_ievent_cli`'s `IDENTITY_CARD`
sent `"jwt": "null"` to the backend, defeating the
`empty_string()` check in `c_ievent_srv` that drives the BFF
httpOnly-cookie auth path; `verify_token` then tried to
validate the literal `"null"` as a JWT and failed with
"No OAuth2 Issuer found". Treat `null` and `undefined` as
`""` in `DTP_STRING` to bring JS in line with the C runtime
(where `DTP_STRING` cannot hold a `NULL` pointer).
- **refactor(C kernel): log hygiene for monitor stats**. Several
warning counters in the global-warnings dashboard were noisier
than they needed to be:
* `c_auth_bff`: 4xx HTTP responses logged as warning instead
of info (5xx still error). Sudden 4xx bursts now show in
dashboards that filter on warning severity.
* `c_prot_mqtt`, `c_prot_mqtt2`: malformed CONNECT frames
(client-side protocol issues) downgraded from error to
warning; rejection messages tightened so v1 and v2 paths
bucket into the same precise counter. Dump the offending
gbuf when `handle__connect` returns < 0 so the bad CONNECT
can be inspected in the trace.
* `c_authz`, `c_ievent_srv`: dropped the duplicate
"Authentication rejected" warning in `c_ievent_srv` (each
`result < 0` path in `c_authz::mt_authenticate` already
logs its own); audited `mt_authenticate` so every
`result < 0` contributes to the per-msg stats counter;
fixed a `peername`-empty branch whose `msg` field was
leaking into the unrelated `dst_service`-not-found stats
bucket.
* `ydaemon`: translated the lone Spanish `msg` field
("Soy el Matador" → "I am the killer") so monitors and
search tools group cleanly under the English-only
convention.
- **feat(lib-yui): toolbar brand / avatar / dropdown item types
with per-item `show_on`**. Three new toolbar item kinds
validated by `shell_toolbar_helpers.js` and rendered by
`c_yui_shell.js`: `type:"brand"` (logo image + wordmark, with
optional action — passive `<div>` if action is omitted),
`type:"avatar"` (circular initials rendered from a
host-registered provider via the new
`yui_shell_set_avatar_provider` /
`yui_shell_refresh_avatars` helpers), and
`action.type:"dropdown"` (panel mounted on the popup layer
with `divider` entries, focus-trap, escape-stack push/pop and
capture-phase click-outside dismissal — closes on `scroll`
and `resize` to match native `<select>` UX). `show_on` now
applies per item, not just per area. CSS for all three
shipped, SHELL.md §3.4 cheatsheet rewritten and §10
"Implemented" updated. 23 new unit tests for the validators,
27 chromium e2e specs still pass.
- **build(linux-ext-libs): nginx / openresty link against system
libs; ncurses switched to widec for UTF-8**. The vendored
OpenSSL / PCRE2 stay only for yuneta's own static binaries
(`ytls`, `yev_loop`); nginx and openresty now embed
`libssl` / `libcrypto` / `libpcre` / `libz` from the host,
same as the distro-packaged nginx — closes a latent
Makefile-clobbering bug in `re-install-libs.sh`. Ncurses
re-enabled `--enable-widec` (v1.11) so `ycli` and `mqtt_tui`
render UTF-8 emoji / accents instead of `M-x` escape
sequences; consumers migrated to `<ncursesw/...>` and call
`setlocale(LC_ALL, "")` before `initscr()`. Also:
`MAKEFLAGS=-j$(nproc)` for parallel builds, mbedtls Debug →
Release, and explicit Release+static+PIC flags across mbedtls
/ jansson / pcre2 / libbacktrace / argp-standalone.
- **fix(c_auth_bff): shrink `legacy_base` buffer to silence
`-Wformat-truncation`**. PATH_MAX-sized `legacy_base` plus
`/token` or `/logout` suffix into a PATH_MAX destination
tripped GCC's truncation analysis. A legacy Keycloak base
URL is realistically well under 1 KB.
- **fix(lib-yui): TomSelect re-initialisation guards**. The
"Tom Select already initialized on this element" exception
could be thrown when `build_topic_modal` ran twice — the
query for `.select2-multiple` was matching inputs in earlier
modals still attached to the popup-layer. Scope the query
to the freshly built `$element`; also add a defensive skip
when the element already has a `tomselect` instance.
- **feat(gobj-c, gobj-js): EV_ON_OPEN_ERROR — close before open**.
When a connection-oriented gobj closes before ever opening (TCP
connect failed, TLS cert refused, non-101 handshake response,
handshake timeout, firewall) it now publishes a separate
`EV_ON_OPEN_ERROR` instead of `EV_ON_CLOSE`, preserving the
EV_ON_OPEN→EV_ON_CLOSE FSM contract for subscribers that only
handle close in their connected state. Declared as a kernel
event in `g_ev_kernel.{h,c}` and wired in:
* `kernel/c/root-linux/src/c_ievent_cli.c` (IEvent client)
* `kernel/c/root-linux/src/c_websocket.c` (low-level WS)
* `kernel/js/gobj-js/src/c_ievent_cli.js` (browser client)
Mirrors the browser WebSocket split (.onopen/.onclose/.onerror).
Flagged with `EVF_NO_WARN_SUBS` so backend FSMs that ignore it
don't trip the no-subscribers warning; interactive frontends
opt in. Retry policy unchanged: the connection-responsible
gobj keeps reconnecting forever while running — only the parent
(by stopping the gobj) decides to give up. Each emission also
writes a `log_warning` (`MSGSET_CONNECT_DISCONNECT` in C)
including the remote yuno identity / url / peername — gives
logcenter and other monitors a precise per-attempt alert that
a silent retry loop is in progress.
- **fix(lib-yui): bare-route redirect skips decorative items**.
`navigate_to()` was using `submenu.items[0].route` as the
fallback for a level-1 container — undefined when item 0 is a
`type:"header"` / `type:"divider"`, which caused the bare route
to fall through to "no target". Use the first item with a
`route` instead; `submenu.default` still wins. SHELL.md §3
updated.
- **feat(lib-yui): item tooltips**. Nav and toolbar items accept a
`tooltip` field (fallback: `aria_label`); rendered as the HTML
`title` attribute on the generated `<a>`/`<button>`.
- **feat(yuno-skeleton): `js_gui` template**. New skeleton type for
JS GUI yunos — Vite + lib-yui declarative shell with locales/
(en+es), public/ web assets, 5 placeholder primary areas, and a
burger drawer hosting Account + Help. Registered in
`__skeletons__.json` (type: Yuno; vars: version, description,
author, author_email, license_name).
- **feat(gobj-js, lib-yui): translatable tooltips**. Nav and toolbar
items rendered by lib-yui now also emit `data-i18n-title="<key>"`
next to their `title` attribute, and `refresh_language()` in
gobj-js gained a second pass that walks `[data-i18n-title]` and
re-translates the `title`. Hover tooltips swap language alongside
the visible labels.
- **feat(gobj-js, lib-yui): translatable aria-labels**. Nav and
toolbar renderers now also emit `data-i18n-aria-label="<key>"`
next to their `aria-label` attribute (toolbar root, action items,
brand, avatar, dropdown panel, dropdown rows, and nav items), and
`refresh_language()` walks `[data-i18n-aria-label]` to rewrite
`aria-label`. Screen-reader names now follow the active locale.
- **fix(lib-yui): toolbar dropdown anchor drift on scroll/resize**.
The `position:fixed` panel coordinates are frozen at open time
from `getBoundingClientRect()`; any layout shift previously left
the panel detached from its trigger. Match native `<select>` UX
and dismiss the dropdown on scroll (capture, passive — catches
every ancestor scroller) and on window resize.
- **refactor(gui_treedb): apply locale convention**. Trimmed
en.js/es.js to the 19 keys actually called from src/ (auth_bff
protocol IDs + the half-dozen `t(...)` calls in
`c_yuneta_gui.js`); deleted ~140 aspirational entries that had
no caller. Renamed `remote-service` → `remote service`,
`connection-backend-refused` → `connection to backend refused`
(rule: spaces, not kebab); fixed top-level `nombre:` → `name:`
to match the rest of the codebase. Added
`keySeparator: false` + `nsSeparator: false` in
`setup_locale()` so a future dotted key (e.g. device-namespace)
doesn't fall silently to nested-lookup. Same
`scripts/validate-locales.mjs` + `prebuild` wiring as
wattyzer. Auth_bff snake_case codes kept as-is (wire
contract, see `c_auth_bff.c`).
- **feat(yuno-skeleton): locale convention + validator**. The
`js_gui` template now ships `scripts/validate-locales.mjs`
(asserts every i18n key is ASCII + lower-case + present in every
locale) wired as `npm run validate-locales` and `prebuild`.
en.js/es.js header banners spell out the convention so new
yunos inherit it from day one.v7.3.1 -- 30/Apr/2026¶
- **breaking(auth): standard OIDC migration of `c_auth_bff` and
`c_task_authenticate`**. Both gclasses now resolve IdP endpoints
in the same priority order:
1. Explicit `token_endpoint` + `end_session_endpoint` attrs
(full URLs, skips discovery — one fewer round-trip).
2. `issuer` attr — task chain prepends a GET of
`<issuer>/.well-known/openid-configuration` and caches the
resolved endpoints in priv before the auth flow runs.
3. Refuse to start.
Any conformant OIDC IdP works (Keycloak, Auth0, Cognito, Azure AD,
Authentik, ...). Hardcoded Keycloak path scheme removed.
- **`c_task_authenticate` and its 6 callers** (`c_cli`, `c_mqtt_tui`,
`c_ycommand`, `c_ystats`, `c_ytests`, `c_ybatch`) had their
legacy `auth_url`+`auth_system` attrs **removed outright** and
the `azp` attr **renamed to `client_id`** to match the form
parameter actually sent on `/token` and `/logout`.
- **CLI flag set** in `ycommand` / `ystats` / `ytests` / `ybatch` /
`mqtt_tui` is now `-I/--issuer`, `-T/--token-endpoint`,
`-E/--end-session-endpoint`, `-Z/--client-id`. Old `-K/--auth_system`,
`-k/--auth_url` and `-Z/--azp` (renamed) are gone.
- **`c_auth_bff` keeps `idp_url`+`realm`** as a deprecated path
(warning fired at `mt_create`); removal scheduled once one
release has shipped with the warning in place. See
[`TODO.md`](TODO.md) for the remaining smoke tests against
non-Keycloak IdPs and the open ROPC-vs-PKCE question.
- **feat(gobj, gobj-js): `SDF_DEPRECATED` attribute flag**. New
sdata flag (`0x00000100`) to mark a gclass attribute as deprecated.
Both the C runtime and the JS runtime emit a warning when a
deprecated attribute is set during gobj creation, naming the
gclass and the attr. First adopter: `c_authz::authz_yuno_role`
(use `authz_service` instead).
- **test(c_task_authenticate)**: new self-contained suite under
`tests/c/c_task_authenticate/` (`test1_discovery`,
`test2_explicit_endpoints`, `test4_discovery_failure`). Mock IdP
gclass with `override_*_body` knobs for failure injection;
shared `test_main.c` boilerplate; the driver subscribes to
`EV_ON_TOKEN`, asserts the result code, and dies.
- **test(c_auth_bff)**: new `test17_legacy_idp_url` covers the
`idp_url`+`realm` deprecation path that tests 1–16 missed.
Captures the deprecation warning at `LOG_OPT_UP_WARNING` and
drives the full login flow against the same mock-Keycloak.
- **feat(lib-yui): declarative app shell `C_YUI_SHELL` + `C_YUI_NAV`**.
A JSON-driven replacement for `C_YUI_MAIN` + `C_YUI_ROUTING`, shipped
alongside the legacy stack (no migration planned — see
[`SHELL.md` §10](kernel/js/lib-yui/SHELL.md)). New GUIs can adopt
the new shell; existing GUIs keep using the old one unchanged.
- **Layered grid**: 6 z-stacked layers (`base`, `overlay`, `popup`,
`modal`, `notification`, `loading`) and 7 zones (`top`, `top-sub`,
`left`, `center`, `right`, `bottom-sub`, `bottom`) inside `base`,
all driven by a single declarative JSON config.
- **Six menu layouts**: `vertical`, `icon-bar`, `tabs`, `drawer`,
`submenu`, `accordion`. Same menu may render differently per
zone via `render[zone]`. Auto-expand of the active branch on
accordion when the route changes.
- **`show_on` parser**: zone visibility per Bulma breakpoint with
the operators `>=`, `<=`, `<`, `>`, enumeration and `|`. Pure
module (`shell_show_on.js`), 13 `node --test` unit tests.
- **Three lifecycle modes per item** (`eager` / `keep_alive` /
`lazy_destroy`) decide when the routed view is created and
destroyed.
- **Single router**: `C_YUI_NAV` publishes `EV_NAV_CLICKED`; the
shell publishes `EV_ROUTE_REQUESTED` (intent, audit witness)
and `EV_ROUTE_CHANGED` (fact). Hash-based 2-level routing,
no dependency on `C_YUI_ROUTING`.
- **Drawer overlay** on the `overlay` layer with focus-trap
(Tab/Shift+Tab cycling, focus restoration on close), backdrop
click closes via `EV_DRAWER_CLOSE_REQUESTED` (canonical close
path with focus-trap release + escape-stack pop).
- **Escape priority chain**: `priv.escape_stack` is a LIFO of
`{layer, handler}`; the global `keydown` listener calls only
the top entry. Modal-over-drawer closes the modal first.
Public API `yui_shell_push_escape` / `yui_shell_pop_escape`
for app-level overlays.
- **Modal / notification API** on top of the shell layers
(`yui_shell_show_info` / `show_warning` / `show_error` /
`show_modal` for non-blocking; `yui_shell_confirm_ok` /
`confirm_yesno` / `confirm_yesnocancel` for blocking dialogs
that resolve a Promise). Each modal/dialog auto-pushes onto
the Escape stack and installs a focus-trap. Bulma `.modal-card`
/ `.notification` markup verbatim. Generic focus-trap moved
to `shell_focus_trap.js` with 10 unit tests.
- **Canonical i18n via `data-i18n` + `refresh_language`**: every
translatable text node carries `data-i18n="<canonical key>"`;
apps switch language by calling
`refresh_language(shell.$container, t)` from `@yuneta/gobj-js`,
the same flow `c_yui_main.js` uses in `change_language()`.
Modals/dialogs accept `opts.t` so they render in the active
language at open time AND retranslate live afterwards.
- **Generalised secondary-nav loop**: `instantiate_menus()` walks
every menu mounted via a `"menu.<id>"` host whose items declare
a `submenu` (not just `menu.primary`). Synthesised menu_id is
`secondary.<owning_menu_id>.<item.id>`, scoped so two
primary-style menus can share item ids without colliding.
- **`gcflag_no_check_output_events`** on the shell so the toolbar
can publish arbitrary user-defined events
(`action.type:"event"`) without each app having to extend the
shell's `event_types` table.
- **Hard contracts**: every view gclass MUST expose `$container`
in `mt_create`; every navigation through an empty/unknown route
logs `log_error` and surfaces a placeholder banner; every
try/catch logs via `log_warning` (no silent swallow).
- **`validate_config()`**: system-boundary guard run at the top
of `mt_start`. Rejects malformed configs with a visible
"invalid config" banner instead of producing a half-built
shell. Checks: object/array shapes, zone-id membership in the
7 valid zones, `host` syntax (`toolbar` | `menu.<id>` |
`stage.<id>`), stage zones declared in `shell.zones`, and
cross-menu route-target uniqueness (warn when two menus claim
the same target).
- **Playwright e2e harness**: 22 spec files × 3 browsers
(chromium + firefox + webkit) = 69 tests covering boot /
navigation / drawer / modals / multimenu / validator /
lifecycle / breakpoint / live-i18n. CI workflow
`.github/workflows/lib-yui.yml` runs unit + e2e on PRs and
pushes touching `kernel/js/lib-yui/**` or
`kernel/js/gobj-js/**`. `kernel/js/lib-yui/install-e2e-deps.sh`
helper installs the apt packages WebKit links against
(`libgstreamer-plugins-bad1.0-0`, `libavif16`).
- **Test-app**: standalone harness in `kernel/js/lib-yui/test-app/`
with three presets (`default`, `?preset=accordion`,
`?preset=multimenu`) plus a deliberately-broken `?preset=invalid`
used by the validator regression test. `C_TEST_LANG`
controller demonstrates the canonical pattern for reacting to
custom toolbar events (language toggle, hello toast, ask
dialog).
- **Docs**: [`SHELL.md`](kernel/js/lib-yui/SHELL.md) (design,
configuration JSON, GClasses + events, modal/notification API,
Escape chain, internationalisation),
[`TODO.md`](kernel/js/lib-yui/TODO.md) (status of every task on
the new shell), updated `lib-yui/README.md` with the
"Which app shell to use?" decision tree.
- **CLAUDE.md**: new "GClass section layout" addendum (JS skeleton
banners + canonical CHILD/SERVICE subscription model + Always
braces rule + EVF_NO_WARN_SUBS) so future agents stay on the
rails the user established for this work.v7.3.0 -- 18/Apr/2026¶
- **feat(ytls, c_yuno, c_agent): TLS certificate hot-reload with
three-layer defence-in-depth**. Lets a Yuneta host keep thousands of
persistent TLS connections alive across a Let's Encrypt renewal,
with no deploy-hook single point of failure.
- **ytls**: new [`ytls_reload_certificates()`](docs/doc.yuneta.io/api/ytls/ytls.md)
that rebuilds the backend context (OpenSSL `SSL_CTX` or mbed-TLS
`mbedtls_state_t` bundle), validates it, and atomically swaps it
in. Live sessions hold their own refcount on the previous context,
so already-established connections keep working until they close.
Invalid material rolls back cleanly — traffic is never interrupted
by a bad reload. `ytls_get_cert_info()` returns
`{subject, issuer, not_before, not_after, serial, days_remaining}`
for the live context, not just the file on disk.
- **c_agent**: new cert auto-sync timer (attr
`cert_sync_interval_sec`, default 900 s) that re-reads
`/yuneta/store/certs/` via `sudo -n copy-certs.sh`; when any
`size+mtime` changes, broadcasts `reload-certs` to every running
yuno. Exposes `cert-sync-now` / `cert-sync-status` commands and
self-heals if the certbot deploy hook fails silently.
- **c_yuno**: periodic expiry monitor (attr `timeout_cert_check`,
default 3600 s) that walks every `C_TCP_S` / `C_UDP_S` listener
and logs `gobj_log_warning()` at `cert_warn_days` (default 7) and
`gobj_log_critical()` at `cert_critical_days` (default 2).
Alert-only — the sync layer owns the reload responsibility.
- **c_tcp_s / c_udp_s**: per-listener `reload-certs` and `view-cert`
commands, routable via `ycommand -c 'command-yuno command=reload-certs
service=__yuno__'` or `gobj=<name>` for a single listener.
- **packages**: `/etc/letsencrypt/renewal-hooks/deploy/reload-certs`
hook copies certs, reloads the web server and broadcasts the
yuno-level reload. Each step runs with `set +e`; output is logged
to `/var/log/yuneta/deploy-hook.log` and the hook writes its last
run timestamp to `/var/lib/yuneta/last-deploy-hook-run` so
`cert-sync-status` can spot a hook that never runs.
- **tests**: `tests/c/ytls/test_cert_reload`,
`test_cert_info`, `test_cert_reload_mem` (1000 reloads, zero leak)
and `tests/c/yev_loop/yev_events_tls/test_yevent_reload_live`,
`test_yevent_reload_stress` (50 reloads with a live session).
- **docs**: new guide [`guide/guide_cert_management.md`](docs/doc.yuneta.io/guide/guide_cert_management.md)
covers the end-to-end story, layered design and file / permission
layout; `guide/guide_ytls.md` gains a hot-reload section.
- **feat(gobj): `gobj_set_manual_start()` + `gobj_flag_manual_start`**.
A gobj can now opt out of the automatic `start-tree` walk so its
parent keeps ownership of lifecycle but decides *when* to bring it
up. Used in `c_auth_bff` to keep `gobj_idprovider` dormant until the
BFF has validated its configuration.
- **feat(ycommand)**: major interactive / scripting overhaul.
- TAB completion of command names, parameter names and boolean values,
from a remote `list-gobj-commands` cache fetched at connect time
(routed through `service=__yuno__`) and from a local command table
for `!cmd` built-ins.
- Inline parameter hints in gray (`<name=type>` required,
`[name=type]` optional, already-typed params dropped).
- Connect-time informative prompt (`<role>^<name>> `) and schema-driven
table rendering in both interactive and non-interactive modes (use
the `*cmd` prefix to force raw-JSON form).
- `Ctrl+R` / `Ctrl+S` incremental history search, `Ctrl+L` clear screen,
bash-style `!!` / `!N` history expansion, erasedups history.
- c_cli-style local commands via the `!` prefix: `!help` (alias `!h` /
`!?`), `!history`, `!clear-history`, `!exit` / `!quit`,
`!source <file>` (alias `!.`). Full keybinding + syntax reference
available as `!help` and in `utils/c/ycommand/README.md`.
- Command chaining with `cmd1 ; cmd2 ; cmd3` (quote/brace-aware split),
`-cmd` ignore-fail (ybatch convention), stdin piping
(`cat batch.ycmd | ycommand -u ws://...`). A single shared
command queue drains one command at a time, waiting for the previous
response before sending the next.
- `did-you-mean` suggestions on `command not available` errors,
Levenshtein-matched against the cache.
- Positional command form (`ycommand kill-yuno id=foo`, equivalent to
`-c`). The `-c` flag still wins when both are present.
- **feat(c_editline)**: new public helpers shared by every editline
client — `editline_set_completion_callback` /
`editline_set_hints_callback` / `editline_add_completion` /
`editline_history_count` / `editline_history_get`. New events
`EV_EDITLINE_REVERSE_SEARCH` / `EV_EDITLINE_FORWARD_SEARCH` for
incremental history search; candidate list + description is rendered
on TAB when multiple options exist.
- **fix(c_editline)**: after the user selects a TAB candidate, the
keystroke that committed the selection (Enter, Backspace, printable)
is now re-dispatched so the action takes effect in the same press
instead of requiring a second press.
- **fix(ycommand)**: `on_read_cb` no longer drops trailing bytes of a
batched read that matched a keytable entry, so rapid TAB+value typing
no longer needs a second press.
- **feat(ycli)**: TAB completion brought in line with ycommand, adapted
to the multi-window ncurses UI.
- `!cmd<TAB>` completes local `c_cli` commands; `cmd<TAB>` (no `!`)
completes remote commands of the yuno attached to the focused
display window. Cache is per-connection, fetched silently on
`EV_ON_OPEN` via `list-gobj-commands` and dropped on
`EV_ON_CLOSE`.
- Multi-candidate list is rendered in a temporary ncurses popup
above the editline (no more blocking `read(STDIN_FILENO)` inside
the yev_loop callback); cycling is driven through the normal FSM
(TAB / Up / Down navigate, Enter commits to the edit line only,
Esc / Ctrl+G / Backspace cancel, printable keys commit + insert).
- Scrollable popup with a status row (`N/M ↑ K above ↓ L below`)
rendered in dim attributes so A_REVERSE on the selected row can
never bleed into it.
- Inline hints (`<req=type>` / `[opt=type]`) in gray (A_BOLD on
COLOR_BLACK = bright-black / gray in most terminals).
- **feat(c_editline)**: new `EV_EDITLINE_CANCEL` event for escape-style
cancellation of reverse-i-search and TAB-popup sub-modes; `refreshSearchLine`
now draws through ncurses (`wmove/waddnstr/wrefresh`) on `use_ncurses`
clients instead of bypassing the pane via `printf`.
- **feat(ycli / ycommand)**: `Ctrl+K` switched to readline semantics —
delete from cursor to end of line (`EV_EDITLINE_DEL_EOL`).
`Ctrl+U` / `Ctrl+Y` remain "delete whole line"; `Ctrl+L` is the
clear-screen shortcut (previously shared with `Ctrl+K`).
- **docs**: added `utils/c/ycommand/README.md`, `TODO.md` and updated
`docs/doc.yuneta.io/{utilities,yunos,modules}.md` to cover the new
features.
- **API change(ghttp_parser)**: `ghttp_parser_reset()` is **removed** from
the public API. It was a foot-gun: calling it from inside an llhttp
callback (as `on_message_complete` used to do) corrupted llhttp's state
machine and silently swallowed pipelined messages. Callers that need a
pristine parser for a new connection now use the destroy+create cycle
(see `c_prot_http_sr::ac_connected`, `c_prot_http_cl::ac_connected`,
`c_websocket::ac_connected`). The llhttp settings vtable is now
initialised once, lazily, via `llhttp_settings_init()` in
`ensure_settings_initialized()`.
- **feat(ghttp_parser)**: new `ghttp_parser_finish()` that signals
end-of-stream (`llhttp_finish()`) to the parser. Fixes a latent bug
where HTTP/1.0 responses (or HTTP/1.1 `Connection: close` responses
without `Content-Length` / `Transfer-Encoding: chunked`) never fired
`on_message_complete` because the peer's socket close was the only
message terminator. Wired up in `c_prot_http_cl::ac_disconnected`
(the critical case for response parsers), `c_prot_http_sr::ac_disconnected`,
and `c_websocket::ac_disconnected`.
- **fix(ghttp_parser)**: on `HPE_PAUSED_UPGRADE`, `ghttp_parser_received()`
now returns the actual number of bytes llhttp consumed (computed via
`llhttp_get_error_pos()`) instead of lying that it consumed the whole
buffer. This lets the caller re-route any tail bytes that belong to
the new protocol (e.g. a WebSocket frame piggy-backed on the same TCP
segment as the upgrade request) to the next handler.
- **CRITICAL fix(ghttp_parser)**: HTTP/1.1 pipelining was silently broken —
`on_message_complete()` called `ghttp_parser_reset()`, which in turn called
`llhttp_init()` from inside the llhttp callback, corrupting the parser's
internal state machine so every subsequent message in the same buffer was
swallowed without a log. Affects every yuno serving or consuming HTTP
over keep-alive when more than one message is in flight on a single
connection (c_prot_http_sr, c_prot_http_cl, c_websocket). Fix: reset the
per-message app fields inline in `on_message_complete` without touching
llhttp; leave `ghttp_parser_reset()` for the other (non-callback) call
sites. Surfaced by the new test suite `tests/c/c_auth_bff/test8_queue_full`.
- **refactor(c_auth_bff): IdP-agnostic naming, single-job task, queue +
routing hardening**. The BFF used to be visibly wired to Keycloak
(`kc_*` attrs, stats, logs). Code, attrs and stats now use the
generic `idp_*` prefix; any OIDC provider fits. The outbound IdP
gobj chain is now named `<bff-name>-idp` for trace clarity.
- **Pending queue** migrated from a fixed-size `PENDING_AUTH *` ring
to a `dl_list`, drained one job at a time. Configurable per
instance via `pending_queue_size` (default 16, clamped to
`[1, 1024]`). Overflow bumps `q_full_drops` and the browser sees
a mapped `error_code`; peak depth is exposed as `q_max_seen`.
- **Flush-on-disconnect**: when a browser closes mid-round-trip the
BFF flushes its pending queue for that channel; late IdP replies
for disconnected clients are dropped (`responses_dropped` counter)
instead of being forwarded. Each task also carries a per-browser
generation so a cross-user token leak cannot occur.
- **Single-job task, teardown-safe close**: the C_TASK instance
holds a single job at a time; `mt_stop` drains the inbound
`C_PROT_HTTP_SR + C_TCP` chain and the outbound `gobj_http` so a
SIGTERM with live browser connections no longer logs
"Destroying a RUNNING gobj".
- **Outbound watchdog**: per-instance attr `idp_timeout_ms`
(default 30000, 0 disables) armed via a `C_TIMER0` child right
after the outbound HTTP client is created and cleared in
`ac_end_task`. On fire, responds 504 to the browser and drains
the task; closes the "IdP silence → channel wedged forever"
deadlock. New `idp_timeouts` stat counter.
- **IdP health signal fix**: count any 2xx IdP reply as `idp_ok`;
previously only 200 counted, so every successful `/logout`
(Keycloak returns spec-compliant 204 No Content) poisoned the
ratio as an `idp_error`.
- **Logout routing fix**: route the logout reply to the bottom
browser channel, not to the dangling `_browser_src` from an
earlier round-trip.
- **`mt_stats` filter** mirrors the default `stats_parser.c`
two-stage matcher (full name OR underscore-prefix) and is
case-insensitive, so `gobj_stats(bff, "idp_", ...)` returns the
idp_* set as expected. `redact_for_trace()` key matching is also
case-insensitive so HTTP headers like "Cookie"/"cookie"/"COOKIE"
are all masked.
- **Stats moved to PRIVATE_DATA + `mt_stats`** for zero hot-path
cost; the gclass now also exposes a stats/queue-state command
through the normal command interface.
- **Stable `error_code`** in every BFF response (snake_case, e.g.
`invalid_refresh_token`, `idp_unreachable`, `queue_full`) — the
GUI uses this as its i18n translation key. Action-aware error
mapping wired through `gui_treedb`.
- **Log hygiene**: 4xx IdP replies are logged as `INFO`, not
`ERROR` (a wrong password is not a server error), with
`MSGSET_PROTOCOL`. New `messages` / `traffic` trace levels; 👤
BFF log prefix and ⏩/⏪ direction arrows across BFF traces.
- **Own orchestrator GClass** at the top of the `auth_bff` yuno
(replaces the citizen-yuno shortcut) and `gobj_idprovider` is
tagged `gobj_flag_manual_start` so it stays dormant until the
BFF validates its configuration.
- `gobj_http` single-instance invariant is now asserted in debug
builds to catch re-entrancy regressions.
- **perf(auth_bff)**: new `perf_auth_bff` ping-pong-style live
throughput benchmark (`performance/c/perf_auth_bff/`). Default
10 s run, ~180 000 ops on the reference box; registered as ctest.
- **test(c_auth_bff)**: 16-binary suite self-contained under
`tests/c/c_auth_bff/` with a scriptable mock Keycloak
(`c_mock_keycloak`): signed HS256 JWTs, configurable latency /
status / body override. Covers login, callback, refresh, logout,
validation errors, IdP 401, slow IdP, queue pipelining + overflow,
browser cancel mid-round-trip, cancel-then-retry, cross-user stale
replies, expired refresh, 405 / missing body / unknown endpoint.
Gates the watchdog, `browser_alive`, flush-on-disconnect and
ghttp_parser fixes.
- **test(c_llhttp_parser)**: sanity suite for the vendored llhttp
library and the `ghttp_parser` wrapper (`tests/c/c_llhttp_parser/`).
- **stress(auth_bff)**: new concurrent stress runner
(`stress/c/auth_bff/`) that exercises the pending queue, the
watchdog and the flush-on-disconnect path.
- **fix(c_prot_http_sr)**: omit response body on 1xx / 204 / 304
replies (RFC 7230). The parser path was emitting a body for these
status codes, confusing downstream clients and tripping some
proxies.
- **fix(c_task)**: `volatil` gobjs now self-destroy at end-of-work —
making the long-standing `// auto-destroy` comment actually true.
The outbound HTTP client used by the BFF is created `volatil` so
teardown is explicit and framework-free (PR #95). Also silences
the `-Wcomment` warning in the auto-destroy comment and dedups
`TRACE_MESSAGES` / `TRACE_MESSAGES2` output.
- **fix(lib-yui)**: restore `publi_page` iframe rendering for
logged-out users — a regression in the login split hid the public
landing page behind the auth screen.
- **fix(ytls/openssl)**: guard `flush_clear_data` against a
re-entrant `sskt` free under specific TLS teardown paths.
- **build(libjwt)**: yuno skeleton `CMakeLists.txt` templates now
link `${JWT_LIBS}` out of the box (PR #92).
- **refactor(gobj)**: drop TLS knowledge from `gobj-c`, inject it
from the ytls layer via a new `gobj_add_global_variable()`
extension point. Removes the `CONFIG_HAVE_OPENSSL/MBEDTLS` `#if`
blocks from `gobj_global_variables()` and keeps the core
backend-agnostic — `root-linux`'s `yunetas_register_c_core()`
publishes `__tls_library__` and `__tls_libraries__` at startup.v7.2.1 -- 07/Apr/2026¶
- TLS: change Kconfig from radio (choice) to checkboxes — both OpenSSL and mbedTLS can be
enabled simultaneously for runtime backend selection per connection
- TLS: add `__tls_libraries__` global variable (reports all compiled backends)
- Documentation: add Test Suite page, fix glossary warnings, improve gobj-js and lib-yui READMEs
- Remove obsolete defconfig and REVIEW.md
- Fix duplicate measure_times declarations in yev_loop.hv7.2.0 -- 04/Apr/2026¶
- Fully static glibc binaries (CONFIG_FULLY_STATIC): GCC and Clang, with custom
static resolver (yuneta_getaddrinfo) and NSS replacements (static_getpwuid, etc.)
- mbedTLS support as alternative TLS backend (~3x smaller static binaries vs OpenSSL)
- Fix mbedTLS bad_record_mac: accumulate TLS records before writing
- Add TRACE_TLS trace level and mbedTLS debug callback for TLS diagnostics
- JS kernel restructured: gobj-js (7.1.x) and lib-yui (7.1.x) published to npm
- Replace bootstrap-table+jQuery with Tabulator in gui_treedb
- Vite 8 build for lib-yui (ES/CJS/UMD/IIFE bundles)
- MQTT 5.0: will properties, user properties, topic alias, subscription identifiers
- Fix MQTT QoS 2 infinite loop and flow control (receive-maximum, keepalive)
- OAuth2 BFF (auth_bff yuno) with PKCE, httpOnly cookies, security hardening
- TreeDB: compound link improvements, undo/redo history sync, new tr2search/treedb_list utils
- G6 graph visualization: C_G6_NODES_TREE and C_YUI_JSON_GRAPH GClasses
- Fix c_watchfs: memory leak, event name mismatch (EV_FS_CHANGED), buffer bugs
- Fix c_fs: memory leak in destroy_subdir_watch
- Fix XSS vulnerabilities in gui_treedb webapp
- Kconfig: add CONFIG_C_PROT_MQTT, organize protocol modules submenu
- Remove deprecated musl compiler optionv7.0.1 -- 29/Mar/2026¶
- Release 7.0.1
- JS kernel (yunetas npm package) published as v0.3.0
- Updated and documented .deb packaging (packages/)v7.0.0 -- 28/Sep/2025¶
- Publish first 7.0.0 for productionv7.0.0-b17 -- 26/Sep/2025¶
- fix remote console (controlcenter) blocked when paste textv7.0.0-b15 -- 22/Sep/2025¶
- fix yuneta_agent: wrong assignment of ips to public servicev7.0.0-b14 -- 11/Sep/2025¶
- improve .deb
- yuno-skeleton to /yuneta/bin and skeletons to /yuneta/bin/skeletons
- check inherited files only for daemonsv7.0.0-b12 -- 7/Sep/2025¶
- now you can select openresty or nginx in .debv7.0.0-b10 -- 2/Sep/2025¶
- jwt in remote connectionv7.0.0-b9 -- 2/Sep/2025¶
- Remote control (controlcenter) okv7.0.0-b8 -- 29/Aug/2025¶
- GObj: fix bug with rename eventsv7.0.0-b7 -- 29/Aug/2025¶
- Fixed: avoid that yunos (fork child) inherit the socket/file descriptors from agent.