Three admitted learning cohorts and automatic continuation
The fixed demonstration criteria for TODO items 1 and 2 are complete. The accepted model lineage has three prospective learning admissions: conversation, audit and storage. The latest two were prepared, funded, trained, audited and promoted automatically on one native research chain. Every measured previously correct retained answer survived each admission.
These are bounded source-backed knowledge tasks with broader assistant retention checks, operated by one administrator. They establish a working learning and admission method. Independent permissionless hosting, sustainable verification and usable public chat remain the existing TODO items 4–6. The public 0.4.0 network is a separate deployment.
Measured learning
| Admitted cohort | Ordinary singles, before → after | Combined answers, before → after | Retained correct answers lost | Native promotion height |
|---|---|---|---|---|
| Conversation | 1 → 15/16 | 1 → 14/16 | 0 | 8,843, earlier chain |
| Audit | 2 → 15/16 | 1 → 14/16 | 0 | 10,503, continuation chain |
| Storage | 1 → 16/16 | 0 → 14/16 | 0 | 28,685, continuation chain |
Each cohort ran its prescribed 128 updates and evaluated only its terminal checkpoint. The unchanged gates require at least 80% single-answer accuracy, 75% combined-answer accuracy, a positive lower confidence bound on paired single-answer improvement, zero retained-correct losses, and 75% minimum accuracy in each broader retained role. Audit's gain interval has lower bound 0.625; storage's is 0.8125. Three fresh complete numerical audits agreed for each promotion: 105 stages for conversation, 145 for audit, and 177 for storage.
Storage's retention evaluation includes 105 knowledge, 24 skills and 16 conversation cases. It preserved all 128 previously correct answers and gained two additional retained answers. Skills remained 20/24 and conversations 14/16. Ordinary mixed questions combine newly learned knowledge with earlier experts. Routing receives user wording and conversation context, without evaluator labels. The complete ordinary replies include every fresh test response, including the remaining failures.
The three admissions span two research geneses and two committed serving policies. Conversation was accepted before the serving repair. Its accepted weights and complete evaluation history were imported into the continuation; balances and rejected feed weights were excluded. Audit and storage then passed without another source, genesis, gate or checkpoint-selection change. Rejected trials and the opened development repair receive no successful-cohort credit. The earlier failures and repair remain recorded.
Why the method works in this trial
The shared backbone and interpreter remain partitioned across owners. New knowledge is trained into an additional expert tail while accepted model paths remain available. The accepted graph now contains six experts: directory, protocol, planner, conversation, audit and storage. Its nine logical owners ran on seven physical GPU hosts; the allocation rejects placing a complete backbone on any one host.
Selection first decides whether a new expert is eligible, preserving the accepted selector on rejection. Inside an admitted expert, a frozen question-pair model compares the ordinary request with four original training paraphrases per intent. The index contains questions; answers are generated by the trained expert weights. Modal and compound-question preservation prevents dropping request clauses. The 567,755,777-parameter reranker is an additional frozen model on logical owner 3, and its cost is included in both comparison arms.
Native promotion binds the complete answering graph: backbone, expert weights, selector, question access, planning/composition, tokenizer and generation rules. Correctly executed training earns its prescribed reward; quality alone decides whether the resulting graph can replace serving.
Equal-resource comparison
Both arms received seven g5.xlarge hosts, 300 GiB per host and a 9,000-second interval. Both executed the same 128-update numerical trajectory, prefix work, three fresh audits of every window, and three complete quality audits. Two serving replicas consumed the remaining interval using the same frozen queue.
| Result | Add isolated audit expert | Replace existing planner expert |
|---|---|---|
| New single / combined answers | 15/16; 14/16 | 15/16; 14/16 |
| Previously correct retained answers lost | 0 | 4 |
| Complete quality gate | Passed | Failed retention |
| Requests served within the interval | 208 | 1,297 |
| Generated tokens in those requests | 4,002 | 20,848 |
The measured reason to add capacity is preservation. The replacement arm served more requests within its interval. These totals include different amounts of time spent preparing, training, verifying and serving; they are not isolated inference-throughput measurements. Last requests overran their deadlines by at most 6.38 seconds. This is one declared workload and replacement strategy, with complete resource counters, rather than a lifetime-cost optimum. Shadow comparison work issued zero NEURO.
The admission rule remains: add an isolated expert by default; allow updates or consolidation when complete answering quality and retention demonstrate benefit within the resource budget. Parameter count alone does not authorize growth.
Automatic operation and settlement
The new chain rejected its one-update bootstrap after three matching 113-stage audits. Its source curator rejected a separately hashed feed containing an unsupported substituted answer. It then admitted both genuine data jobs, advanced immutable source cursors, trained with replay of actually trained windows, obtained complete funded audits and promoted both passing graphs. The native manifest stayed unchanged throughout that sequence.
The publisher resumed its durable journal across 3,040 actual process starts. There were 226 during-work probes of accepted serving, plus phase-boundary probes. The earlier controller/native-process restart preserved checkpoint, issuance and serving across a settled-boundary restart; the host-restoration continuation preserved the earlier chain and accepted conversation graph. Those are separate recovery experiments, not claims of a host failure in this uninterrupted run.
Full application replay reproduced 28,912 headers and 869 accepted signed transactions, with zero rejected transactions and the exact final state. The chain issued 257 trial NEURO for 256 full-cohort updates plus bootstrap, once per prescribed work identity. A paid response from the final accepted graph settled for six atoms, refunded 970 unused atoms and created no additional tokens.
The curator checks this pinned, source-backed corpus, training/evaluation separation, duplicates and admitted history. It does not establish truth or poisoning resistance for arbitrary web sources. The broader retention cases are small fixed assistant screens, not evidence of general ChatGPT-level capability. All validators, curators and GPU owners in this run share one administrator.
Reproduce the evidence
- Complete execution evidence: public genesis, complete blocks, final state, all producer/audit outputs, source/data objects, checkpoint publication receipts, comparison logs and ordinary replies.
- Matching numerical source, 49d56b3.
- Inputs committed before training.
- Earlier conversation admission and failed-feed history, with its own matching-source reference and replay driver.
The compact measured record and publication manifest bind these claims to their immutable artifacts. Verify the archive SHA-256, then extract source and evidence into separate directories. With the pinned execution dependencies installed, run:
PYTHONPATH=/absolute/path/to/matching-source/src python /absolute/path/to/evidence/replay_public.py
PYTHONPATH=/absolute/path/to/matching-source/src python /absolute/path/to/evidence/report_ordinary_campaign.py \
--home /absolute/path/to/evidence --output /absolute/path/to/reconciled-evidence.jsonApplication replay needs no GPU or signing key. Numerical reproduction also requires the committed A10G runtime and publicly retained tensor objects, identified in the catalog, auxiliary inventory and publication receipts. Every published numerical boundary was hash-verified before pruning. The evidence archive passed full public readback, and its extracted contents were checked against their complete file inventory and quality/ledger commitments.
All seven experiment instances, their volumes and the experiment security group were retired at 11:50 UTC on September 19. The compute upper estimate is $59.59; EBS, S3 and transfer are separate. The earlier failed allocations and their costs remain in their original records. No new GPU trial was needed to verify or report this completed result.
