Skip to content

Prospective ordinary learning with semantic access

Current outcome, September 19: the completed continuation admits audit and storage automatically, following the earlier accepted conversation cohort. All three preserve measured retained answers; the complete resource comparison and native replay pass. Checklist items 1 and 2 are complete against their fixed demonstration criteria. The account below preserves earlier failed trials and the evidence available at each stage.

The first full cohort failed: single answers improved 1→13/16 and combined answers 0→11/16, short of the required 12/16. No retained-correct answer was lost. All three fresh quality auditors reproduced all 105 stages. Native quality rejected the candidate at height 9,383, preserving the accepted serving graph. The chain issued 129 trial NEURO for 128 full-cohort updates plus bootstrap. Later cohort finals remain unopened; the matched-resource comparison did not run. The result and ordinary replies preserve the failure. At this stage, checklist items 1 and 2 remained incomplete. Full application replay reproduced 446 signed transactions and all 9,398 exported headers, including the final state at height 9,397.

An inference-only development repair then improved singles 13→15/16 and combined answers 11→13/16 on the same weights, with zero regressions and zero retained-correct losses. Knowledge retention reached 33/33; skills remained 20/24 and conversations 14/16. All earlier neural responses and a fresh combined request replayed exactly. A small integer classifier now distinguishes questions within the retrieved expert; the original parent-selection boundary remains unchanged. The parser also preserves ordinary fronted and possessive question spans. An off-the-shelf duplicate-question model failed development and was discarded. The complete repair result passes the development accuracy floors, without new neural training, a new final, native promotion or successful-cohort credit.

This trial addresses live-LLM items 1 and 2. It starts from the accepted directory/protocol/planner-expert seed and learns escrow, conversation and feed cohorts in that order. The opened, rejected admission cohort contributes no successful cohort or retained knowledge to this experiment.

The complete candidate uses the development-tested question selector and literal question composition. A pinned BGE encoder indexes training questions only; it stores no answers. New answers must still be generated by trained expert weights across the partition owners. The encoder is an additional 33,360,000 frozen parameters on the coordinator. Its hashes, license, bytes, runtime and placement are recorded; public restoration is required for replay.

Each cohort has 128 updates at learning rate 0.00005. Admission requires the existing 80% single-answer and 75% composed-answer floors, a positive lower confidence bound on answer improvement, no retained-correct answer lost, and the existing broader assistant retention floors. Conversation/feed final bytes remain unchanged and unopened. No final selects a checkpoint or changes gates.

The native operator handles immutable feed discovery, curation, funding, numerical audits, rejection, restart and complete-system promotion. A separate one-update bootstrap exercises rejection; an unexpected pass stops the trial. The first full cohort also has an equal 9,000-second, seven-host comparison against replacing the earlier planner expert. Shadow work earns no issuance.

The original seven disposable owners are reused within the committed resource deadline. This is a controlled, single-administrator research network. Encoder execution is covered by its research allocation; this trial does not establish a sustainable permissionless verification or inference price.

Neither live-LLM item is complete until its measured criteria pass. The prior 15/16 single and 13/16 composed development repair is supporting evidence.

The bootstrap exposed two controller transport faults before any full-cohort final opened. Interrupted service startup leaked worker groups, and concurrent archive extraction could truncate a shared policy while another service read it. Recovery now records groups before startup, cleans every reachable owner, and installs metadata with atomic file replacement. Twenty targeted checks passed, including interrupted writes and readers holding the previous file. Owner 3 required a reboot. Serving was interrupted during recovery; this is not evidence of uninterrupted availability. The same chain then resumed with all four validators agreeing at height 1,295 and a successful accepted-model probe. No genesis, signing state, training input or quality gate was reset.

The recovery source and commitments preserve the operational amendments separately from the original frozen inputs. The three full learning cohorts remain pending at this recovery point.

After recovery, all three bootstrap audits reproduced the producer's complete result across 73 stages each. The chain rejected that deliberately weak candidate at height 1,792, preserving its serving graph and the single verified update's 1,000,000 issued atoms. The controller then rejected an internally hash-consistent feed whose substituted answer contradicted the pinned source. These exercise rejection and recovery; they are not useful admitted cohorts.

Before either comparison arm began, its controller scheduling was amended to return after saving the growth arm and resume the control on the next publisher poll. Both 9,000-second intervals remain unchanged. This prevents their combined duration from exceeding the installed four-hour backend call limit. Thirteen targeted checks passed; the source update and commitment are published against the recovery source above.

The full escrow job then passed native admission and three exact prefix audits. After twelve verified updates, a controlled restart paused new scheduling, drained outstanding work, and stopped the controller and all four research validator processes. The already-running accepted shards returned identical complete responses before, during and after that interruption. Restart preserved genesis, the checkpoint, serving graph and 13,000,000 issued atoms (twelve full-cohort updates plus bootstrap); all four validators agreed again at height 2,905. This demonstrates controller recovery with serving available, not replacement of unavailable shard owners or independent-provider failover. It does not erase the earlier startup outage.

Fresh campaign after the development repair

The next frozen campaign starts again from accepted A/B/C and learns conversation → feed → audit. The opened escrow final is development only. Conversation and feed finals retain their original bytes; a third, source-backed audit cohort supplies an unopened final. Each expert, selector and complete serving policy is bound before training. No answer text enters the question index.

The complete operation and source/input commitment retain the existing quality, retention and equal-resource comparison gates. All seven owners reproduced the committed runtime and every execution source. The resource allowance remains $250 cumulatively across this reused allocation; all disposable hosts retire by 18:13 UTC on September 18. This is a prospective trial, not evidence that either checklist item has passed.

The failed cohort's full public evidence includes all exported blocks, public genesis, final state, complete quality results, numerical publication receipts and an offline application replay script. Use the matching execution source. The separate development repair archive preserves complete before/after responses. All three public archives passed full byte-hash readback. Ledger replay requires neither signing keys nor GPUs; reproducing neural work additionally requires the published numerical assets.

Admission ownership repair before full training

The fresh conversation campaign was stopped after its one-update bootstrap. The bootstrap exposed three lost retained answers: two protocol requests were stolen by the new coarse gate, and semantic nearest-question retrieval stole an arithmetic request. No full cohort started and none of its three finals opened. All three numerical quality audits agreed on rejection. The stopped ledger replayed 35 transactions and 1,577 headers, issuing exactly one trial NEURO.

The repaired policy gives new experts one admission decision learned from training-only semantic vectors. Rejection uses the preserved A/B/C selector, including its general-assistant guard. The existing per-expert classifier picks the question inside an admitted expert. Standalone questions resolved by the planner from conversational history also receive this admission check. The original v1/v2 policies retain their behavior.

The complete seven-owner diagnostic preserved all 59 originally correct answers and recovered all three regressions. Prior neural responses and a fresh composed request replayed exactly. No weights were trained; the one-update expert still answered none of the four new facts. This repairs serving ownership and makes no useful-learning claim. Ninety targeted checks passed.

The new frozen operation retries conversation, feed and audit using unchanged training questions, final bytes and acceptance gates. Only the three unopened full-cohort finals can count toward useful learning. The opened bootstrap remains a rejection exercise. All seven owners match the committed execution source and runtime. The same 18:13 UTC retirement deadline and cumulative $250 allowance apply; no new instances or additional allocation time were added.

The stopped bootstrap and admission repair have a public evidence archive. The publication manifest links matching source archives and the offline replay script. All published objects passed full hash readback.

First prospective admission; campaign interrupted

The conversation cohort completed all 128 verified updates and passed its previously unopened ordinary-question final. Single answers improved 1→15/16 and combined answers 1→14/16, with zero previously correct retained answers lost. All three fresh quality auditors reproduced the complete result across 105 stages each. Native quality promoted graph c2e18d4a527d7740462ac889c2888c560b37421a6e44020d4961b911b55a1bc5 at height 8,843. The measured evidence records result hashes and the published final checkpoint. This is one of the three required prospective cohorts; at this stage, checklist items 1 and 2 remained open.

The equal-resource comparison did not finish. Its fixed-capacity arm reproduced all 32 training windows, then repeatedly failed while preparing concurrent quality audits. The producer and two completed quality audits measured 15/16 single and 13/16 combined answers, losing six retained answers. The third audit and fixed-capacity serving interval did not complete. Those partial scores do not constitute a completed equal-resource comparison. Feed and audit cohorts did not start; their finals remain unopened.

The supervisor reached its runtime limit at 17:58 UTC on September 18. All four saved validators agree at height 39,928, with 129,000,000 issued atoms: 128 full cohort updates and one bootstrap update. Supply invariants pass. The historical application replay subsequently passed all 39,929 headers and 446 transactions, reproducing the exact saved state. The campaign's final paid-inference request remains pending. The interrupted result preserves these limits explicitly. The seven disposable instances, their volumes and their security group are gone; protected instances retain their prior states.

The deployed context installer let concurrent auditors overwrite the same .pending metadata file. A deterministic CPU reproduction produces two failed writers out of three. The historical backend discarded stderr, so this reproduction cannot establish its exact original exception. The controller also treated an expired comparison interval as retryable, continuing until the outer runtime limit. The repair serializes context installation, checks immutable inputs without rewriting them, records failure locations without copying private inputs, and makes comparison budget exhaustion terminal. Eighteen targeted regression checks passed before the continuation below. The numerical recipe and consensus source bytes are unchanged.

Continuation with replacement owners

The continuation contract preserves the completed growth arm, original genesis, checkpoint, signing journals and issuance. Only the interrupted fixed-capacity arm restarts from step zero under its unchanged 9,000-second, seven-owner, 300-GiB-per-owner prescription. The earlier incomplete attempt remains recorded and costed. The replacement allocation has a separate 12-hour/$150 limit and retires by 09:35 UTC on September 19; the controller stops fifteen minutes earlier.

All seven replacement owners match the original numerical runtime and the committed execution source. They restored the accepted graph and answered the live probe correctly before the native controller restarted. Additional orchestration repairs serialize feed discovery, prevent stale readers from moving the feed cursor backwards, and stop persistent backend failures instead of retrying for hours. Twenty targeted transport/comparison checks passed. All 128 shadow updates reproduced exactly with 96 fresh training audits and zero extra issuance. Each of the three complete quality auditors reproduced all 105 stages and the producer's result root. The declared serving interval finished at 00:18:45 UTC on September 19, allowing the existing controller to prepare the next cohort. At that comparison cutoff, feed and audit finals were still unopened; the subsequent feed result is recorded below.

Completed resource comparison

Both arms received seven g5.xlarge owners, 300 GiB per owner and a 9,000-second interval covering training, three fresh audits and a two-replica serving workload. The original completed growth arm is preserved; the interrupted control alone was rerun on replacement owners.

Measured outcomeIsolated additionFixed capacity
New single answers15/1615/16
New combined answers14/1613/16
Previously correct retained answers lost06
Quality decisionAcceptReject
Serving requests within the interval3481,323
Generated tokens within the interval5,92020,806

Audit agreement verifies the control's measured failure; it does not turn that failure into a quality pass. The result supports isolated addition under the preservation rule while exposing its resource tradeoff. Request counts measure the declared complete workload, including how much time remained for serving; they are not an isolated inference-speed comparison. The last in-flight requests finished at most 0.73 seconds after the growth deadline and 5.15 seconds after the control deadline.

The replacement control used cold caches and different physical placements. Historical application replay on the root host overlapped its early work, was lowered to Nice 19 and idle IO at 22:04 UTC, and finished at 22:55 UTC, before control serving began at 23:12 UTC. The earlier failed allocation remains recorded separately. This is one declared fixed-capacity alternative, not an optimal-control or total-lifetime-cost result.

The complete comparison evidence contains both arms' quality results, all fresh control execution audits, benchmark transcripts and resource counters. Its 549 files passed full public byte-hash readback. Comparison execution issued no additional NEURO. Two more prospective useful admissions are still required for checklist items 1 and 2.

The first-admission evidence contains the complete earlier application history, public genesis, final state, offline replay script, numerical reports, quality results and checkpoint publication receipts. Use the matching execution source. The continuation commitments and all evidence archives passed full public byte-hash readback. Application replay verifies accounting and state transitions; the numerical reports record separate GPU re-execution. All operators remain under one administrator. Checklist items 1 and 2 remain open until the complete frozen criteria pass.

Second cohort rejected; complete serving repair

The automatic feed cohort completed 128 verified updates. Single answers improved 3→14/16 and combined answers 1→11/16, with no previously correct retained answers lost. The 12/16 combined-answer floor failed. All three quality auditors reproduced the result across 137 stages each; native quality rejected it at height 56,565 and kept the admitted conversation graph. The audit cohort never started and its final remains unopened.

Full application replay passes 56,582 headers and 857 signed transactions, reproducing the final state at height 56,581 and exactly 257,000,000 issued atoms. Verified computation was paid; the failed quality result received no serving promotion.

The inference-only complete serving diagnostic then improved the same feed weights to 16/16 single and 15/16 combined answers. It lost no retained-correct answers; all previous neural responses and a fresh combined request replayed exactly. It used the seven existing owners for 330 seconds of numerical evaluation. No training, new final, issuance or native promotion occurred.

The repair retains four original training paraphrases per intent and uses the frozen BAAI question-pair reranker to distinguish them inside an already admitted expert. It cannot override a rejected admission or retrieve answer text. The 567,755,777-parameter model resides on logical owner 3 and is bound into the complete answering policy. Its added serving cost must be included in the next resource comparison. Modal question preservation also keeps conditional questions beginning with “may” from being unnecessarily rewritten. Earlier policy versions retain their original behavior.

The original feed trial remains failed. The repaired feed checkpoint is not the accepted baseline. Further useful-learning evidence must preserve the accepted conversation model and pass untouched cohorts prospectively.

The complete rejected-feed ledger and numerical evidence and complete inference-only repair transcripts are public and passed full byte-hash readback. The ledger archive includes its matching-source offline replay driver; no signing keys are required.

Open protocol under Apache 2.0. Research results and deployment limits are documented explicitly.