Skip to content

Ordinary access to the retained experts

This inference-only candidate repairs selection before another learning cohort. The committed prescription binds the complete service, unchanged expert checkpoints, selector inputs and execution source. It uses the same 15 exposed development conversations as the ordinary serving diagnostic. It opens no new final and cannot complete the three-cohort learning milestone.

Request-preservation candidate

The next frozen candidate keeps the measured router and every expert unchanged. A conservative syntax check sends a clear, single-turn standalone question verbatim to the selector; compound, contextual and ambiguous inputs retain neural decomposition. If a decomposed question still contains a referential pronoun, one bounded neural call proposes an explicit subject. Validation permits only pronoun substitutions grounded in the conversation, preserving other words, question count and order. An invalid repair fails closed. This grounding restriction does not prove that the chosen referent is semantically correct.

The complete service binds the policy and its source. Metering reserves the possible repair and charges every actual neural call, including failed repairs; direct requests incur no planner call. This is an experimental English request policy, not a general natural-language parser or an elected permanent planner.

One inference-only allocation uses the same exposed development cases and unchanged scorer. The engineering decision requires zero planning/selection failures in both ordinary requests and standalone controls, retention of all seven previously correct cases, replay and complete-call metering. The original all-answer quality gate remains unchanged: C's three incorrect standalone answers are expected to keep it closed. No next cohort is started. The same five-host, two-hour absolute deadline and $25 cap apply.

The first preservation attempt completed with 9/15 correct, all 21/21 standalone routes correct, exact automatic/forced replay and valid complete-call metering. Five cases still fail knowledge; one fails reference repair because the interpreter copied the unresolved pronoun unchanged. The engineering gate failed. Every previously correct case remained correct.

The subject-resolution revision asks the interpreter only for explicit replacement subjects. Code performs the substitution; the model cannot rewrite, add or drop questions through this interface. The same grounded-edit validator applies. Its one-allocation gate requires preservation of all nine passing cases, with unchanged experts, selector and scoring. It remains an exposed development repair, not a new final.

The subject-resolution revision completed at 10/15, with no selection, decomposition or assembly failures, all seven retained cases passing, 21/21 correct standalone routes and exact automatic/forced replay. Six owners saved identical complete transcripts. The bounded engineering gate passed; the unchanged all-answer quality gate failed on C's three factual errors, affecting five cases. The subsequent C-only learning prescription keeps this answering policy fixed.

The original assembled service used growing router 9df5d8fa with base c4b1a0a8, rather than the earlier ordinary-question base 3766fcc3. Its C gate predicted the correct class on six ordinary C questions but rejected every one on confidence; all six framed versions were accepted. The new gate also overrode a correct protocol route. These are different selection mechanisms.

The candidate retains base 3766fcc3 unchanged. Its new C gate fits unwrapped and framed training questions, with earlier protocol, directory and general examples as negatives. A separate binary guard may return confident general questions to the parent. That guard contains no diagnostic names or answers. The first CPU candidate retained the sky-question error: its earlier general training inputs covered structured records but lacked ordinary explanations. The next candidate adds 80 explicitly recorded general-assistant training questions. Both CPU candidates are retained as development evidence.

The preparer excludes all diagnostic requests, gold questions, already observed planner rewrites and semantic final question strings from fitting after Unicode normalization and punctuation removal. It reads only the earlier selected training split and the exact C training inventory. Its labels initially describe training domains. The outcome-label helper prefers an already correct route, selects C when only C is correct, and supplies no target when both fail. Outcome signals from the development controls are not fed back into this fit.

C remains a fact expert named planner; it is distinct from a query-planner adapter. Decomposition uses the same prompted interpreter. C's input contract now supplies its training-time domain framing to each standalone question, after selection. This does not influence the selector. Automatic and forced controls use the same expert input handling, including directory argument interpretation.

One allocation runs automatic ordinary conversations, automatic standalone gold questions, and explicitly forced-expert standalone controls. Complete prompts, routes, neural outputs and control results are retained. Duplicate standalone requests may reuse a recorded result within this immutable service. All owners must agree; the first automatic response and forced control are replayed.

Only automatic results count as ordinary serving. Forced answers diagnose knowledge and input-contract failures. The access checks require all standalone gold routes, preservation of the previously passing general response, complete ordinary answers and exact replay. A failed check stays failed. No NEURO is issued, no native promotion occurs, and no expert training is permitted. The five-host allocation retains its two-hour absolute shutdown, one-hour inference limit and $25 cap; instances, volumes and its temporary security group are retired after evidence preservation.

Completed result

The five-host allocation ran frozen source 8af5a685c556ac77cddbc0df077333467fa6743e. All six logical owners saved identical complete transcripts for all 15 ordinary conversations and their standalone controls. The strict access gate failed.

Primary diagnostic outcomePrevious serving screenAccess candidate
Passed17
Selection failure120
Knowledge failure05
Decomposition failure23
Assembly failure00

The zero earlier knowledge failures meant the intended experts were not reached; it did not establish their knowledge. Forced controls now answer 10/13 distinct questions correctly: all seven retained questions and three of the six C questions. C confuses a 16-update window with the 4096-step schedule, the 2048-record job limit with 64, and planner_audit with claim_planner. These three atomic errors cause five failed ordinary cases. All six C controls have exactly the same prompt hashes and output token IDs as the matching earlier semantic development responses, including its three wrong answers.

CPU routing succeeds on all 21 gold calls covering 13 distinct questions. The complete automatic gold path reaches exactly one expected expert in 17/21 calls; the others fail planning by adding an unrequested second question or producing an invalid plan. The ordinary two-protocol plan also omits an explicit subject. Those are still failures under the frozen diagnostic. The previously passing general response stays correct. No forced success is counted as an automatic pass, and the overall access gate remains closed.

After collecting every trace and executing its replay calls, the driver failed while aggregating a previously unhandled case: the ordinary response passed but one automatic standalone control failed decomposition. The old scorer returned passed=false with no primary category. CPU recovery assigns that existing failure its gold-control category. Every historical pass/fail flag is unchanged; all six owners' full traces agree. The regression is covered by a test, and the driver now persists replay results before aggregation. This run's replay verdicts were not persisted, so neither first-conversation replay nor forced replay is claimed as passed. The separate comparison with earlier C tokens is recorded independently.

All five instances, volumes and the temporary security group were deleted. The compute upper estimate is $1.16, with storage and transfer separate. The three protected hosts retain their earlier states. No neural weights were trained, no new final was opened, no tokens were issued, and no serving promotion occurred.

The result manifest pins the local evidence archive. The study directory is .neuroshard/ordinary-access-study-20260917b/. It contains the original source, frozen inputs, all owner archives, complete traces, CPU recovery, prior-C comparison and retirement evidence. To rescore an extracted archive, extract source.tar.gz into a source directory, apply scorer-fix.patch there, and run verify_results.py --source /path/to/source. This performs metadata and scoring checks without GPU inference. The evidence has not been uploaded or pushed.

The next learning cohort remains on hold. Selector access improved on this development screen; the knowledge and planning failures still prevent complete ordinary access and do not close TODO items 1 or 2.

Open protocol under Apache 2.0. Research results and deployment limits are documented explicitly.