Skip to content

Complete answering admission

Status: implemented and locally verified for a new disposable native candidate. The existing deployed chains retain their source, genesis, state and signing keys.

The passing ordinary research service binds expert weights, a learned selector, question handling, input contracts and generation. Native expert admission still promotes the older graph interface. Items 1 and 2 require a candidate's quality evaluation and subsequent serving to execute the same complete system.

Commitment and execution

A native graph may carry an optional answering descriptor. It commits to a content-addressed policy containing the existing planned-service configuration, with its graph references omitted. Materialization derives those references from the graph's model fields, avoiding a circular hash. The full native graph root therefore commits to the backbone, expert checkpoints, tokenizer, router, question handling, input contracts, decoding and execution sources together.

The policy is an immutable object rather than an embedded router matrix. The current three-expert configuration alone occupies about 2 MB; repeating it in native proposals would consume the transaction budget. Every numerical executor must retrieve and verify the complete policy before serving or auditing.

The descriptor also commits to bounded call/token limits and any planner adapter ownership. These are derived from the policy and checked again when it is loaded. Ledger checks bound escrow and compute claims without evaluating neural selection. An execution quorum must reproduce the full response before payment. Existing graph-only profiles keep their current interpretation and prices.

Training claims continue to pay prescribed numerical work only. A separate quality claim can promote the completed graph after the whole job settles. Materializing a checkpoint updates its model commitment while preserving the precommitted policy. Policy updates require a new, reviewed cohort proposal and the same complete-system quality gate; they cannot change an accepted service through local files or runtime flags.

Quality and admission

The new ordinary quality mode executes baseline and candidate on raw committed conversations. Scoring metadata reaches only the scorer. New knowledge requires individual and combined-answer accuracy plus positive paired gain; cumulative retention protects previously correct answers. All broader assistant and conversation anchors are executed. Identical weights alone cannot establish preservation when selection or question handling changes.

The ordinary gate scores ordered short answers and the actual visible reply. Case sensitivity and any alternate answers are frozen scoring metadata, checked against the immutable source and carried forward unchanged. Each retained role also has a positive accuracy floor, so preserving an entirely wrong baseline cannot pass. This gate measures bounded factual, instruction-following and conversation tasks; it does not establish open-ended ChatGPT-level quality.

Native source cursors, exact provenance, document exclusion, actually trained replay identities and accumulated evaluation history remain authoritative. Immutable publication supplies availability and provenance, not semantic truth. Installed curator policy must independently check source correspondence and reject unapproved, repeated, conflicting or contaminated records.

Completion experiment

Before training, commit three new source cohorts, separate ordinary evaluation inventories, fixed broader assistant anchors, training recipes, selector fitting inputs, generation, gates and a capacity rule. The exposed 15-case diagnostic is retention evidence and cannot count as a fresh final. The complete answering policy and every training/evaluation prescription must be committed before the corresponding run; opened failures remain failures.

Run those cohorts through one operated native genesis and the durable publisher controller, using real sharded LLM computation and funded replay. Demonstrate three consecutive useful promotions, publisher restart, rejected data and a rejected candidate while the accepted answering service remains available. No cohort may depend on manually editing genesis, model roots or source cursors.

The capacity rule compares isolated addition with updating existing capacity under a declared equal total-resource allowance. Count training, verification, hosting, serving and transfer for the measured interval. Preserve accepted behavior in either arm. Adding parameters alone is not a success criterion.

These are the existing TODO 1/2 requirements, not additional completion items. Controlled operators can establish the implementation and learning claims; independent ownership and a public provider market remain the separate item 4.

Required verification

  • Reject missing or changed policies, graphs, tokenizer contracts and source bindings before execution; forbid serving a different policy after promotion.
  • Exercise real ordinary, combined and conversational answering through the quality auditor and the subsequent inference path.
  • Check whole-response replay, bounded pricing, unused-escrow refunds and supply conservation; failed quality or unavailable bytes must not promote serving.
  • Replay the operated ledger, compare validators and verify restart does not duplicate jobs, nonces or numerical rewards.
  • Preserve source, checkpoints, prescriptions, inputs, transcripts, cost and cleanup records so the experiment can be reconstructed.

Local implementation evidence

Five actual CPU shard processes execute the committed policy, agree on complete responses, replay ordinary quality including multi-turn retention, and refute a forged visible answer. Their response also exercises native reservation, signed whole-response receipts, funded audit rejection/acceptance, payments and unused escrow refunds without issuing training rewards. The ledger transitions remain tensor-free. A separate source integration publishes ordinary held-out records, prepares and seals the full native job, and rejects substituted scoring labels.

The focused answering, quality, source, graph, lifecycle and metering check passed 67 tests; the subsequent ordinary source-to-admission integration also passed. These use small numerical fixtures to verify the bridge. No three-cohort LLM result or public promotion is implied; TODO items 1 and 2 remain open.

Open protocol under Apache 2.0. Research results and deployment limits are documented explicitly.