Skip to content

Ordinary serving diagnostic

The next useful measurement is what the complete planned serving path does on ordinary development questions. It is not another isolated-expert comparison under the explicit First… Second… grammar, and it is not a new independent final.

The prescription reuses the isolated semantic-expert terminal checkpoint 618e3eb0ab52eb688cfa1d17961a124b940ba9267f3e0bc316e151748af7e14c inside the planned graph service. Fifteen development conversations cover retained directory/protocol/parent questions and planner-cohort development facts. Gold standalone questions run on the same frozen service. Scoring labels never enter the request.

Each failure is labeled in order: knowledge if a gold standalone question is already wrong; otherwise decomposition, selection, or assembly from the ordinary response. Isolated additions remain the default later experiment when this screen says knowledge is intact and capacity should grow. Updates and consolidation remain candidates only if they preserve accepted answers. This diagnostic does not elect the planner as the solution.

CPU scoring of recorded traces needs no GPU. The GPU driver is inference-only: five physical GPUs carrying six logical owners ([0,1,2,3,4,4]), one hour of neural serving, a $25 cap, no training, no new final, automatic retirement. The opened semantic final cannot be reused as this test. A passing classification is evidence about which component to change; it is not item 1 or native promotion.

Open protocol under Apache 2.0. Research results and deployment limits are documented explicitly.