User reference / Use the evidence
Interpretation & comparisons
Use the evidence to improve prompts and collaboration while keeping agreement, quality and correctness distinct.
Slow Thinker makes the collaboration inspectable. To improve it, start with a concrete observation: a useful objection disappears, one agent adopts another's structure very early, or another round adds cost without an identifiable substantive change. The report helps locate the evidence; the next experiment tests a proposed explanation.
| Question | Evidence to inspect |
|---|---|
| Who shaped the response?Reviewed E 0.1.5 · A 0.1.4 · R 0.1.16 | Influence matrix, cited step descriptions, process lineage and contributions. Check conceptual continuity yourself where titles change. |
| Did an objection get resolved?Reviewed E 0.1.5 · A 0.1.4 · R 0.1.16 | Original objection, later steps, analyst regressions and process evaluation. Disappearance is not resolution. |
| Did another round help?Reviewed E 0.1.5 · A 0.1.4 · R 0.1.16 | Round assessments, direct proposal comparison, preserved constraints and outcome criteria, then incremental cost/time. |
| Did voting select the strongest proposal?Reviewed E 0.1.5 · A 0.1.4 · R 0.1.16 | Voter justifications, vote margin, fallback/tie rules, blind ranking and independent review. |
| Is a model defending or imitating?Reviewed E 0.1.5 · A 0.1.4 · R 0.1.16 | Several rounds and repeated cases, including what evidence it changes in response to. A single sequence does not establish a stable model trait. |
| What should change next?Reviewed E 0.1.5 · A 0.1.4 · R 0.1.16 | Specific stage prompts, model mix, context budget, voter criteria or an explicitly implemented process variant. |
Keep the case and evaluation criteria fixed when comparing councils. Change one major factor at a time where practical: model combination, reasoning settings, refinement count, voter count or a stage prompt. Preserve the source revision and component versions. Repeat runs to observe variation rather than treating one stochastic output as a reliable effect.
The current history groups exact prompt identities. Different task prompts create different groups. Changes to stage templates are code changes, not task metadata; record their source revision because a configuration key alone does not identify those templates.
Convergence in the title chart means textual stabilization. Convergence in a yellow round badge is an analyst's interpretation. Improvement is a judgment against criteria. Correctness requires evidence appropriate to the challenge. None of these implies the next automatically.
Sustained substantive improvement is a reason to investigate a longer recurrence, including a critic introduced at a chosen stage. It is not proof that recurrence will converge to a successful answer. A stable group can preserve shared mistakes. The current executor has a fixed refinement count and no automatic critic loop or validated quality-based stopping condition.
Show the starting proposals, the final alternatives and the concrete changes that matter to the client's case. Explain which observations are calculated and which are model assessments. Compare the collaboration cost separately from the extra cost of studying it. Include a simpler baseline and a defensible way to judge whether any improvement justifies the additional latency, expense and implementation complexity.
The most useful demonstration is an inspectable comparison with a clear result and known limits. A high vote share, long plan, colourful lineage chart or enthusiastic analyst is not enough by itself.
Source files used for this reference
Reviewed at 7361ea0b71a5.