What's actually built, tagged honestly.
Sourced from 16 recorded founding sessions and a direct inventory of the working codebase as of 2026-08-14. Every item below carries a status tag. Nothing marked Spec'd or Vision is styled to look like something that ships today.
The founding meeting defined one integrated platform.
Two of nine ship today; two more partially. The rest is the roadmap, not a promise of a date.
| # | Pillar | Status |
|---|---|---|
| 1 | Moderation & refereeing | Vision |
| 2 | Reputation system | Vision |
| 3 | Tournament hosting & management | Vision |
| 4 | Education support | Vision |
| 5 | Fact-checking | Live |
| 6 | AI judging | Partial — analysis ships, judging doesn't |
| 7 | Analytics | Live |
| 8 | Scorecard interface | Partial — shipping, mid-rework |
| 9 | Gamification | Vision |
How a debate gets in.
Live Discord voiceLive
Bot joins a channel (/join, /leave, /status). Discord provides one audio stream per speaker, so identity is ground truth rather than inference — the highest-quality path.
WebRTC roomsPartial
Already feeding the system via an external transcriber posting to the same endpoint.
Uploaded recordingsLive
Audio or video file, transcribed and diarized by Deepgram.
Upload from a linkLive
Paste a YouTube or podcast URL; the server fetches and processes it.
Interjection recoveryLive
Diarized turns split on word-level speaker change, not just silence. On a real recording this recovered 49 interjections across 404 words — 22% of words had been filed under the wrong speaker.
Cross-platform captureLive — architecture
The server is capture-agnostic: anything that can POST finalized utterances is a first-class client. The Discord bot is one client of the API, not the system. Purpose-built apps for mobile, Teams, TeamSpeak, or console are Vision — the architecture that would carry them already exists.
The live map, built continuously as people talk.
A tree of one-to-two sentence syntheses.
Live synthesisLive
Every ~30s, new transcript is folded in: extend a nugget, branch a child, start a new root on a topic pivot, or cross-link to a related thread.
Typed nuggetsLive
Five types: claim · position · pushback · question · evidence. Claims carry a checkable flag, true only when there's something concrete a fact-checker could verify.
Tangent detectionLive
Meta-discussion, process disputes, and asides are flagged and collapsible to chips without leaving the tree.
Cross-linksLive
Connects nuggets addressing the same point from different branches.
Board controlsLive
Scan auto-follows the newest nugget. Focus keeps the current thread and its parent/child nuggets expanded while everything else collapses small. Hide tangents collapses meta-discussion to chips. Compact tightens spacing. Arrange switches tree or grid. Fade dims older nuggets by age. Chrono scrub steps back through the board's own history. Positions are draggable and persist; an export freezes the exact layout.
Background analysis during the debate.
Each independently toggleable — different cost profile, different failure mode.
Crux finderLive — default off
Identifies where speakers actually diverge versus what they appear to be arguing about. The product is named after this.
Contradiction watchLive — default off
Flags a speaker contradicting something they said earlier in the same debate, both quotes anchored.
Question tracking / evasionLive — rough
Tracks every question put to a speaker and whether it was ever answered — answered / partial / evaded, escalating on repeat evasion. Currently over-fires and is being tuned; it now has an off switch.
Live positionsLive
A running one-line statement of where each speaker currently stands.
On-demand tools, per nugget, operator-triggered.
Fact-check & academic fact-checkLive
Checks a claim against the web, returns a verdict and sources; the academic variant weights scholarly sources.
SteelmanLive
The strongest honest version of the argument.
ReflectLive
An analytical read on the nugget.
Toulmin analysisLive
Decomposes into claim / evidence / warrant and names the missing pieces — explicitly reporting when no evidence was offered.
Cross-exam questionsLive
Follow-ups that would actually test the claim.
Fallacy checkLive
Against a closed 24-item taxonomy, not a model free-associating: split into formal (structural) and informal (context-dependent), where an informal finding is dropped entirely unless the model supplies a context note justifying it. Every finding carries a steelman pair — the fallacious move and the non-fallacious version of the same point. Findings are user-dismissible; dismissals persist and drop out of all downstream counts.
Analysis desk & Life SupportLive
Two faces of the Lab: the analysis desk is the working surface; Life Support is a full-screen debate vitals monitor — per-speaker lanes, talk-share waveform, live pulse — built for streaming to an audience.
Getting names and attribution right.
AI name inferenceLive
Proposes real names from self-introductions and direct address. Proposals only — nothing applies until confirmed.
Manual renameLive
Sweeps the new label through every already-generated artifact, so nothing is left saying “Speaker 3.”
Attribution repairLive
For diarization errors: an LLM finds suspicious fragments, voiceprint analysis scores whether the audio matches the assigned speaker, and the operator listens and approves. Fully reversible — every batch snapshotted and undoable.
Talk-share statsLive
Live per-speaker time, turns, percentage.
Mid-rework. The site says so on purpose.
GenerationLive
Main claims, evidence quality, fallacies, concessions, unanswered questions, shared agreements, key exchanges, real disagreements — every item anchored to the utterances behind it.
Responsiveness IndexLive — live sessions only
Share of questions put to each speaker that they engaged with. Counted, not model-produced. A partial answer scores half — better than dodging, short of answering.
Room Control IndexLive — both capture paths
Who commanded the floor, independent of whether anything said was sound: talk share, turn length, question pressure. Composure and tone are deliberately not scored.
Escape hatchesLive
insufficient_evidence, format_invalid, inconclusive are reachable and render above everything else.
HTML exportLive
A standalone, fully self-contained file: scorecard, full transcript, and the idea board with your layout preserved. Zero external requests. Openable and shareable forever.
Support IndexSpec’d
Did they ground what they asserted? Claims with real grounding (source, mechanism, worked example, chain of inference, data) over total claims.
Ledger DeltaSpec’d
Concessions extracted minus conceded — and whether the opponent actually cashed in what they extracted.
Format Integrity + confidence veilSpec’d
The comparability checks and diagonal-hatch treatment described on the home page.
Speaker rolesSpec’d
participant · commentator · moderator · audience · narrator — so a post-hoc narrator or a moderator isn't scored as a debater.
Good Faith IndexSpec’d
Built from countable, already-detected events: evasion strikes, self-contradictions, moving goalposts, loaded questions. Distinct from composure, which stays unscored — bad faith leaves evidence, demeanour doesn't.
Verdict layerSpec’d
Per-dimension leaders, named with their counted evidence, plus one overall winner computed from host-set weights agreed before the debate begins — suppressed and replaced with the reason when too few indices are measurable or format-integrity flags fire hard. Per-dimension leaders still show even then.
Two-track headlinePartial
The performance-vs-support scatter described on the home page. The headline computes today, but with two of four indices built it can only speak to responsiveness spread.
Running a room.
Rules & scoring configLive
Per-debate civility level, cussing rules, fact-check limits per debater, and which scorecard metrics are active. Rules are rendered into the analysis prompts and enforced arithmetically in code — point deductions computed in Python, never by the model.
Model selectionLive
Haiku / Sonnet / Opus, with a separate model for the scorecard, so live analysis runs cheap while the post-debate pass runs deep.
Cost trackingLive
Every Claude and Deepgram call recorded with token counts and computed cost, per session and lifetime. Real-time processing runs roughly $1.40–$1.70 per hour-long debate today.
Present modeLive
One toggle strips the interface to what an audience should see.
What's rough, what's not started, and what can drift.
Almost nothing in this category admits its own measurements can drift. Stating it plainly here is the strongest available proof that the principle on the home page is real.
Works well today
Live transcription and diarization; the Nugget Board and its controls; crux and contradiction detection; the entire Lab toolset; speaker naming and attribution repair; cost tracking; the self-contained HTML export.
Rough today
Question/evasion tracking over-fires and is being tuned — it now has an off switch. The scorecard is mid-rework, with two of four indices built.
Not started
The judging and league layer: rubrics, ballots, calibration, standings, tournaments, Elo, badges. Everything in the vision's moderation, reputation, tournament, education, and gamification pillars.
Known limits, stated plainly
Question tracking runs live, so uploaded recordings currently have no Responsiveness score — they show “not measured,” never a zero. Diarization accuracy degrades as speaker count rises on uploads; live capture is unaffected, because each speaker has their own channel. Scorecards produced under different extraction versions are not comparable, and the system version-stamps them rather than silently comparing across a change.
Back to the argument.
The three ideas this is built around — two-track scoring, the confidence veil, and what it refuses to measure.