Agoracrux
Build status

What's actually built, tagged honestly.

Sourced from 16 recorded founding sessions and a direct inventory of the working codebase as of 2026-08-14. Every item below carries a status tag. Nothing marked Spec'd or Vision is styled to look like something that ships today.

LiveBuilt, working, in daily use
Live — roughBuilt and working, but misfires; being tuned
PartialSome of it ships — the text says which
Spec’dDesigned in detail, not built
VisionIntent. No code. Founding-conversation material
The nine pillars

The founding meeting defined one integrated platform.

Two of nine ship today; two more partially. The rest is the roadmap, not a promise of a date.

As of 2026-08-14
# Pillar Status
1Moderation & refereeingVision
2Reputation systemVision
3Tournament hosting & managementVision
4Education supportVision
5Fact-checkingLive
6AI judgingPartial — analysis ships, judging doesn't
7AnalyticsLive
8Scorecard interfacePartial — shipping, mid-rework
9GamificationVision
Capture

How a debate gets in.

Live Discord voiceLive

Bot joins a channel (/join, /leave, /status). Discord provides one audio stream per speaker, so identity is ground truth rather than inference — the highest-quality path.

WebRTC roomsPartial

Already feeding the system via an external transcriber posting to the same endpoint.

Uploaded recordingsLive

Audio or video file, transcribed and diarized by Deepgram.

Upload from a linkLive

Paste a YouTube or podcast URL; the server fetches and processes it.

Interjection recoveryLive

Diarized turns split on word-level speaker change, not just silence. On a real recording this recovered 49 interjections across 404 words — 22% of words had been filed under the wrong speaker.

Cross-platform captureLive — architecture

The server is capture-agnostic: anything that can POST finalized utterances is a first-class client. The Discord bot is one client of the API, not the system. Purpose-built apps for mobile, Teams, TeamSpeak, or console are Vision — the architecture that would carry them already exists.

The Nugget Board

The live map, built continuously as people talk.

A tree of one-to-two sentence syntheses.

Live synthesisLive

Every ~30s, new transcript is folded in: extend a nugget, branch a child, start a new root on a topic pivot, or cross-link to a related thread.

Typed nuggetsLive

Five types: claim · position · pushback · question · evidence. Claims carry a checkable flag, true only when there's something concrete a fact-checker could verify.

Tangent detectionLive

Meta-discussion, process disputes, and asides are flagged and collapsible to chips without leaving the tree.

Cross-linksLive

Connects nuggets addressing the same point from different branches.

Board controlsLive

Scan auto-follows the newest nugget. Focus keeps the current thread and its parent/child nuggets expanded while everything else collapses small. Hide tangents collapses meta-discussion to chips. Compact tightens spacing. Arrange switches tree or grid. Fade dims older nuggets by age. Chrono scrub steps back through the board's own history. Positions are draggable and persist; an export freezes the exact layout.

Live watchers

Background analysis during the debate.

Each independently toggleable — different cost profile, different failure mode.

Crux finderLive — default off

Identifies where speakers actually diverge versus what they appear to be arguing about. The product is named after this.

Contradiction watchLive — default off

Flags a speaker contradicting something they said earlier in the same debate, both quotes anchored.

Question tracking / evasionLive — rough

Tracks every question put to a speaker and whether it was ever answered — answered / partial / evaded, escalating on repeat evasion. Currently over-fires and is being tuned; it now has an off switch.

Live positionsLive

A running one-line statement of where each speaker currently stands.

The Lab

On-demand tools, per nugget, operator-triggered.

Fact-check & academic fact-checkLive

Checks a claim against the web, returns a verdict and sources; the academic variant weights scholarly sources.

SteelmanLive

The strongest honest version of the argument.

ReflectLive

An analytical read on the nugget.

Toulmin analysisLive

Decomposes into claim / evidence / warrant and names the missing pieces — explicitly reporting when no evidence was offered.

Cross-exam questionsLive

Follow-ups that would actually test the claim.

Fallacy checkLive

Against a closed 24-item taxonomy, not a model free-associating: split into formal (structural) and informal (context-dependent), where an informal finding is dropped entirely unless the model supplies a context note justifying it. Every finding carries a steelman pair — the fallacious move and the non-fallacious version of the same point. Findings are user-dismissible; dismissals persist and drop out of all downstream counts.

Analysis desk & Life SupportLive

Two faces of the Lab: the analysis desk is the working surface; Life Support is a full-screen debate vitals monitor — per-speaker lanes, talk-share waveform, live pulse — built for streaming to an audience.

Speaker management

Getting names and attribution right.

AI name inferenceLive

Proposes real names from self-introductions and direct address. Proposals only — nothing applies until confirmed.

Manual renameLive

Sweeps the new label through every already-generated artifact, so nothing is left saying “Speaker 3.”

Attribution repairLive

For diarization errors: an LLM finds suspicious fragments, voiceprint analysis scores whether the audio matches the assigned speaker, and the operator listens and approves. Fully reversible — every batch snapshotted and undoable.

Talk-share statsLive

Live per-speaker time, turns, percentage.

The scorecard

Mid-rework. The site says so on purpose.

Shipped

GenerationLive

Main claims, evidence quality, fallacies, concessions, unanswered questions, shared agreements, key exchanges, real disagreements — every item anchored to the utterances behind it.

Responsiveness IndexLive — live sessions only

Share of questions put to each speaker that they engaged with. Counted, not model-produced. A partial answer scores half — better than dodging, short of answering.

Room Control IndexLive — both capture paths

Who commanded the floor, independent of whether anything said was sound: talk share, turn length, question pressure. Composure and tone are deliberately not scored.

Escape hatchesLive

insufficient_evidence, format_invalid, inconclusive are reachable and render above everything else.

HTML exportLive

A standalone, fully self-contained file: scorecard, full transcript, and the idea board with your layout preserved. Zero external requests. Openable and shareable forever.

Designed, not built

Support IndexSpec’d

Did they ground what they asserted? Claims with real grounding (source, mechanism, worked example, chain of inference, data) over total claims.

Ledger DeltaSpec’d

Concessions extracted minus conceded — and whether the opponent actually cashed in what they extracted.

Format Integrity + confidence veilSpec’d

The comparability checks and diagonal-hatch treatment described on the home page.

Speaker rolesSpec’d

participant · commentator · moderator · audience · narrator — so a post-hoc narrator or a moderator isn't scored as a debater.

Good Faith IndexSpec’d

Built from countable, already-detected events: evasion strikes, self-contradictions, moving goalposts, loaded questions. Distinct from composure, which stays unscored — bad faith leaves evidence, demeanour doesn't.

Verdict layerSpec’d

Per-dimension leaders, named with their counted evidence, plus one overall winner computed from host-set weights agreed before the debate begins — suppressed and replaced with the reason when too few indices are measurable or format-integrity flags fire hard. Per-dimension leaders still show even then.

Two-track headlinePartial

The performance-vs-support scatter described on the home page. The headline computes today, but with two of four indices built it can only speak to responsiveness spread.

Operator tooling

Running a room.

Rules & scoring configLive

Per-debate civility level, cussing rules, fact-check limits per debater, and which scorecard metrics are active. Rules are rendered into the analysis prompts and enforced arithmetically in code — point deductions computed in Python, never by the model.

Model selectionLive

Haiku / Sonnet / Opus, with a separate model for the scorecard, so live analysis runs cheap while the post-debate pass runs deep.

Cost trackingLive

Every Claude and Deepgram call recorded with token counts and computed cost, per session and lifetime. Real-time processing runs roughly $1.40–$1.70 per hour-long debate today.

Present modeLive

One toggle strips the interface to what an audience should see.

Honest current state

What's rough, what's not started, and what can drift.

Almost nothing in this category admits its own measurements can drift. Stating it plainly here is the strongest available proof that the principle on the home page is real.

Works well today

Live transcription and diarization; the Nugget Board and its controls; crux and contradiction detection; the entire Lab toolset; speaker naming and attribution repair; cost tracking; the self-contained HTML export.

Rough today

Question/evasion tracking over-fires and is being tuned — it now has an off switch. The scorecard is mid-rework, with two of four indices built.

Not started

The judging and league layer: rubrics, ballots, calibration, standings, tournaments, Elo, badges. Everything in the vision's moderation, reputation, tournament, education, and gamification pillars.

Known limits, stated plainly

Question tracking runs live, so uploaded recordings currently have no Responsiveness score — they show “not measured,” never a zero. Diarization accuracy degrades as speaker count rises on uploads; live capture is unaffected, because each speaker has their own channel. Scorecards produced under different extraction versions are not comparable, and the system version-stamps them rather than silently comparing across a change.

Back to the argument.

The three ideas this is built around — two-track scoring, the confidence veil, and what it refuses to measure.