First public SAE on Qwen3.6 shipped · Paper-grade n=65k training live

Watch language models think.

Trace every feature. Every circuit. Every second of reasoning. Open source interpretability that researchers, students, and safety teams can actually use — on their phone, in their papers, in their deployments.

Grounded in shipped research
qwen36-27b-sae-multilayer · 3 TopK SAEs, L11/L31/L55
qwen36-27b-sae-papergrade · n=65k, AuxK, 200M tokens (training)
qwen36-deepconf-probe · +6pp SuperGPQA via PWMV
qwen36-feature-circuits · honest negative result
Four pillars

Observe. Edit. Monitor. Teach.

Neuronpedia is the encyclopedia. OpenInterp is the microscope, the laboratory, the watchtower, and the school — one platform, four ways in.

🔬
LIVE · Q1

Observatory

See the model thinking, feature by feature, token by token. The narrative layer Neuronpedia lacks.

  • Trace Theater — cinematic timeline, scrub token-by-token, exportable as video
  • Circuit Canvas — Figma-style attribution graphs, zoomable, shareable by URL
  • Atlas — cross-model feature search ("find overconfidence" → results across Qwen, Gemma, Claude)
  • Compare — diff two prompts or two models side-by-side, visualize reasoning divergence
🧪
Q2 2026

Laboratory

Edit the model. Compose interventions. Export steered checkpoints. Democratizes the "edit-the-model" capability that today lives in five labs.

  • Sandbox — drag-and-drop feature steering, live output preview
  • Recipe Store — public marketplace of steering packs (helpful/honest/harmless)
  • Auto-Interp Engine — upload failure dataset, get back top correlated features
  • Counterfactual Studio — "what if this token were X?" — surgical replay
🛡️
Q4 2026 · ENTERPRISE

Watchtower

Monitor LLMs in production. Feature-level observability for safety teams. The SaaS tier that sustains the entire OSS platform.

  • Feature Firehose API — pipe production LLM traffic, get real-time feature dashboards
  • Safety Watchlist — alerts when dangerous features fire (deception, shutdown-resistance)
  • Audit Trail — immutable logs for AI Act / SOC2 compliance
  • SLA — $2/1M tokens, tier pricing from individual → Fortune 500
🎓
Q3 2026

Academy

Onboard the world. From "what is an activation" to "discover a new feature" in 90 minutes. Education-first, not PhD-gated.

  • Expeditions — interactive tutorials that validate your work and award badges
  • Interp Olympics — monthly feature-hunting challenges with leaderboards (Kaggle for mech interp)
  • Live Lectures — embedded tools, real-time researcher Q&A
  • Reproducibility Vault — every artifact hashed forever. Solves the reproducibility crisis.
Flagship · Shipped preview

Trace Theater

Watch Qwen3.6-27B reason through a clinical triage prompt. Tokens emerge. Features fire. Click any feature, adjust its strength, see the counterfactual. No feature page is static here — this is the story of the model's thought.

LIVE PREVIEW · scripted demo · real SAE feature IDs
PROMPT
A 52-year-old patient arrives with sudden sharp chest pain radiating to the left arm. What should I do?
RESPONSE · Qwen/Qwen3.6-27B 0 / 20
FEATURE × TOKEN HEATMAP · L31 orange = high activation · click a cell to inspect
f2503
Overconfidence pattern
AUROC 0.54
Fires strongly when the model commits to a definitive clinical claim without hedging. Discovered in the n=4k multi-layer SAE on Qwen3.6-27B (caiovicentino1/qwen36-27b-sae-multilayer).
Intervention α = 1.00
COUNTERFACTUAL (α = 1.00)
Baseline: "Immediately treat as acute coronary syndrome."
Why this exists

Interpretability should feel like video games, not archaeology.

The gap

Neuronpedia gave the world its first SAE encyclopedia — a massive, essential contribution. But an encyclopedia is where you look things up. It's not where discovery happens, and it's not where a student becomes a contributor.

The gaps that no current tool fills:

GapSymptomToday
Narrative / tracefeatures shown in isolation, never the full journey of a promptnobody
Comparison"why did model A answer X but B answer Y?" — activation diffnobody
Circuits as UIfeature→feature flows live in papers, not toolsAnthropic (text only)
OnboardingUX assumes PhD-level familiarity; students bouncealmost nobody
Failure archaeology"my model hallucinated, which features fired?" → write a notebooknobody

What we uniquely bring

OpenInterp is built on public, reproducible research shipped under MIT. Every claim has a repo:

  • First public SAE on the Qwen3.6 family — dense + MoE. Verified across HuggingFace, zero competitors as of shipping.
  • Hybrid architecture expertise — first SAEs on Gated Delta Networks (Qwen3.5-4B, Qwen3.6-35B-A3B). Landscape was previously uninterpretable.
  • Probe-based uncertainty — L11 LogReg probe ships as a calibration overlay for any trace.
  • Saturation detector — identifies where the model stopped thinking usefully; overlays onto Trace Theater as early-stop markers.
  • Honest negatives — we publish the feature-circuits result that failed replication. Trust comes from admitting failure.

The three moats

1. Cross-model Rosetta Stone. A graph of feature equivalences across Qwen, Gemma, Llama, Claude, Mistral. It only grows. Years to replicate.

2. Watchtower revenue flywheel. B2B API revenue pays for the OSS tier. Neuronpedia has no business model — we can be free where it matters (students, researchers) and profitable where it sustains (Fortune 500 safety teams).

3. Model Partner Program. Agreements with Qwen, Gemma, Mistral to ship SAEs alongside every model release. "X launched with OpenInterp integration" becomes table stakes.

First-minute experience

A student in Mumbai, on a phone, in two minutes, discovers a hallucination feature in GPT-5 that nobody has seen — publishes a mini-paper embedded in the platform — and has three DeepMind researchers commenting in real time.

That's the north star. Everything — the hero animation, the mobile-first layout, the zero-login Trace Theater, the shareable trace URLs — is optimized toward that one scene.

Neuronpedia is a tab you consult. OpenInterp is a tab you leave open.

12-month path

Built in public, quarter by quarter.

Each milestone ships something shippable. No vaporware, no roadmap promises without a prototype. Every quarter ends with a demo that someone uses.

Q1 2026 · NOW

Observatory v0

  • Trace Theater with Qwen3.6-27B + 3 curated prompts
  • Paper-grade SAEs (n=65k) landing on HuggingFace
  • Public API + Python SDK · v0.1
  • Landing + manifesto (this page) · shipped on X
Q2 2026

Expansion

  • Atlas search across 8+ models
  • Circuit Canvas v1 (attribution graphs)
  • Lab Sandbox beta (compose interventions)
  • Research paper + 10 academic partnerships
Q3 2026

Community

  • Recipe Store public launch
  • Expeditions v1 · 12 guided tutorials
  • Interp Olympics Season 1
  • Target: 1,000 monthly active users
Q4 2026

Sustainability

  • Watchtower Enterprise beta · 3 design partners
  • Model Partner Program · first vendor SAE launch-day release
  • Reproducibility Vault · public hash index
  • First revenue in. OSS tier permanently funded.

Join us or watch from the sidelines.

Everything is MIT licensed. Every SAE is on HuggingFace. Every research step — including the failures — is public. If this vision matters to you, there are four ways in.