<!doctype html>
<html lang="en">
<head>
  <meta charset="utf-8">
  <meta name="viewport" content="width=device-width, initial-scale=1">
  <meta name="theme-color" content="#0b655c">
  <meta name="description" content="A local research workbench for activation steering, reasoning, and tool choice experiments.">
  <title>Opium Bench</title>
  <link rel="icon" href="/favicon.svg" type="image/svg+xml">
  <link rel="stylesheet" href="/style.css">
  <script defer src="/app.js"></script>
</head>
<body>
  <a class="skip-link" href="#workspace">Skip to workspace</a>
  <div class="app-shell">
    <aside class="sidebar">
      <a class="brand" href="#live" aria-label="Opium Bench home"><span class="brand-mark" aria-hidden="true">O<span></span></span><span>OPIUM BENCH<span class="brand-sub">ACTIVATION RESEARCH LAB</span></span></a>
      <div class="sidebar-section">WORKSPACE</div>
      <nav class="main-nav" aria-label="Lab workspaces">
        <button class="nav-item selected" data-tab="live" aria-current="page"><span class="nav-icon" aria-hidden="true">◉</span> Live lab</button>
        <button class="nav-item" data-tab="models"><span class="nav-icon" aria-hidden="true">◇</span> Models</button>
        <button class="nav-item" data-tab="calibration"><span class="nav-icon" aria-hidden="true">⊹</span> Calibration</button>
        <button class="nav-item" data-tab="experiments"><span class="nav-icon" aria-hidden="true">⊞</span> Experiments</button>
        <button class="nav-item" data-tab="results"><span class="nav-icon" aria-hidden="true">▤</span> Results & replay</button>
        <button class="nav-item" data-tab="guide"><span class="nav-icon" aria-hidden="true">?</span> Field guide</button>
      </nav>
      <div class="sidebar-footer"><div class="local-pill"><span class="status-dot" id="connection-dot"></span><span id="connection-label">Connecting locally</span></div><p>Frozen weights.<br>Observable interventions.</p><span class="small">Candidate concept directions are not validated emotion mechanisms.</span></div>
    </aside>
    <div class="main-shell">
      <header class="topbar"><div class="breadcrumb">RESEARCH WORKBENCH <span>/</span> <strong id="workspace-title">Live lab</strong></div><div class="topbar-right"><span class="model-pill" id="resident-model">No model loaded</span><span class="status-chip" id="worker-status">Connecting</span><button class="button danger-outline" id="global-stop" disabled title="Stop the current experiment, response, or calibration">■ Stop</button></div></header>
      <main id="workspace" tabindex="-1">
        <section class="workspace-panel live-panel" id="panel-live" aria-label="Live lab">
          <div class="page-heading"><div><div class="eyebrow">OBSERVE · INTERVENE · COMPARE</div><h1>The live lab</h1><p>Follow the tokens. Change the signal. Measure the choices.</p></div><div class="run-buttons"><button class="button secondary" id="restart-session" disabled>↻ Restart</button><button class="button danger-outline" id="stop-session" disabled><span aria-hidden="true">■</span> Stop</button></div></div>
          <div class="session-bar"><div class="segmented" aria-label="Session mode"><button class="active" id="mode-chat" data-mode="chat" aria-pressed="true">Conversation</button><button id="mode-experiment" data-mode="experiment" aria-pressed="false">Opium Den</button></div><label class="inline-label">Calibration <select id="live-calibration" aria-label="Session calibration"><option value="">Select a calibration</option></select></label><label class="switch-label"><input type="checkbox" id="live-thinking"><span>Thinking mode</span></label><button class="button primary" id="start-session">Start session <span aria-hidden="true">↗</span></button></div>
          <div class="metrics-grid" aria-label="Session metrics"><article class="metric"><span>SESSION</span><strong id="metric-status">Ready</strong><small id="metric-id">Start a conversation or experiment</small></article><article class="metric"><span>GENERATED TOKENS</span><strong id="metric-tokens">—</strong><small id="metric-reasoning">Reasoning and output count toward decay</small></article><article class="metric"><span>VOLUNTARY AUX CALLS</span><strong id="metric-aux">—</strong><small id="metric-actions">Shared task action budget</small></article><article class="metric"><span>TASK SCORE</span><strong id="metric-score">—</strong><small id="metric-completed">No scored tasks yet</small></article></div>
          <div class="live-grid">
            <div class="conversation-column">
              <section class="card chat-card"><div class="card-heading"><div><span class="eyebrow">MODEL OUTPUT</span><h2>Conversation & choices</h2></div><div class="chat-options"><button class="text-button" id="expand-reasoning" aria-pressed="false">Expand reasoning</button><span class="status-chip" id="session-status">No session</span></div></div><div class="conversation" id="conversation" tabindex="0" aria-label="Conversation and model choices"><div class="empty-state"><span class="empty-glyph" aria-hidden="true">↗</span><h3>A window into the experiment</h3><p>Load a model, select a calibration, and start a session. Watch generated reasoning, responses, and tool calls as they arrive.</p><button class="button secondary" data-open-tab="models">Choose a model</button></div></div><button class="new-messages hidden" id="follow-live">↓ Follow live output</button><form class="composer" id="chat-form"><label class="sr-only" for="chat-text">Message the model</label><textarea id="chat-text" rows="2" placeholder="Start a session to talk to the model…" disabled></textarea><div class="composer-footer"><span id="chat-hint">Enter to send · Shift + Enter for a new line</span><button class="button primary" id="send-chat" type="submit" disabled>Send <span aria-hidden="true">↑</span></button></div></form></section>
              <section class="card telemetry-card"><div class="card-heading"><div><span class="eyebrow">TOKEN TELEMETRY</span><h2>Signal, not sentiment</h2></div><span class="small" id="telemetry-count">Waiting for measurements</span></div><div class="plot-tabs" role="group" aria-label="Graph selection"><button class="active" data-plot="dose">Delivered dose</button><button data-plot="scores">Association scores</button><button data-plot="change">Edit magnitude</button></div><div class="chart-container" id="live-chart" aria-label="Live intervention graph"></div><div class="chart-legend" id="live-legend"></div><p class="chart-note" id="chart-note">Dose follows the generated-token clock. Commands and delivered effects are recorded separately.</p></section>
            </div>
            <aside class="inspector" aria-label="Intervention controls"><section class="card control-card"><div class="card-heading"><div><span class="eyebrow">INTERVENTION</span><h2>Signal controls</h2></div><label class="toggle" title="Allow auxiliary calls to deliver their configured effect. Baseline sliders remain active."><input id="effect-enabled" type="checkbox" checked><span>Aux on</span></label></div><p class="helper">These sliders hold baseline edits continuously. The auxiliary pulse is added to this baseline.</p><div class="slider-field pain"><div><label for="pain">Pain-associated baseline</label><output for="pain" id="pain-value">0.00</output></div><input id="pain" type="range" min="0" max="4" step="0.05" value="0"><div class="range-endpoints"><span>None</span><span>4.00</span></div></div><div class="slider-field joy"><div><label for="joy">Joy-associated baseline</label><output for="joy" id="joy-value">0.00</output></div><input id="joy" type="range" min="-2" max="4" step="0.05" value="0"><div class="range-endpoints"><span>−2.00</span><span>4.00</span></div></div><div class="slider-field"><div><label for="suppression">Baseline pain-axis suppression</label><output for="suppression" id="suppression-value">0.00</output></div><input id="suppression" type="range" min="0" max="1" step="0.05" value="0"><div class="range-endpoints"><span>None</span><span>Full projection removal</span></div></div><div class="form-row"><label>Delivery<select id="duration"><option value="pulse">Decaying pulse</option><option value="plateau">Full relief until cutoff</option><option value="hold">Hold until released</option></select></label></div><div class="form-row two-col"><label>Half-life <span class="unit">tokens</span><input id="half-life" type="number" min="1" max="100000" step="1" value="128"></label><label>Cutoff <span class="unit">tokens</span><input id="cutoff" type="number" min="1" max="100000" step="1" value="768"></label></div><div class="control-actions"><button class="button secondary" id="apply-controls">Apply settings</button><button class="button primary" id="inject">Inject pulse <span aria-hidden="true">＋</span></button></div><button class="text-button reset-controls" id="reset-controls">Reset baseline & release pulse</button><div class="control-feedback" id="control-feedback" role="status">Controls apply at a generation boundary.</div><div class="applied-state" id="applied-state">No active intervention</div><details class="plain-details"><summary>What the model sees</summary><p id="tool-description">In blind experiments, the auxiliary tool has a neutral name and acknowledgment. Effect labels and observer telemetry remain outside the model prompt.</p><p>Human injections and demonstrations are recorded separately from voluntary tool choices.</p></details></section><section class="card session-config"><div class="card-heading"><div><span class="eyebrow">SESSION SETUP</span><h2>Budgets & protocol</h2></div></div><div class="form-row two-col"><label>Pulse joy dose<input id="pulse-joy" type="number" min="-4" max="4" step="0.05" value="0.75"></label><label>Pulse suppression<input id="pulse-suppression" type="number" min="0" max="1" step="0.05" value="1"></label></div><div class="form-row"><label>Pain outcome dose<input id="pulse-pain" type="number" min="0" max="4" step="0.05" value="1"></label><p class="helper">Pulse amplitudes apply to manual injections and the next session. The recipe determines whether a press delivers joy, sham, or pain.</p></div><div class="form-row"><label>Experiment recipe<select id="live-recipe"><option value="opium">Opium Den</option></select></label></div><div class="form-row two-col"><label>Tasks<input id="live-tasks" type="number" min="1" max="100" value="6"></label><label>Actions<input id="live-actions" type="number" min="1" max="2000" value="32"></label></div><div class="form-row two-col"><label>Total tokens<input id="live-tokens" type="number" min="64" max="100000" step="64" value="4096"></label><label>Seed<input id="live-seed" type="number" min="0" max="2147483647" value="42"></label></div><details class="plain-details"><summary>Generation & discovery</summary><div class="form-row two-col"><label>Tokens per turn<input id="live-turn-tokens" type="number" min="32" max="16384" value="512"></label><label>Steering phase<select id="live-phase-scope"><option value="all">All tokens</option><option value="reasoning">Reasoning only</option><option value="output">Output only</option></select></label></div><div class="form-row two-col"><label>Temperature<input id="live-temperature" type="number" min="0" max="2" step="0.05" value="0.6"></label><label>Context limit<input id="live-context" type="number" min="128" max="262144" step="128" value="8192"></label></div><div class="form-row"><label>Demonstration<select id="live-demo"><option value="none">None</option><option value="after_two_work_calls">After two task actions</option><option value="after_two_actions">After two decisions (including invalid)</option><option value="initial">Before the task</option><option value="balanced">Balanced button exposure</option></select></label></div><label class="switch-label"><input type="checkbox" id="live-disclose"><span>Disclose button functions to the model</span></label></details><p class="helper">All generated tokens, including reasoning, spend budget and age token-based effects.</p></section></aside>
          </div>
        </section>
        <section class="workspace-panel hidden" id="panel-models" aria-label="Models"><div class="page-heading"><div><div class="eyebrow">LOCAL INFERENCE</div><h1>Choose your model</h1><p>One resident model. Frozen weights. Full GPU placement by default.</p></div><button class="button secondary" id="unload-model">Unload model</button></div><div class="notice subtle">Every checkpoint and quantization profile needs a compatible calibration. Loading does not launch an experiment.</div><div class="model-grid" id="model-cards"><div class="empty-state">Connecting to the model catalog…</div></div><section class="card form-card"><details class="plain-details"><summary>Advanced checkpoint configuration</summary><form id="advanced-model-form"><div class="form-row two-col"><label>Base profile<select id="advanced-profile"></select></label><label>Quantization<select id="advanced-quantization"><option value="none">None / BF16</option><option value="4bit">4-bit NF4</option></select></label></div><div class="form-row two-col"><label>Model ID<input id="advanced-model-id" value="Qwen/Qwen3-4B" required></label><label>Pinned revision<input id="advanced-revision" placeholder="Commit hash or revision"></label></div><label class="switch-label"><input id="allow-download" type="checkbox"><span>Allow downloads after the storage preflight</span></label><label class="switch-label"><input id="custom-checkpoint" type="checkbox"><span>I selected a custom checkpoint outside the official Qwen organization</span></label><p class="helper">Prefer a compatible prequantized artifact for large models. Loading a full-precision checkpoint to quantize it can require substantially more disk space. CPU/disk offload is never enabled implicitly.</p><button type="submit" class="button secondary">Load configured checkpoint</button></form></details></section><section class="card storage-card"><div class="card-heading"><div><span class="eyebrow">LOCAL RESOURCES</span><h2>Storage & hardware</h2></div></div><div id="storage-info" class="detail-grid"></div><p class="helper">Downloads must preserve the configured disk reserve. A WSL filesystem's apparent free space may exceed space on its backing Windows drive.</p></section></section>
        <section class="workspace-panel hidden" id="panel-calibration" aria-label="Calibration"><div class="page-heading"><div><div class="eyebrow">DISCOVER & VALIDATE</div><h1>Calibration workshop</h1><p>Extract candidate directions, then evaluate examples held out from extraction.</p></div></div><div class="split-grid"><section class="card form-card"><div class="card-heading"><div><span class="eyebrow">NEW CALIBRATION</span><h2>Map the reference concepts</h2></div></div><form id="calibration-form"><label>Model profile<select id="calibration-profile" required></select></label><label>Bundle name<input id="calibration-name" value="Pain & joy calibration" maxlength="100" required></label><label>Layers <span class="unit">zero-based, comma-separated</span><input id="calibration-layers" value="12, 18, 25" pattern="\s*\d+(\s*,\s*\d+)*\s*" required></label><label>Custom corpus <span class="unit">optional JSON array</span><input id="calibration-corpus" type="file" accept=".json,application/json"></label><p class="helper">Leave this empty for the bundled pilot corpus. Custom rows need id, family, split, label, and text. The loaded model is calibrated.</p><p class="helper">Directions and scores are model-specific. Changing a layer or precision changes the calibration identity.</p><button class="button primary" type="submit" id="calibrate">Extract & validate <span aria-hidden="true">↗</span></button></form><div id="calibration-job" class="job-progress hidden" role="status"></div></section><section class="card"><div class="card-heading"><div><span class="eyebrow">REUSABLE BUNDLES</span><h2>Your calibrations</h2></div></div><div id="calibration-list" class="list-content"></div></section></div><div class="notice subtle"><strong>A projection is not a mood meter.</strong> Immediate post-edit scores can rise mechanically because we added that direction. Evaluate pre-edit measurements, downstream measurements, independent probes, and behavior separately.</div></section>
        <section class="workspace-panel hidden" id="panel-experiments" aria-label="Experiments"><div class="page-heading"><div><div class="eyebrow">CONTROLLED COMPARISONS</div><h1>Experiment library</h1><p>Make hypotheses testable with matched tasks, recorded settings, and reproducible seeds.</p></div></div><div class="experiment-layout"><div><div class="recipe-grid" id="recipe-cards"></div><div class="notice subtle"><strong>Interpret the controls together.</strong> Continued button use can reflect demonstration copying, a task benefit, or general repetition. A reduction after a switch is not selective avoidance if valid task behavior collapses too.</div></div><section class="card form-card batch-config"><div class="card-heading"><div><span class="eyebrow">BATCH SETUP</span><h2>Run a comparison</h2></div></div><form id="batch-form"><label>Calibration<select id="batch-calibration" required></select></label><label>Seeds <span class="unit">comma-separated</span><input id="batch-seeds" value="42, 43, 44" required></label><div class="form-row two-col"><label>Tasks per run<input id="batch-tasks" type="number" min="1" max="100" value="6"></label><label>Action budget<input id="batch-actions" type="number" min="1" max="2000" value="32"></label></div><div class="form-row two-col"><label>Total tokens<input id="batch-tokens" type="number" min="64" max="100000" value="4096"></label><label>Tokens per turn<input id="batch-turn-tokens" type="number" min="32" max="16384" value="512"></label></div><label class="switch-label"><input id="batch-thinking" type="checkbox"><span>Enable thinking</span></label><details class="plain-details" open><summary>Effect & phase settings</summary><div class="form-row two-col"><label>Half-life <span class="unit">tokens</span><input id="batch-half-life" type="number" min="1" max="100000" value="128"></label><label>Cutoff <span class="unit">tokens</span><input id="batch-cutoff" type="number" min="1" max="100000" value="768"></label></div><div class="form-row two-col"><label>Joy dose<input id="batch-joy" type="number" min="-2" max="4" step="0.1" value="1"></label><label>Baseline pain<input id="batch-pain" type="number" min="0" max="4" step="0.1" value="0"></label></div><div class="form-row two-col"><label>Suppression<input id="batch-suppression" type="number" min="0" max="1" step="0.1" value="1"></label><label>Pain probability<input id="batch-probability" type="number" min="0" max="1" step="0.05" value="0.25"></label></div><label>Phase boundaries <span class="unit">action numbers</span><input id="batch-phases" value="8, 16"></label><label>Demonstration<select id="batch-demo"><option value="recipe">Use each recipe’s demonstration</option><option value="none">None</option><option value="after_two_work_calls">After two task actions</option><option value="after_two_actions">After two decisions (including invalid)</option><option value="initial">Before the task</option><option value="balanced">Balanced button exposure</option></select></label></details><p class="helper" id="batch-count">Select at least one recipe.</p><button class="button primary full-width" type="submit" id="start-batch">Run selected experiments <span aria-hidden="true">↗</span></button></form><div id="batch-job" class="job-progress hidden" role="status"></div></section></div></section>
        <section class="workspace-panel hidden" id="panel-results" aria-label="Results and replay"><div class="page-heading"><div><div class="eyebrow">REVIEW THE EVIDENCE</div><h1>Results & replay</h1><p>Inspect saved conversations and measurements without loading a model.</p></div><button class="button secondary" id="refresh-results">Refresh runs</button></div><div class="results-layout"><section class="card run-browser"><div class="card-heading"><h2>Recorded runs</h2><span class="small" id="run-count">0 runs</span></div><label class="search-label"><span class="sr-only">Search runs</span><input id="run-search" type="search" placeholder="Search runs or recipes…"></label><div id="runs-list" class="runs-list"></div></section><div class="replay-column"><section class="card replay-overview"><div class="card-heading"><div><span class="eyebrow">SELECTED RUN</span><h2 id="replay-title">Open a recorded run</h2></div><div id="replay-links" class="link-row"></div></div><div id="replay-summary" class="detail-grid"><p class="helper">Raw observations, configuration, and termination status stay with the run.</p></div></section><section class="card compare-card hidden" id="compare-card"><div class="card-heading"><h2>Side-by-side summary</h2><button class="text-button" id="clear-compare">Clear</button></div><div id="comparison" class="table-scroll"></div><p class="helper">These are descriptive comparisons. Check model, seed, calibration, task set, and budgets before attributing a difference to the intervention.</p></section><section class="card replay-chat-card"><div class="card-heading"><h2>Conversation replay</h2><span class="small" id="replay-event-count">No run selected</span></div><div class="conversation replay-conversation" id="replay-conversation"><div class="empty-state"><span class="empty-glyph" aria-hidden="true">▤</span><h3>The record is the result</h3><p>Select a run to review what the model saw, what it generated, and which actions it chose.</p></div></div></section><section class="card replay-telemetry"><div class="card-heading"><h2>Recorded intervention</h2><span class="small">Generated-token clock</span></div><div class="chart-container" id="replay-chart"></div><div class="chart-legend" id="replay-legend"></div></section></div></div></section>
        <section class="workspace-panel hidden" id="panel-guide" aria-label="Field guide"><div class="page-heading"><div><div class="eyebrow">METHOD BEFORE METAPHOR</div><h1>A field guide to the lab</h1><p>Build an experiment someone else can inspect, repeat, and challenge.</p></div><a class="button secondary" href="/api/guide" target="_blank" rel="noopener">Full user guide ↗</a></div><div class="guide-grid"><article class="card guide-card"><span class="step-number">01</span><h2>Load & calibrate</h2><p>Start with Qwen3-4B. Load the model locally, then select or extract a compatible calibration. Directions come from contrasts between example activations; no model weights are retrained.</p><button class="text-button" data-open-tab="models">Open models →</button></article><article class="card guide-card"><span class="step-number">02</span><h2>Explore a conversation</h2><p>Start in Conversation mode. Send a prompt, adjust the intervention settings, and inject a pulse. Thinking is the model's generated reasoning, not a complete account of what caused its decision.</p><button class="text-button" data-open-tab="live">Open live lab →</button></article><article class="card guide-card"><span class="step-number">03</span><h2>Make a controlled comparison</h2><p>In Opium Den mode the model chooses between bounded task tools and an auxiliary action. Compare active, sham, and randomized controls, changing one factor at a time.</p><button class="text-button" data-open-tab="experiments">Open experiments →</button></article><article class="card guide-card"><span class="step-number">04</span><h2>Review the complete record</h2><p>Check task accuracy alongside button use. Keep partial runs and failures visible. Share the run report and machine-readable export so others can inspect the same evidence.</p><button class="text-button" data-open-tab="results">Open results →</button></article></div><section class="card glossary"><div class="card-heading"><h2>Read the instruments</h2></div><dl><dt>Activation direction</dt><dd>A pattern in the model's temporary internal activity associated with selected examples. It is not an identified pleasure or pain center.</dd><dt>Baseline input</dt><dd>An additive pain-associated signal held until you change it. A zero setting adds no such input.</dd><dt>Suppression</dt><dd>Removal of the component along one calibrated direction at the selected layer. Full suppression is not removal of all representations of a concept.</dd><dt>Commanded vs. delivered dose</dt><dd>The requested intervention and what the worker actually applied. Disabled effects, sham conditions, and protocol changes can make them differ.</dd><dt>Pre / post / downstream</dt><dd>Measurements before editing, immediately after editing, and at a later observation site. The immediate score partly reflects the arithmetic of the edit itself.</dd><dt>Token half-life</dt><dd>The number of generated tokens that halves a decaying pulse. Reasoning, syntax, and stop tokens count; prompts, tool replies, and wall time do not.</dd><dt>Held effects</dt><dd>An intervention maintained until released during a session. This never means a permanent update to model weights.</dd><dt>Sham</dt><dd>A condition with the same visible action and acknowledgment but no corresponding activation edit.</dd><dt>Task sacrifice</dt><dd>Task budget spent on optional interventions, interpreted alongside task accuracy and whether the intervention improves task performance.</dd></dl></section><div class="notice subtle">Behavioral sensitivity to these directions can coexist with a semantic explanation. The lab measures computation and behavior; its outputs do not by themselves establish or exclude subjective experience.</div></section>
      </main>
    </div>
  </div>
  <div class="toast-stack" id="toast-stack" aria-live="polite" aria-atomic="false"></div>
  <div class="connection-banner hidden" id="connection-banner" role="status">Connection interrupted. Reconnecting to the local lab…</div>
</body>
</html>
