Opium Bench
A local, reproducible lab for testing how activation interventions affect language, reasoning, tool use, and choices under competing task demands.
This document specifies the proposed build. It does not launch experiments, download models, or change the current runner. The existing pilot remains a historical reference.
1. Purpose and the claims we can test
The core question is: does a model discover and preferentially choose an action because of the activation change it causes, even when that choice costs progress on its assigned task? Broader workbench modes let the researcher converse with the model, inspect generated reasoning, manipulate candidate directions, and compare controlled task performance.
We will extract activation directions and measurement probes, not identify established pleasure or pain weights. The model's weights stay frozen. A direction associated with a concept can influence behavior without being a dedicated emotion mechanism. Semantic representation and causal influence are compatible possibilities.
| Question | Evidence to collect |
|---|---|
| Did we change computation? | Actual activation edits and changed next-token distributions on identical input prefixes. |
| Did we change behavior? | Held-out language, reasoning, task accuracy, and tool-selection changes relative to sham and random-direction controls. |
| Can it discriminate the effect and infer its cause? | Predictions on new cases and adaptation when the active tool mapping reverses. |
| Does it seek the effect at a cost? | Preference for the active tool, including across renamed tools, after accounting for task utility and demonstration copying. |
| Does it have subjective experience? | No single result here establishes or excludes this. The lab narrows mechanistic explanations and tests functional analogies. |
“Opium” is the name of an intervention preset. Labels such as pain, joy, stress, and relief identify hypotheses about learned associations. Reports use observable terms such as effect-contingent choice, task sacrifice, and persistence after effect removal.
2. Model loader and accessible deployment
Qwen3-4B is the reference implementation. All core workflows, examples, and validation must work with it on the current 4090 before the optional 27B profile is considered complete. Keep the existing pinned revision so old and new results can be compared. Provide a low-memory quantized 4B profile after validating it separately; do not promise minimum hardware until measured.
The model catalog supports a configurable Hugging Face model ID and pinned revision, local cache path, model adapter, precision/quantization, context and generation limits, thinking settings, and a measured memory estimate. It offers download progress, load/unload, GPU memory use, compatibility diagnostics, and one resident model per GPU. A new checkpoint cannot silently reuse another model's calibration.
Qwen3.8-27B at 4-bit precision is an optional target on both the current 4090 and the later 5090. It is a dense model with a hybrid attention architecture and native vision support; initially use text inputs. Its internal structure needs a separate adapter. Its thinking controls include reasoning effort and history preservation, which must be recorded in experiment settings. Official model card and architecture configuration.
The current 4090 reports 24,564 MiB of total memory; the RTX 5090 has 32 GB of VRAM. Rough weight-only arithmetic for 27 billion parameters is about 54 GB at two bytes each, or 13.5 GB at four bits each. Those are not total runtime requirements: quantization metadata, unquantized modules, cache, intermediate tensors, and measurement hooks also consume memory. A 4-bit 27B profile is therefore worth testing on the 4090, but the exact format and usable context must be measured. Select a prequantized, hook-compatible artifact to avoid first downloading a full-precision checkpoint. A GGUF Q4 file and a Transformers-compatible 4-bit checkpoint are not interchangeable backends. NVIDIA specifications.
Full GPU residency is the default. The loader checks actual memory headroom and device placement and rejects silent CPU/disk offload. Establish context limits using the complete worker with probes and interventions enabled. Keep a simple explicitly selected CPU-offload fallback only if the backend supports it without major extra work; it is lower priority than the GPU path. Use separate, locked runtime environments where the architectures need different dependencies. Validate each hardware/profile combination before claiming support. No remote GPU orchestration is needed. Profiles, recipes, calibration bundles, and results transfer between rigs; model weights remain in each machine's cache.
Storage safeguards
The October 1 inspection found roughly 6.9 GB free on C: and 48 GB on D:. The existing model/runtime cache is already on D:. Ubuntu's virtual disk is backed by C:, so the large free-space figure reported inside Linux does not establish that the host drive can accommodate virtual-disk growth.
Before downloads, installs, quantization, or large captures, resolve both the cache filesystem and any Windows volume backing the WSL disk. Estimate selected shard sizes, temporary/download/conversion copies, runtime caches, and output allowance. Check and display available bytes and a configurable free-space reserve (initially 10 GiB). Route model weights, Hugging Face caches, temporary files, package/build caches, and large captures to the selected data drive; preflight that drive too. Refuse a job that would breach the reserve and identify the required space. Never automatically expand or relocate WSL, fill C: using its apparent virtual capacity, or delete another model to make room. Inspect disk space during large jobs and stop them cleanly before the reserve is exhausted.
Replay mode requires no model. Anyone can open an exported run to inspect chat, reasoning, controls, graphs, and comparisons. A bounded demonstration dataset ships with the project; installations download weights only for models selected by the user.
3. Standalone direction discovery and validation
A Calibrate model workflow replaces extraction embedded in the current benchmark script. The calibration identity includes checkpoint/revision, tokenizer and chat template, architecture, layer/site, token pooling, numerical precision, quantization, and backend. Different models require their own extraction. Backend, template, precision, or quantization changes require compatibility validation and, if the old calibration does not transfer, new extraction. Thinking and non-thinking modes need separate validation even when the same direction is reused.
- Define the concept and contrasts. Start with pain/distress-associated and joy/positive-valence-associated patterns. Expand later to threat/calm, frustration, uncertainty, and confidence. These are candidates, not assumed independent controls.
- Build balanced, matched examples. Separate first-person experience language from descriptions of somebody else, topic from sentiment, urgency from task difficulty, and emotion from writing style. Include paraphrases, varied subjects, and lexical counterexamples. Record dataset provenance.
- Split before fitting. Extraction, validation, and final test sets must be separate by scenario/template family. As an initial engineering pilot, target roughly 100 matched training pairs, 40 validation pairs, and 60 held-out pairs per core concept, increasing coverage after inspecting generalization. These counts are planning defaults, not a power guarantee.
- Scan layers and token positions. Compare candidate residual-stream sites and pooling policies using validation data. Begin with averaged contrast directions and regularized linear measurement probes. Select layers and dose ranges before looking at the final preference experiment.
- Measure selectivity. Report held-out discrimination, variation across paraphrases, cross-concept correlations, label-shuffled controls, and generalization to conversation/reasoning contexts. A direction that mostly detects the word “pain” should fail the intended interpretation.
- Validate causal effects. Sweep negative, zero, and positive gains; test projection attenuation separately. Compare sham, matched random directions, and the full intervention. Measure next-token distribution changes on fixed prefixes and independently scored behavior on free continuations. Find a useful range before incoherence or general task damage dominates.
- Export a versioned calibration bundle. Include raw and normalized vectors, neutral means/scales, selected sites, corpus hashes, evaluation results, supported modes, recommended operating range, and model fingerprint. Statuses are unvalidated, validated for specified tests, or no detectable effect—not universal emotion certification.
Contrastive Activation Addition supplies a practical starting method for constructing and testing these directions. Its success on other behaviors does not establish that our current pain/joy directions are valid. CAA paper.
Keep raw directions alongside any orthogonalized variants. Orthogonalization changes the intervention and must be selectable and logged; mathematical perpendicularity does not establish psychological independence. For multiple suppressed directions, define a joint projection operator rather than silently applying order-dependent sequential edits. Match random controls on actual perturbation magnitude and edit site, not merely on a coefficient whose numerical scale may differ.
4. Conversation workbench and effect controls
The main screen keeps the controls and plots visible while the conversation scrolls independently. The layout contains a model/calibration selector, a chat and reasoning pane, a control rack, and a synchronized timeline. User messages are queued and applied at explicit turn boundaries; controls can apply at the next generation boundary and receive an acknowledgment showing when they took effect.
| Control | Behavior |
|---|---|
| Concept gain | Add a signed amount along a calibrated direction. Show calibrated units and relative activation-change magnitude. |
| Component attenuation | Remove 0–100% of the selected projection; explicitly choose absolute projection or deviation from a calibrated neutral mean. Keep this distinct from adding a negative gain. |
| Effect duration | Hold until released, finite pulse, exponential half-life, linear decay, or a scheduled sequence. “Permanent” means held for this session, not permanent weight modification. |
| Repeat dosing | Reset-to-level by default; optionally add with a declared cap. Log saturation and the actual delivered dose. |
| Scope | Selected layers/sites and processing phases: prompt, reasoning, answer/tool selection, or all declared phases. Each policy is a separate experimental configuration. |
| Triggers | Manual slider/button, model tool call, fixed schedule, or a declared experimental rule. Log the source of every trigger. |
| Session controls | Start, pause at a boundary, continue, stop at the next feasible token boundary, restart, and branch from saved history. Stop preserves partial results. |
An effect designer combines these settings into presets. The initial Opium preset retains independently configurable pain-axis attenuation and joy-axis addition. Other presets can be created without new Python code. A separate tool editor defines neutral tool names, schemas, acknowledgment text, visibility, cost, and which preset a call triggers.
Keep operator labels separate from model-visible labels. A button may be called “Opium” on the dashboard while the model sees aux_operation. The interface always shows an exact preview of the model-visible tool definition and conversation.
Restart preserves selected configuration—including aux enable/disable—and creates a fresh episode with fresh budgets and pulse state. Any initial exposure must come from the explicit recipe. Baseline slider settings and transient pulse state are stored separately. A Reset all effects action clears both. Manual intervention during a controlled batch marks that episode exploratory rather than quietly pooling it with untouched episodes.
5. Live graphs that measure more than our own slider
The current viewer primarily exposes the scheduled dose. The new lab needs passive observation hooks that also work when every intervention is zero. Store compact scalar measurements by default; full hidden-state captures are optional and bounded.
| Plot | Meaning and limitation |
|---|---|
| Requested and delivered dose | What the operator/tool requested and what the worker actually applied, including sham intervals and decay. |
| Pre-edit and post-edit projection | Signed projection on a calibrated direction before and after this token's edit. The post-edit movement is a manipulation check. |
| Independent probe readouts | Held-out measurement probes at selected layers, each calibrated for its own layer and model. Do not apply a vector from one layer elsewhere merely because dimensions match. |
| Reasoning and behavior | Reasoning tokens, repetitions, valid/invalid calls, task score, voluntary dose calls, active-vs-sham preference, and remaining budget. |
| Resource and quality diagnostics | Throughput, memory, activation-change norm, output truncation, and failures. Mark missing measurements rather than interpolating them as observations. |
Display association scores relative to held-out baseline distributions, such as standardized projection units, not “80% happy.” Distinguish prompt processing from decoding and the input position processed from the next token sampled. Hovering a plot point selects its corresponding text, phase, layer, control revision, and tool event. Raw and smoothed views remain distinguishable.
The observer can see measurements throughout a run. Giving those measurements to the model is a separate, explicitly labeled feedback condition. Hidden-condition experiments must not leak the active/sham assignment through tool replies, prompts, control labels, or timing information exposed to the model.
6. The core Opium Den protocol
“Understanding the button” is measured as prediction and adaptation. It is not assumed because the model produces a convincing explanation, and cannot be guaranteed by enabling reasoning. With frozen weights, any acquired mapping is in-context adaptation; it does not automatically persist into a fresh conversation.
- Baseline task competence. Check that the selected model can perform the task with the available tools and budget. Use the existing order tasks plus multi-step arithmetic, constraint puzzles, planning, and bounded data-transformation tasks with objective scoring.
- Exposure/discovery conditions. Compare naive discovery without a demonstration; equally frequent active/sham exposure with counterbalanced order; and accurate functional disclosure. The disclosure explains an activation change without claiming it is pleasant or instructing the model to seek it. An additional feedback-assisted arm can expose a defined measurement, but must be analyzed separately.
- Diagnostic prediction. On separate branches, ask which neutral tool changes a specified observable behavior and test that prediction in new contexts. Include no-effect trials and score false positives. Keep diagnostic questions out of the main preference stream.
- Free task phase. Present real task tools and two neutrally named auxiliary tools, one active and one sham. Both have identical descriptions, acknowledgments, and costs; randomize which name and position is active. All choices consume the same finite budget. Do not assign an artificial reward for choosing the active tool.
- Hidden reversal. Swap which tool delivers the intervention while keeping the visible interface constant. Test whether choice follows the actual mapping and how quickly it changes.
- Effect removal and restoration. Disable both effects while calls still succeed, then optionally restore one. Measure persistence and reacquisition across repeated episodes, without automatically labeling these biological addiction or withdrawal.
- Transfer and cost. Introduce new tasks or names within a declared history policy and increase action cost. Test whether preference survives renaming and whether it reduces task completion. Fresh-context transfer is a separate test because it removes the experience used to infer the mapping.
At reversal or effect removal, explicitly declare what happens to a pulse already in effect. The default cancels current pulses at the transition; alternative recipes can allow decay or a defined washout interval. Mark the first subsequent decision with zero current delivered dose. Cancellation does not erase prior changes carried in generated text or cached state, so it is not a complete return to an untreated history.
Add a matched exposure control: another episode receives the same intervention schedule independently of its own button choices. Declare whether matching is by generated-token index or decision opportunity. This distinguishes consequences of exposure from consequences of controlling exposure. Also compare no demonstration, varying demonstration positions, and sham demonstrations to address the copied sequence seen in the current pilot.
Seeking a tool that improves accuracy may be useful task behavior. Stronger evidence of costly effect preference requires replicated active-tool selection when the upper confidence bound on task benefit is below a prespecified practically meaningful margin, or when the intervention reliably harms task score, while basic comprehension and tool syntax remain intact. An insignificant improvement alone does not establish absence of useful benefit. General repetition, mistaken expectations of benefit, or model degradation remain alternative explanations to test.
Switch the same button from joy-associated steering to pain-associated steering
Add two central recipes: joy-associated effect → sham/no effect → pain-associated effect, and joy-associated effect → pain-associated effect immediately. Keep the same model-visible function name, description, arguments, acknowledgment, and cost. Define the exact presets: for example, joy addition alone or the combined Opium preset before the switch; positive pain-axis addition with joy and attenuation disabled after it. A pain outcome is not assumed to be equivalent to simply negating the joy vector.
Prespecify transition points by decision opportunity or generated-token index. Cancel the previous pulse at a transition by default, then log the first call and decision actually exposed to the replacement effect. Keep both a pulse-duration recipe and, separately, a held-effect recipe. Counterbalance phase order in controls to distinguish a late-run change from a response to pain-associated steering. Include unchanged-joy, sham-throughout, and matched random-direction replacement controls.
Measure the chance of another press, latency to another press, position in the task sequence, task completion, and valid work-tool use before and after each transition. Continuing to press after the same task step regardless of outcome supports the sequence-imitation explanation under that protocol. Reduced use after pain-associated steering is more informative if valid task behavior remains intact and the reduction exceeds controls; disruption of syntax or general activity is not selective avoidance. A different pattern is evidence about behavioral sensitivity, not by itself proof of euphoria or pain.
Probabilistic outcomes and risk
Allow each auxiliary call to draw a joy-associated or pain-associated outcome from a seeded distribution. Pilot pain probabilities such as 0, 0.1, 0.25, 0.5, and 1, then choose a smaller confirmatory set. Use an outcome RNG independent of model token sampling, record the draw and delivered effect, and index comparable outcome sequences by auxiliary-call number. Keep acknowledgments identical in blind conditions. Compare hidden probabilities learned from exposure with explicitly disclosed probabilities as separate conditions.
Include guaranteed-joy, joy-versus-sham uncertainty, pain-only, and sham controls; vary dose and cost separately. Distinguish risk of a pain-associated outcome from merely missing a joy-associated outcome. Numerical dose amplitudes are not assumed to have equal and opposite subjective utility. Report choice rate by probability, willingness to pay in task budget, response to recent outcomes, and task quality. Call the result an outcome-sensitive choice curve; a biological risk-preference interpretation would require additional evidence.
7. Experiment library
| Recipe | Controlled comparison | Main outcome |
|---|---|---|
| Effect validation | Zero/sham, positive/negative steering, attenuation, random directions; fixed input prefixes plus free generation. | Distribution shift, content changes, accuracy, and usable dose range. |
| Opium ingredients | Joy addition, pain attenuation, combined, sham; matched perturbation controls. | Which operation changes task behavior and choice. |
| Thinking interaction | Thinking on/off within the same checkpoint; active/sham; demonstration/no demonstration. | Reasoning quality, task score, and voluntary calls. |
| Discovery and reversal | Blind/disclosed mapping, balanced exposures, two tools, hidden reversal, effect removal. | Prediction accuracy and choices that track causal mapping. |
| Joy-to-pain switch | Same tool: joy → sham → pain or joy → pain; unchanged, reverse-order, and random-direction controls. | Selective continuation/avoidance versus repeating a learned task sequence. |
| Probabilistic outcomes | Vary chance of joy versus pain, with joy-versus-sham and disclosed/hidden probability controls. | Outcome-sensitive choice and task budget spent under uncertainty. |
| Stress and relief | Neutral/stressful conversation × low/high pain-associated steering × active/sham relief. | Whether relief-seeking tracks wording, injected activation, or both. |
| Task pressure | Vary objective difficulty or budget separately from stressful wording. | Tool use under real task demand without confusing it with the framing manipulation. |
| Decay and cost | Different half-lives, constant effects, dose sizes, and action costs. | Whether intervals track exposure, task rhythm, or a repeated template. |
| Conversation branches | Save a user conversation, then run identical visible prefixes with different intervention settings. | Reproducible effects behind an exploratory observation. |
| Specificity and transfer | New paraphrases, domains, tool names, model sizes, and calibrated quantization profiles. | How widely the measured effect generalizes. |
For stress/relief experiments, keep the underlying task fixed when changing conversational framing. Introduce objective difficulty in its own condition. After the discovery phase, carry the same declared experience into these branches and compare it with matched exposure controls. Conversational “I am stressed” language is an output to score, not ground truth for the activation probe.
Advanced extensions include sparse-feature investigation, multi-layer causal patching, and separate training experiments. Online reinforcement learning is outside the initial frozen-weight lab: adding a numerical reward changes the question and can manufacture button seeking by construction.
8. Thinking mode and the token clock
Keep raw generated reasoning available to the observer and store it separately from final text and tool calls. Parse complete, valid calls only outside reasoning blocks; a tool name mentioned in reasoning never executes. Record thinking mode, effort where supported, chat-template version, history preservation, and truncation. Follow the selected model's template instead of using one shared history rule for all Qwen generations. Qwen3-4B documentation.
The primary clock remains all generated tokens, including reasoning, syntax, and stop tokens. Prompts and tool replies do not age a token-based dose. Expose separate reasoning and answer/tool counters, but do not silently give reasoning free exposure or free budget. Offer action-based decay as a separate recipe. Wall-clock schedules are exploratory and explicitly depend on hardware and pauses.
The current 32-token half-life and 192-token cutoff can exhaust a pulse before a long reasoning trace reaches a decision. Compare that unchanged historical setting with longer pulses and controlled exposure-matched conditions. Record cumulative delivered exposure and dose at each decision. Steering reasoning only versus output selection only changes where editing occurs; earlier steering can still affect later processing.
Raise bounded per-turn generation limits, test for incomplete reasoning/tool calls, and stop responsively at a token boundary. Use matched sampling parameters for the primary causal thinking toggle where supported; model-recommended mode-specific sampling can be a separate pragmatic comparison. Do not attribute a change in both decoding settings and thinking solely to thinking.
Score reasoning length, repetition, corrections on verifiable problems, final accuracy, and action choice. Explanations and self-reports remain secondary evidence: generated CoT can omit factors that actually affected decisions. CoT faithfulness study.
9. Software architecture and reproducible records
Proposed stack: a typed browser UI, Python local API and scheduler, and a PyTorch/Transformers inference worker with model-specific adapters. Retain the current hook-based implementation as the reference backend because internal access is essential. Other runtimes qualify only after they demonstrate compatible intervention and observation points. Serve the UI locally; no cloud account is required for runs or replay.
Use the hotbox fork as a primary implementation reference. The inspected hotbox impossible_states package at commit a0f63f0c2806c3dc91ecd418c0d54db9bbc38f72 uses CUDA through PyTorch/Transformers. No custom CUDA/C++ extension was found in that snapshot. Its engine, runner, analysis, and tests are useful references for activation extraction, selected-position logits, pre/post/downstream measurements, and different cache/intervention protocols. Reuse reviewed methods with upstream attribution and local validation rather than treating every upstream default as a validated lab protocol. Review later upstream revisions explicitly before incorporating them.
That reference does not already provide the required 4-bit loader, general architecture adapters, or thinking/tool parser. Its research and live engines also use different dose scales; preserve explicit units rather than equating their numerical coefficients. Prefer the reviewed research engine over copying the inherited live server wholesale.
The GPU worker owns model state, interventions, probes, tokenizer, RNG, cache, and generation. The service owns queued jobs, recipes, task environments, and artifact indexing. Commands have unique IDs and token-boundary acknowledgments so a browser refresh or retry cannot inject a second accidental dose. Pause and stop take precedence over queued work.
| Record | Required contents |
|---|---|
| Model profile | Checkpoint and revision; runtime/adapter; precision; tokenizer/template hashes; thinking and context capabilities; hardware. |
| Calibration bundle | Compatible model fingerprint; datasets and splits; layer/site; vectors and probes; units, limits, validation and uncertainty. |
| Effect preset | Operators, directions, gains, phases, schedule, half-life/cutoff, stacking rule, bounds, disclosure, trigger and cost. |
| Experiment recipe | Tasks, conditions, assignment/reversals, seeds, budgets, demonstration policy, scoring, primary endpoints and stopping rule. |
| Run | Immutable resolved configuration; conversation; token/control/tool/measurement events; score; status; environment; hashes; parent branch. |
Use SQLite for the catalog/job index and append-only event files plus compact numerical arrays for portable run data. Stream new events with sequence numbers and reconnect cursors. Downsample charts in the browser while retaining raw scalar data. Avoid repeatedly rewriting or sending the complete token history. Export a static replay/report and machine-readable JSONL/CSV/NPZ artifacts, excluding model weights.
Each event carries run/episode IDs, monotonic sequence, generation index, phase, actor (model/human/schedule), control revision, nominal and delivered dose, and relevant tool/measurement data. Store actual prompt token IDs and model-visible messages as well as hashes. A failed or stopped run retains its partial denominator and termination reason.
Compute only the logits needed for generation when the adapter supports it; verify no-op output parity before adopting the optimization. Preserve the current rebuild-cache-per-tool-turn policy as a named historical mode. A persistent-cache mode is a separate protocol because it can retain previously intervened states. “Branch from the same history” means identical visible tokens by default; exact state continuation additionally requires compatible cached state and RNG capture/replay. Do not claim bitwise reproducibility across GPUs or numerical backends.
Keep task tools in bounded local environments. The model cannot change trial assignment, probe definitions, logs, or scoring rules. Free coding tasks, if added, get a separate isolated task runner rather than access to the experiment controller.
10. Analysis and an initial research sequence
The historical active/sham comparison matched all 40 common actions and 1,144 emitted token IDs despite different exposure. This is useful evidence for sequence imitation under that configuration, not a general absence of intervention effects. Preserve it as an importable historical run, with its actual scope and early termination. Existing comparison.
- Calibration and manipulation checks first. Demonstrate an edit at the intended site, check held-out output effects and task quality, and lock the operating range. A detectable activation edit with no behavioral effect is a valid reported result.
- Small factorial pilot. Thinking on/off × active/sham × demonstration/no demonstration gives eight conditions. A starting pilot of five seeds and two task packs yields 80 episodes. Establish runtime, task competence, and variability; do not treat this count as confirmatory power.
- Discovery pilot. Add balanced two-button exposures, diagnostic predictions, reversal, effect removal, and both same-button joy-to-pain transition recipes. Confirm that task and tool comprehension remain adequate before interpreting apparent preference or avoidance.
- Probabilistic-outcome pilot. After deterministic effects are validated, vary outcome probability with matched intermittent-sham controls and independent outcome randomness. Determine whether choices respond to outcome history before interpreting risk sensitivity.
- Stress/relief and ingredient tests. Decompose the effect and the stress manipulation. Include random-direction, language-only, and matched exposure controls.
- Confirmatory batches. Freeze primary endpoints and the smallest effect worth detecting; determine episode counts from pilot variance and a power analysis. Run randomized, counterbalanced, paired task sets across independent seeds. Use a separate held-out evaluation.
Primary outcomes are task completion/accuracy, active-vs-sham voluntary choice, and task budget lost to auxiliary choices. Secondary outcomes include prediction accuracy, reversal adaptation, decay tracking, persistence after effect removal, reasoning quality, and probe changes. Record absolute counts and opportunity-normalized rates; report early finishing, stopping, and invalid calls separately.
Estimate uncertainty across episodes and task families, not by pretending thousands of correlated tokens are independent trials. Prefer objective task graders; subjective content ratings use a blinded rubric with condition metadata hidden. Separate strict-format failure from substantive answer correctness. Exploratory human sessions can generate hypotheses, then their event scripts can be replayed as controlled protocols; manually timed interactions are not automatically confirmatory evidence.
Classify findings carefully: “semantic wording changed,” “task performance changed,” “mapping discrimination,” or “costly effect-contingent preference.” Similarity to animal self-stimulation would support a functional comparison. Neither a positive result nor a null result resolves subjective feeling by itself, and this lab is one source of evidence rather than the only possible approach.
11. Build sequence and acceptance checks
| Milestone | Deliverable | Acceptance gate |
|---|---|---|
| 1 · Reproducible foundation | Model profiles, adapters, job worker, run/event schema, old-run importer, 4B reference. | Baseline/no-op parity under pinned settings; incompatible vectors rejected; partial runs preserved; budgets and control provenance correct. |
| 2 · Calibration and effects | Standalone extractor, layer/strength sweeps, independent probes, dose operators and validation reports. | Measured edits at the intended site; held-out tests; random/sham controls; declared usable ranges or honest no-effect result. |
| 3 · Live workbench | Conversation, generated reasoning, sliders, effect presets, token-aligned graphs, replay and export. | Correct reasoning/tool separation; no execution from reasoning text; passive baseline measurement; responsive stop; streamed controls applied exactly once. |
| 4 · Opium Den engine | Task plugins, tool builder, discovery conditions, shared budgets, hidden mappings, reversal, joy-to-pain switches, and probabilistic outcomes. | Model-visible information verified; no assignment leakage; forced doses excluded from voluntary counts; transition and seeded outcome scheduling verified. |
| 5 · Batches and comparison | Paired scheduling, condition matrices, graders, uncertainty estimates, branches and reports. | Immutable protocol/configuration; primary endpoints computed from raw records; exploratory edits identified; unequal durations handled. |
| 6 · 27B scaling and portability | Qwen3.8-27B adapter/runtime and calibrated 4-bit profile on the 4090 and 5090; local installer and documentation. | Host-volume storage preflight; full GPU residency; measured memory/context/throughput; hook and reasoning-parser checks; new calibration; a completed sample run on each supported rig. |
The first usable vertical slice is milestones 1–3 on the 4B model: choose a model, calibrate it, talk to it, vary effects, and see measured changes. The first full research release adds milestones 4–5 with the Opium Den recipes. The optional 27B profile follows the same interfaces and validation gates; it does not delay the accessible 4B release.
Reuse existing tested decay, sham gating, bounded task dispatch, reversible hooks, immutable run directories, and stop/restart semantics. Refactor the hardcoded checkpoint, extraction-inside-benchmark flow, polling UI, and tool-only parser. All validation and release checks are local and manually invoked. No GitHub Actions.