EGETHROPIC
The workbench
Opium Bench · Research release v0.3

Design the experiment.
Inspect every choice.

A local browser workbench for testing how activation edits affect generated text, task performance, and voluntary tool use. Model weights stay frozen.

Available now

Edit the choices a model sees, preview its actual prompt, run controlled comparisons, and preserve the complete evidence—including unsuccessful attempts.

Design tools and effects.

The version-2 recipe editor separates the visible tool from its hidden effect. Define names, descriptions, bounded argument schemas, acknowledgments, extra action costs, and outcome probabilities. Save reusable tool and effect libraries. Changes to a saved definition do not rewrite an active run’s frozen recipe.

Effects support signed additions, separate pain-axis attenuation, raw or orthogonalized joy directions, and reset or capped additive stacking. Select prefill, reasoning, or output scope. Constant, finite-pulse, linear, and exponential schedules can use generated-token or completed-decision clocks. Decisions, charged action units, and generated tokens have separate accounting; an unaffordable tool is not dispatched.

Preview the resolved messages and function schemas, then apply the loaded model’s tokenizer template without generating tokens. Recipe changes mark an earlier preview stale. During live runs, stream text and tool calls alongside measured edits and association readouts. Manual changes are recorded as exploratory interventions.

Keep extraction and validation separate.

Fast pilot calibration preserves the original procedure. Research v2 adds an authored 960-example corpus with separate scenario families for training, probe fitting, selection, and held-out evaluation. It supports final, mean, and span pooling; mean and ridge readouts; shuffled controls; and family-level uncertainty estimates.

Extracting a direction does not establish a behavioral effect. A separate validation job selects a dose before held-out intervention tests, comparing signed edits, attenuation, combined edits, sham, and random edits matched to actual perturbation norms. Its results form a new immutable bundle.

The independent-rating workflow exports a blinded scoring sheet and a separate observer key. Import checks preserve the original text, rubric, identities, and hashes. Language ratings, coherence, task scores, and numerical edits remain separate measurements. Blinding requires keeping condition information away from raters; the workflow’s existence is not a claim that independent human ratings have been completed.

Freeze the protocol before running.

The research CLI expands smoke or full matrices into fixed recipes, seed streams, task difficulty and wording, episode order, pairing identities, and stage budgets. A dry run shows the expansion and estimated resource use before launch. Resume uses the unchanged expansion; retrying a failure creates a separate attempt. Analysis exposes first/latest attempt policies and keeps failed, stopped, and missing observations visible.

  • Discovery diagnostics branch from complete saved boundaries with their own token allowance. They score predictions about an explicitly specified observable result, including neither/no-effect and abstention options. Answers do not feed back into the main task. An unvalidated criterion cannot establish discovery.
  • Yoked exposure replays recorded source deliveries against a token or decision clock. The recipient’s own auxiliary calls cannot add exposure. Reports retain timing differences, uncovered intervals, and measured edits; equal coefficients do not guarantee equal internal effects.
  • Transfer can carry a complete visible-history prefix into a fresh task with an explicit reset notice. Continuing the original task is a separate contract that preserves its experiment state.

Templates are starting points, not a sample-size justification. Source-bound stages require their declared evidence and observer definitions; unsupported combinations fail explicitly.

Continue, branch, and share the record.

Strict JSON checkpoints save complete decision or idle conversation boundaries. A new branch can preserve task state, effects, and remaining budget, or explicitly add an allowance while retaining history and progress. The parent remains unchanged. The runtime rebuilds its model cache from messages; this does not restore a serialized KV cache or resume half a tool call.

Portable ZIP and lighter JSON exports can be imported and replayed without a model worker. Checked imports do not overwrite existing evidence, execute bundled code, or activate a calibration automatically. Weights and Python environments are excluded. Continuation additionally requires a complete compatible checkpoint, calibration, model, and runtime; some historical records remain replay-only.

Prepare the runtime and storage.

Reviewing evidence needs standard-library Python. The validated inference path uses Linux/WSL and an RTX 4090: a Qwen3-4B BF16 reference environment and a separate pinned Qwen3.8-27B NF4 environment with its matching calibration. Selecting a model does not install its dependencies or make another checkpoint compatible.

Setup helpers plan explicit environment, cache, and temporary destinations, check shared backing-volume reserves, and retain command logs and receipts. Preparation is a dry plan until execution is requested; network access is a separate opt-in. Managed writes and owned workers are monitored so a storage stop can preserve partial evidence. These controls are not an operating-system quota.

The published Blackfrost Q4_K_M experiments used a separately archived Windows/llama.cpp runtime on an RTX 5090. Current v0.3 has no native GGUF backend and no RTX 5090 release acceptance. Those historical results do not certify this release on that hardware.

What the release checks establish.

718 local Python tests, browser checks, and 168 bundled HTTP replays with ML imports blocked passed. The separate research acceptance archive records bounded RTX 4090 execution paths: 4B protocol, diagnostic, yoke, lifecycle, and portable-evidence stages, plus two 27B compatibility cases. Earlier setup and portable-import failures remain in the archive.

The 4B calibration selected dose zero: no tested nonzero dose met its frozen numerical selection bounds. Its four sham continuations were truncated and remain unrated; the unchanged 0.25 protocol cases were engineering stress checks. No voluntary auxiliary calls occurred in the main acceptance cases, so paid auxiliary presses were not validated on the GPU. These checks establish neither semantic efficacy nor subjective experience. Ratings made after inspecting the public condition-bearing evidence must be described as retrospective and unblinded.

A separate fresh-clone verification completed extraction, a direct active/sham pair, and exact HTTP export/import/replay using new data directories. It reused installed dependencies and cached weights; it did not test a clean-machine installation, a new model download, or continuation with an imported model.

The earlier behavioral studies retain their own models, protocols, and denominators. Engineering acceptance is not pooled with those findings. Base-model comparisons and larger powered behavioral studies remain future research.

Source documentation.

This guide describes source revision 5b1236e, whose implementation is unchanged from eeed6a5. The repository carries the exact commands, formats, supported combinations, and release receipts.