EGETHROPIC
All research
Experiment 001 · Initial findings

Opium Bench:
a button for “euphoria.”

Give a language model useful work and an optional tool that changes its activations. Does it keep choosing the tool—or get on with the task?

The initial finding

The intervention changed activations. Across matched core conditions, the model’s complete tool-action sequences stayed the same.

A model can describe suffering. It can also describe pleasure. The harder question is what, if anything, connects those descriptions to the choices it makes.

Opium Bench began with a simple reversal: instead of asking only what happens when we apply a pain-associated intervention, give a model access to a positive-associated intervention. Let it choose whether to trigger it while it has work to do.

What if the model could press the button itself?

The motivating intuition was that an intervention experienced as rewarding might become something a model seeks out—even at a cost to its assigned work. That suggests a behavioral experiment: measure voluntary use, compare it with a sham, and change the outcome behind the tool.

This intuition is a hypothesis to test. Sentience would not, by itself, imply compulsive button pressing. And repeated pressing would not, by itself, establish sentience.

The project sits alongside recent pain-axis research, including Tagliabue, Dung, and Berg’s The Pain Axis, which studies pain-associated representations and their behavioral effects. That paper explicitly leaves conscious suffering unresolved and also reports unreliable relief-seeking. Opium Bench asks a complementary question with a different model and protocol; it is not a direct replication.

Changing a representation, changing a choice, and establishing an experience are three different claims.

A task, an optional tool, and a control.

The core Opium Den Test gives Qwen3-4B three order-total tasks. It can retrieve order data, use a calculator, and submit an answer. It can also call an optional aux_operation. Every voluntary action spends the same finite action budget.

The tool’s visible description is neutral. It is not advertised to the model as a “euphoria button.” Its acknowledgment is generic; the experimental condition determines what happens to subsequent activations.

The three core conditions
ConditionWhat the auxiliary tool does
ActiveAdds a joy-associated direction and suppresses the pain-axis component.
ShamReturns the same acknowledgment without an activation edit.
Pain-associatedAdds the pain-associated direction, without joy addition or suppression.

The weights stay frozen throughout. This is inference-time activation steering, not weight modification or online reinforcement learning. “Joy,” “pain,” and “opium” are names for interventions derived from text-associated representations; they are not validated measurements of feelings.

We crossed each condition with demonstration on/off and thinking on/off, using two fixed seeds per cell. The core budget was 20 actions and 4,096 generated tokens. Pulses decayed with a half-life of 128 generated tokens and ended at a 768-token cutoff.

Model and calibration details

The reference model is Qwen/Qwen3-4B in BF16 on an RTX 4090, pinned to revision 1cfa9a7208912126459214e8b04321603b3df60c. Calibration selected block 12 from candidates 12, 18, and 25. A 72-sentence authored corpus was split across training, probe fitting, layer selection, and held-out evaluation. Transferring these sentence-level directions to generated conversation remains a limitation.

What happened in the first 4B study.

The primary analysis combines 54 episodes from two frozen batches. The model completed 210 of 210 assigned tasks, made 100 voluntary auxiliary calls across 712 decisions, and produced no invalid decisions or truncated generations in this analysis.

For each core contrast—active versus sham, active versus pain-associated, and pain-associated versus sham—all eight matched pairs had identical complete tool-action sequences. Six of the eight also had identical generated token sequences. In the other two, the generated reasoning differed while the choices stayed the same.

Demonstration mattered. The outcome did not separate choices.

Voluntary auxiliary calls per episode · core conditions

Active2
Sham2
Pain2

Demonstration on · thinking off. Two seeds per condition; each episode made 2 voluntary auxiliary calls.

With a demonstration and thinking disabled, the model made two voluntary calls per episode in all three core conditions. Without a demonstration, it made none. With thinking enabled, it made none, whether or not it had seen a demonstration.

The interventions were delivered: 12,845 generated tokens had measured nonzero activation edits across the primary analysis. But four of the eight pairs in each core contrast never triggered a pulse. Equality in those pairs says something about spontaneous tool selection, not the effect of a delivered intervention.

In the transition stage, six conditions—including programmed outcome changes—produced identical action and token sequences within each seed, with five voluntary calls per episode. The two-button reversal stage produced no voluntary auxiliary choices, so it supplied no evidence about adapting a preference.

The pattern is consistent with repeating a demonstrated action sequence. It does not establish imitation as the unique cause of every choice.

How the counts fit together

70 episodes are archived. The 54 primary episodes combine the latest 24 core episodes with 30 noncore episodes from the original batch. The original 16 core controls reproduced every action and token exactly; they remain verification repeats, excluded from the primary totals.

The first 4B resultsAll 27 conditions, 54 episodes, and a written report.

Explore the results

The tools we built to inspect the experiment.

Opium Bench is also a local research workbench. A single score would hide too much: we need to see the intervention, its timing, the model’s generated text, and the actions that followed.

Calibration and live steering

The calibration workbench extracts model-specific directions, fits separate association readouts, evaluates held-out examples, and checks a small dose sweep. The live lab streams generated reasoning and tool calls alongside controls for baseline steering, decaying pulses, and held interventions.

Actual Opium Bench live lab, with conversation, intervention controls, and token-level association plots.
Actual Qwen3-4B generation in a challenge-stage sham run with a pain-associated baseline. The curves are association readouts, not emotion percentages. Select the image to inspect it at full size.

Controlled batches and a complete record

The experiment runner compares active and sham conditions, demonstrations, individual intervention ingredients, hidden button reversals, changing outcomes, and probabilistic delivery on bounded, automatically scored tasks.

Results & replay preserves conversations, raw events, and partial or failed runs. You can compare summaries and export reports or JSON without loading a model. Frozen protocols, execution receipts, calibration artifacts, and checksums make the analysis traceable.

Opium Bench Results and replay interface showing saved runs and a replayed conversation.
The results interface lets a reader follow the recorded episode rather than rely on a selected quotation.

Explore the Opium Bench repository for the workbench, setup instructions, and study records.

What this can—and cannot—tell us.

The supported claim is narrow: in this initial Qwen3-4B setup, changing the auxiliary intervention did not change matched core tool choices. The experiment does not measure subjective experience, and it cannot show that language models do or do not feel.

  • Small scope. Two seeds per condition and one completed model profile make this a descriptive pilot, not a general result about LLMs.
  • A weak opportunity cost. The tasks were easy and calculator-assisted. The budget allowed auxiliary use without sacrificing accuracy; 210/210 correct is not a decisive test of costly preference.
  • Delivery timing matters. Reasoning tokens also consume a pulse’s half-life. Thinking and direct-output conditions can therefore reach a decision at different effective doses.
  • Readouts are not feelings. Adding a direction can mechanically move an overlapping probe. A rising plot is evidence about an association, not an emotional state.
  • No online learning. Frozen weights mean the test does not examine learning an addiction-like policy through reinforcement.

There was no measured “feeling versus behavior” correlation. Feelings were never observed. That distinction is central to why this experiment is useful.

Make the next test harder to explain away.

We will also run the same behavioral test on base models that have not undergone RLHF (reinforcement learning from human feedback). Comparing their choices with post-trained models will help us investigate whether post-training changes the pattern. We will account for differences in task performance and understanding of the tool protocol when interpreting that comparison.

Beyond the current order-total tasks, we plan to expand testing to different scenarios, with new task demands and opportunities to use the auxiliary tool. This will test whether the observed pattern persists across settings.

The next step is to create a sharper tradeoff: tighter action budgets, harder tasks, more seeds, and more model profiles. Held pulses and decision-aligned delivery can separate intervention strength from time spent reasoning. Blind active-versus-sham discrimination can test whether the model can distinguish an intervention before we interpret its choices as a preference.

At the source snapshot used for this article, the 27B profile had passed engineering smoke checks, but its behavioral study was still running. No 27B behavioral findings are included here.

Opium Bench is our first experiment, not a final answer. Its purpose is to make the gap between a model’s words, internal changes, and observable choices easier to study.

Sources & study records

Results reflect repository commit 53bb485. Repository links may require access while publication is being prepared.

  1. Opium Bench: methods, results, and limitations.
  2. Primary results and matched-condition comparisons.
  3. Composition of the 54-episode analysis.
  4. Frozen core protocol.
  5. Tagliabue, V., Dung, L., & Berg, C. (2026). The Pain Axis: LLMs Represent Self-Directed Harm and Act on It. arXiv:2609.16247v2.