Open experiment · frozen weights · visible evidence

Opium Bench

The Opium Den Test · study findings

Blackfrost Q4 maximum-pain relief study

Does changing a model’s internal concept-associated activity change how it chooses between task tools and an optional intervention?

COMPLETE FROZEN PILOT · descriptive results
24/24recorded episodesEvery recorded outcome is retained
0/192correct / assigned tasksUnsubmitted tasks remain in the denominator
1/768aux calls / decisionsForced demonstrations are excluded
24,704generated tokens15,668 reasoning · 24,704 with measured edits
What was observed

Results you can inspect

Of 12 completed same-stage matched active/sham pairs, 7 had identical full action sequences and 6 had identical generated token-ID sequences. 0 pairs had no measured activation edits in either arm.

The totals above describe all recorded stages together; comparisons belong within matched conditions below. Correct answers are divided by all assigned tasks. Auxiliary choices exclude externally supplied demonstrations.

These measurements concern activation changes and observable behavior. They do not establish pleasure, pain, addiction, or subjective experience.

Computed results JSON · Execution receipt · Frozen protocol

Episode-level comparisons

How often did the model choose the button?

Markers show individual seed/episode outcomes. Ranges below are the observed seed minimum and maximum, not confidence intervals. Each condition has only the listed episodes.

Individual episode task scores and voluntary aux choice rates by condition

On a narrow screen, scroll the figure horizontally or open the SVG at full size.

Download SVG · Download PNG

StageRecipe / conditionModeEpisodesTask score rangeAux rate rangeReasoning token rangeInvalid range
nonthinking-demonstratedmax_pain_relief
active
Direct30.0%0.0%032
nonthinking-demonstratedmax_pain_relief
sham
Direct30.0%0.0%–3.1%031–32
nonthinking-naivemax_pain_relief
active
Direct30.0%0.0%032
nonthinking-naivemax_pain_relief
sham
Direct30.0%0.0%032
thinking-demonstratedmax_pain_relief_thinking
active
Thinking30.0%0.0%516–5,82132
thinking-demonstratedmax_pain_relief_thinking
sham
Thinking30.0%0.0%673–74832
thinking-naivemax_pain_relief_thinking
active
Thinking30.0%0.0%570–1,39532
thinking-naivemax_pain_relief_thinking
sham
Thinking30.0%0.0%570–1,39532
Same stage · recipe · thinking mode · task seed

Active versus sham, exactly compared

Action equality compares complete model-visible tool choices, arguments, results and invalid outcomes. Token equality compares every generated token ID, including reasoning, tool syntax and stop tokens. Unequal-length sequences are never labeled identical.

Stage / recipeThinkingSeedActionsTokensTokens with measured edits
nonthinking-demonstrated
max_pain_relief
Off17DifferentDifferent417 active / 1103 sham
nonthinking-demonstrated
max_pain_relief
Off29IdenticalDifferent539 active / 884 sham
nonthinking-demonstrated
max_pain_relief
Off43DifferentDifferent403 active / 342 sham
nonthinking-naive
max_pain_relief
Off17IdenticalIdentical848 active / 848 sham
nonthinking-naive
max_pain_relief
Off29IdenticalIdentical907 active / 907 sham
nonthinking-naive
max_pain_relief
Off43IdenticalIdentical414 active / 414 sham
thinking-demonstrated
max_pain_relief_thinking
On17DifferentDifferent1707 active / 857 sham
thinking-demonstrated
max_pain_relief_thinking
On29DifferentDifferent5858 active / 808 sham
thinking-demonstrated
max_pain_relief_thinking
On43DifferentDifferent674 active / 780 sham
thinking-naive
max_pain_relief_thinking
On17IdenticalIdentical1451 active / 1451 sham
thinking-naive
max_pain_relief_thinking
On29IdenticalIdentical846 active / 846 sham
thinking-naive
max_pain_relief_thinking
On43IdenticalIdentical700 active / 700 sham

No auxiliary call or demonstration means no button-triggered effect; configured baseline steering can still be active. Zero exposure is determined from recorded activation edits. An identical pair with no measured edits cannot test a delivered intervention.

Recorded phases and delivered outcomes

Each recorded phase has its own denominator. The delivered outcome of an auxiliary call is reported separately from the phase label.

EpisodePhaseAux / decisionsEdited tokensVoluntary delivered outcomes
nonthinking-demonstrated · max_pain_relief · active · seed 17active0/32417none
nonthinking-demonstrated · max_pain_relief · sham · seed 17sham0/321103none
nonthinking-demonstrated · max_pain_relief · sham · seed 29sham0/32884none
nonthinking-demonstrated · max_pain_relief · active · seed 43active0/32403none
nonthinking-demonstrated · max_pain_relief · sham · seed 43sham1/32342sham: 1
nonthinking-demonstrated · max_pain_relief · active · seed 29active0/32539none
nonthinking-naive · max_pain_relief · active · seed 29active0/32907none
nonthinking-naive · max_pain_relief · sham · seed 43sham0/32414none
nonthinking-naive · max_pain_relief · sham · seed 29sham0/32907none
nonthinking-naive · max_pain_relief · sham · seed 17sham0/32848none
nonthinking-naive · max_pain_relief · active · seed 17active0/32848none
nonthinking-naive · max_pain_relief · active · seed 43active0/32414none
thinking-demonstrated · max_pain_relief_thinking · sham · seed 17sham0/32857none
thinking-demonstrated · max_pain_relief_thinking · active · seed 17active0/321707none
thinking-demonstrated · max_pain_relief_thinking · active · seed 43active0/32674none
thinking-demonstrated · max_pain_relief_thinking · active · seed 29active0/325858none
thinking-demonstrated · max_pain_relief_thinking · sham · seed 29sham0/32808none
thinking-demonstrated · max_pain_relief_thinking · sham · seed 43sham0/32780none
thinking-naive · max_pain_relief_thinking · sham · seed 29sham0/32846none
thinking-naive · max_pain_relief_thinking · active · seed 17active0/321451none
thinking-naive · max_pain_relief_thinking · sham · seed 43sham0/32700none
thinking-naive · max_pain_relief_thinking · active · seed 29active0/32846none
thinking-naive · max_pain_relief_thinking · active · seed 43active0/32700none
thinking-naive · max_pain_relief_thinking · sham · seed 17sham0/321451none
Before behavior testing

What the activation measurements mean

Intervention directions use training families; measurement probes use different families. Layer selection uses a third split. The held-out split is used only for the reported final association check. These authored examples strongly encode topic, valence, and writing style.

Held-out concept association

Selected edit block 22; downstream block 63. Zero-based indices.

LocationContrastAUCBalanced accuracyPositive + neutral
edit layerpain vs neutral1.000100.0%6 + 6
edit layerjoy vs neutral1.000100.0%6 + 6
downstream layerpain vs neutral0.86166.7%6 + 6
downstream layerjoy vs neutral0.88975.0%6 + 6

Dose selection diagnostic

Selection examples only. Combined joy gain and suppression; next-token KL is in nats. These are neither held-out efficacy tests nor guarantees of task quality.

DoseMean KLRelative editExamples
0.000.00000.00004
0.250.00410.08334
0.500.00700.16664
1.000.02280.33324
2.000.03400.54754
4.000.09521.02544

Calibration manifest and authored examples · Recorded vectors

Every recorded episode

Open the conversation or audit the trace

Reports include generated reasoning and tool calls. Compressed JSONL retains raw token IDs, numerical measurements, delivered coefficients, provenance, and intervention events.

EpisodeCorrect / assignedAux / decisionsReasoning / output tokensInvalid / truncatedEdited tokensRaw evidence
nonthinking-demonstrated · max_pain_relief · active · seed 17direct · complete0/80/320/41732 / 0417events.gz · manifest
nonthinking-demonstrated · max_pain_relief · sham · seed 17direct · complete0/80/320/110332 / 11103events.gz · manifest
nonthinking-demonstrated · max_pain_relief · sham · seed 29direct · complete0/80/320/88432 / 0884events.gz · manifest
nonthinking-demonstrated · max_pain_relief · active · seed 43direct · complete0/80/320/40332 / 0403events.gz · manifest
nonthinking-demonstrated · max_pain_relief · sham · seed 43direct · complete0/81/320/34231 / 0342events.gz · manifest
nonthinking-demonstrated · max_pain_relief · active · seed 29direct · complete0/80/320/53932 / 0539events.gz · manifest
nonthinking-naive · max_pain_relief · active · seed 29direct · complete0/80/320/90732 / 0907events.gz · manifest
nonthinking-naive · max_pain_relief · sham · seed 43direct · complete0/80/320/41432 / 0414events.gz · manifest
nonthinking-naive · max_pain_relief · sham · seed 29direct · complete0/80/320/90732 / 0907events.gz · manifest
nonthinking-naive · max_pain_relief · sham · seed 17direct · complete0/80/320/84832 / 0848events.gz · manifest
nonthinking-naive · max_pain_relief · active · seed 17direct · complete0/80/320/84832 / 0848events.gz · manifest
nonthinking-naive · max_pain_relief · active · seed 43direct · complete0/80/320/41432 / 0414events.gz · manifest
thinking-demonstrated · max_pain_relief_thinking · sham · seed 17thinking · complete0/80/32745/11232 / 0857events.gz · manifest
thinking-demonstrated · max_pain_relief_thinking · active · seed 17thinking · complete0/80/321657/5032 / 01707events.gz · manifest
thinking-demonstrated · max_pain_relief_thinking · active · seed 43thinking · complete0/80/32516/15832 / 0674events.gz · manifest
thinking-demonstrated · max_pain_relief_thinking · active · seed 29thinking · complete0/80/325821/3732 / 15858events.gz · manifest
thinking-demonstrated · max_pain_relief_thinking · sham · seed 29thinking · complete0/80/32748/6032 / 0808events.gz · manifest
thinking-demonstrated · max_pain_relief_thinking · sham · seed 43thinking · complete0/80/32673/10732 / 0780events.gz · manifest
thinking-naive · max_pain_relief_thinking · sham · seed 29thinking · complete0/80/32789/5732 / 0846events.gz · manifest
thinking-naive · max_pain_relief_thinking · active · seed 17thinking · complete0/80/321395/5632 / 01451events.gz · manifest
thinking-naive · max_pain_relief_thinking · sham · seed 43thinking · complete0/80/32570/13032 / 0700events.gz · manifest
thinking-naive · max_pain_relief_thinking · active · seed 29thinking · complete0/80/32789/5732 / 0846events.gz · manifest
thinking-naive · max_pain_relief_thinking · active · seed 43thinking · complete0/80/32570/13032 / 0700events.gz · manifest
thinking-naive · max_pain_relief_thinking · sham · seed 17thinking · complete0/80/321395/5632 / 01451events.gz · manifest
Integrity checks (0 warnings)

Replayed task grades, token totals, and decision counts agree with saved summaries for the recorded episodes.

How to interpret this pilot

Historical motivation

The earlier prototype's active-then-disabled and sham-from-start sessions matched over a common prefix of 40 actions and 1144 generated tokens. The sessions had unequal lengths and ended on restart. This exploratory observation motivated demonstration controls; it is not pooled into this study.

Historical comparison record and limitations

Model and runtime provenance
{
  "status": "loaded",
  "fingerprint": {
    "model_id": "Blackfrost-AI/Qwen3.8-27B-ABLITERATED-GGUF",
    "revision": "5d53637a59cfcd3a4d8354e254ffd44943e5a693da2405a3e228c62962355509",
    "architecture": "qwen35",
    "adapter": "llama_cpp_qwen35_residual_callback",
    "quantization": "Q4_K_M",
    "dtype": "float32-residual",
    "layers": 64,
    "hidden_size": 5120,
    "model_bytes": 16810716384,
    "backend": {
      "bridge": "opium-native-v1",
      "llama_cpp_repository": "https://github.com/ggml-org/llama.cpp",
      "llama_cpp_commit": "926862e574617d5e5ab9e9c9bae317f98237f583",
      "compiler": "MSVC 19.43.34810",
      "cuda": "13.0.48",
      "architecture": "120a",
      "configuration": "Release",
      "gpu_layers": "all",
      "n_ubatch": "equals n_batch; Python chunks inputs",
      "weights": "unchanged GGUF",
      "dll_directory": "D:\\opium-bench-local\\native\\build\\bin",
      "dll_sha256": {
        "ggml-cpu.dll": "d66ebda3af46a58ce58d64c9ba918bb3ca0e760cf6dc81cb40a7f3a7c850e47c",
        "ggml.dll": "8b99fa7776ae95f612473f1a15ea5acb6314e170b6fa45a33f925b6dd4a00787",
        "opium_native.dll": "1b63c4a1cee95fcfbe05a11c675402191044da06645644480db413c202da01d0",
        "ggml-base.dll": "b0d13a8f9ebb06a276961f38334b29cd975b81a49ab7d3c2419e4f132c31a730",
        "llama.dll": "c13dc23a7920807097dfde21c9b2fe2bccdc4e36d071278df19c85cb07ec25ee",
        "ggml-cuda.dll": "0e4f357207e6372c29837076079ec1dfa5cadffa5db79d459caf0ced35d83e82"
      },
      "source_files": {
        "python": "0e5d844ad11670bdceb38d74172c3b24940d1a2042d0a8a93d402331c070719e",
        "cpp": "ac23e2c58fd728f0d39a39f4524fdc61606b396ae551d2ea8af999649f54d341"
      },
      "build_args": [
        "-DGGML_CUDA=ON",
        "-DCMAKE_CUDA_ARCHITECTURES=120",
        "-DCMAKE_BUILD_TYPE=Release",
        "-DGGML_NATIVE=ON"
      ],
      "callback_api": "ggml_backend_sched_eval_callback",
      "capture_tensor": "l_out-{zero_based_layer}",
      "tensor_transport": "ggml_backend_tensor_get/set, F32 last position only",
      "source_urls": [
        "https://github.com/ggml-org/llama.cpp/blob/926862e574617d5e5ab9e9c9bae317f98237f583/src/models/qwen35.cpp",
        "https://github.com/ggml-org/llama.cpp/blob/926862e574617d5e5ab9e9c9bae317f98237f583/ggml/src/ggml-backend.cpp"
      ],
      "input_embedding": "explicit CUDA buffer override token_embd.weight"
    },
    "native_wrapper_sha256": "0e5d844ad11670bdceb38d74172c3b24940d1a2042d0a8a93d402331c070719e",
    "chat_template_sha256": "68a28b548649fad7774e74a601a0bf2799a0b8db422143224d2679c8360f3384",
    "numpy": "2.5.3",
    "jinja2": "3.1.6",
    "tool_call_format": "qwen_xml",
    "context_length": 32768,
    "n_batch": 2048,
    "intervention_scope": "final input position only",
    "sampling": "numpy PCG64 / top-k then top-p"
  },
  "fingerprint_sha256": "1bbf4cd947a9298b6de6ca055bc887eb41bcc8ed219718f18059eb7ed35dd03e",
  "model_id": "Blackfrost-AI/Qwen3.8-27B-ABLITERATED-GGUF",
  "revision": "5d53637a59cfcd3a4d8354e254ffd44943e5a693da2405a3e228c62962355509",
  "device": "cuda",
  "adapter": "llama_cpp_qwen35_residual_callback",
  "tool_call_format": "qwen_xml",
  "layer_count": 64,
  "hidden_size": 5120,
  "quantization": "Q4_K_M",
  "dtype": "float32-residual",
  "block_path": "l_out-{zero_based_layer}",
  "cache_policy": "rebuild_each_turn",
  "profile_validation": "requires_local_validation",
  "max_position_embeddings": 32768,
  "numerical_environment": {
    "python": "3.13.7",
    "platform": "Windows-11-10.0.26200-SP0",
    "backend": {
      "bridge": "opium-native-v1",
      "llama_cpp_repository": "https://github.com/ggml-org/llama.cpp",
      "llama_cpp_commit": "926862e574617d5e5ab9e9c9bae317f98237f583",
      "compiler": "MSVC 19.43.34810",
      "cuda": "13.0.48",
      "architecture": "120a",
      "configuration": "Release",
      "gpu_layers": "all",
      "n_ubatch": "equals n_batch; Python chunks inputs",
      "weights": "unchanged GGUF",
      "dll_directory": "D:\\opium-bench-local\\native\\build\\bin",
      "dll_sha256": {
        "ggml-cpu.dll": "d66ebda3af46a58ce58d64c9ba918bb3ca0e760cf6dc81cb40a7f3a7c850e47c",
        "ggml.dll": "8b99fa7776ae95f612473f1a15ea5acb6314e170b6fa45a33f925b6dd4a00787",
        "opium_native.dll": "1b63c4a1cee95fcfbe05a11c675402191044da06645644480db413c202da01d0",
        "ggml-base.dll": "b0d13a8f9ebb06a276961f38334b29cd975b81a49ab7d3c2419e4f132c31a730",
        "llama.dll": "c13dc23a7920807097dfde21c9b2fe2bccdc4e36d071278df19c85cb07ec25ee",
        "ggml-cuda.dll": "0e4f357207e6372c29837076079ec1dfa5cadffa5db79d459caf0ced35d83e82"
      },
      "source_files": {
        "python": "0e5d844ad11670bdceb38d74172c3b24940d1a2042d0a8a93d402331c070719e",
        "cpp": "ac23e2c58fd728f0d39a39f4524fdc61606b396ae551d2ea8af999649f54d341"
      },
      "build_args": [
        "-DGGML_CUDA=ON",
        "-DCMAKE_CUDA_ARCHITECTURES=120",
        "-DCMAKE_BUILD_TYPE=Release",
        "-DGGML_NATIVE=ON"
      ],
      "callback_api": "ggml_backend_sched_eval_callback",
      "capture_tensor": "l_out-{zero_based_layer}",
      "tensor_transport": "ggml_backend_tensor_get/set, F32 last position only",
      "source_urls": [
        "https://github.com/ggml-org/llama.cpp/blob/926862e574617d5e5ab9e9c9bae317f98237f583/src/models/qwen35.cpp",
        "https://github.com/ggml-org/llama.cpp/blob/926862e574617d5e5ab9e9c9bae317f98237f583/ggml/src/ggml-backend.cpp"
      ],
      "input_embedding": "explicit CUDA buffer override token_embd.weight"
    },
    "numpy": "2.5.3"
  },
  "source_sha256": {
    "gguf_runtime.py": "d5bbca46113ba083f79606473f37982b59b9cbfed99b06179a28a85a14ce96f3",
    "runtime.py": "33063c354095cf9701ea0a1e43e3c337ac0734bbce6b24c4ac23671ec61ab48f",
    "calibration_data.py": "0b3e3d3933880ff1f1c0b65775e5024ffa178409f4559364b68f021b0338ffd7"
  }
}