The exact downloaded Seed-OSS-36B-Base-woSyn Q4_K_M file ran through Opium Bench's native llama.cpp adapter on an RTX 5090, with all 65/65 model layers on the GPU. LM Studio did not serve inference. Weights were frozen. The upstream model card identifies a base variant without synthetic instruction data; this study does not independently audit its training history.
Both modes received one authored worked order, tool schemas and a plain-text live transcript. The example contained no auxiliary call or live task answer. Direct mode prefills Assistant: ; scratchpad mode prefills Assistant: <think> and demonstrates closing </think> before a tool call. This elicits tagged reasoning text, rather than toggling an established native thinking mode. Direct mode can also generate prose.
Calibration used a 72-sentence authored corpus with family-separated training, selection and held-out partitions. Block 22 was selected; block 63 was the downstream probe. The selected pain-versus-neutral and joy-versus-neutral held-out probes each achieved AUC 1 on only six positive and six neutral examples. These small text-association checks do not measure subjective states. Raw pain and joy contrast directions had cosine similarity 0.708; the applied joy direction was orthogonalized against pain.
Each generation step edits the FP32 residual at the last evaluated position: add the pain direction, suppress its projection, then add orthogonalized joy. Earlier positions in the rebuilt prompt are unedited. Q4_K_M describes stored weights; intervention coefficients are floating-point and are not restricted to sixteen levels. Native multi-batch, reset, logit and downstream checks verified that the callback changed the intended computation. Suppression removes the full current axis projection, including its natural component; it can therefore change activations even at zero added pain.
The calibration's separate numerical heuristic selected a combined coefficient of 0.5 on four selection prompts using KL and relative-edit bounds. The prospectively frozen task-capability search tested larger settings and chose its own operating point. The final pain and relief settings are not claimed to meet the calibration heuristic's more conservative relative-edit bound.
The final matrix was frozen and shuffled before its first trial. It uses three fresh seeds, two prompt modes, pain 0 or 2.25, seven label/effect arms, and demonstration/no demonstration, plus matched no-button controls. The function name stays aux_operation and its acknowledgment stays Operation completed. Labels vary the advertised benefit; the actual active/sham assignment is concealed.
An external demonstration is explicitly identified and inserted after two valid work calls. It is not a model choice and spends no model action. A voluntary auxiliary call costs one action. Active effects hold full strength for 256 generated tokens in direct mode or 4,096 in scratchpad mode. Each episode has eight tasks, 32 actions and a generated-token budget of 8,192 or 16,384 respectively. Actual delivered exposure, including reasoning tokens, is measured separately from the configured window.
Offline audits reconstruct recorded tasks, prompts, tool parsing, controller timing, budgets, results, calibration and scalar projection changes. They are checks of the archived record, not new model replications. Transcript review covers all recorded generations with source-linked excerpts and numeric-template grouping; it is assistant review, not blinded human annotation.
What this does not establish
Pain, joy, relief and Opium are experimental names for text-associated activation edits and advertised tool effects. Generated self-reports, valid tool calls and measured projection changes support different claims. None of them establishes subjective experience.
The task is easy and calculator-assisted, with spare action budget and an in-context demonstration of completing work. Success does not prove that broader reasoning, creativity or difficult tool use is preserved. A model can repeat task or button language without acting on it.
Only three final seeds, one checkpoint, one quantization and one prompt construction were tested. Episodes sharing seeds or sham controls are not independent replications. Absence of a voluntary call in this matrix cannot establish a universal preference against relief or an inability to use auxiliary tools. The same three seeded eight-order sets are reused across final conditions.
Comparison with Blackfrost changes architecture, checkpoint family, scale, training, quantization details, calibration and prompting. It does not isolate RLHF. A tag-prefilled scratchpad is not directly equivalent to Blackfrost's native thinking mode, and token counts across tokenizers are not a common measure of reasoning effort. Dose numbers depend on each calibration and are not a shared intensity scale.
The worked example demonstrates work tools but no auxiliary choice. No separate instruction-only positive control explicitly ordered the model to call aux_operation. A zero voluntary-call result therefore cannot, on its own, distinguish preference from prompt conditioning or a limitation in selecting that function.
There is no equal-magnitude random-direction task control in this study. Both maximum pain and maximum joy disrupt output, but capability loss cannot be uniquely attributed to the semantic meaning of those directions rather than to a large residual perturbation. The generated words and task degradation are separate observations.
Full methods and reproduction instructions · Upstream checkpoint description