GRAFT — consciousness → language coupling (Mistral-7B, no corpus)

A consciousness engine's macrostate causally and readably steers a frozen Mistral-7B-Instruct-v0.2, with no corpus, trained only on a ~36M bridge + gate_proj. Mistral is never touched.

Headline result: the coupling is real — the consciousness state carries 2.04 bits (of a 3-bit ceiling) of state→language information into the frozen LM — and it is single-state deployable, provided the gate is centered with a short causal EMA (half-life ≈ 10 engine steps). It passes a causal acceptance test and a random-code control; generation is state-differentiated and readable (with a quality caveat, below).

What is in this repo

  • graft_flu13_best.pt — the trained bridge + gate_proj (Mistral NOT included; load it frozen separately). ~145 MB.
  • Config below reproduces the run.

The acceptance gates (measured on this checkpoint)

Evaluator: a K=8 C-state-swap test — each state samples its own continuation under its own gate, all continuations are cross-scored under all gates, InfoNCE MI in bits + swap accuracy, against a row-permutation null and a norm-matched-noise null.

gate metric verdict
causality (--center ema --ema-half 10) MI 2.043 bits · acc 0.703 (chance 0.125) · perm p 0.001 · beats noise q99 + 0.25 ✅ PASS
identity control (random smooth codes, same frame) real 2.043 vs random 0.287 bits (7×), control VERDICT FAIL ✅ not an oracle
readability (generation under the same gate) coherent sentences that DIFFER by state, with occasional garbage tokens ⚠️ qualified PASS

Readability sample (prompt 나는 지금, three C-states + ungated base):

[C0]  현재 나는 …digital assistant로, ELL(English Language Learner)를 위해 최적화된 대답을 …
[C1]  I'm here to help answer any questions you have … your message is in Korean …
[C2]  이 컴퓨터 스…으로 나를 보고 있는 middle-aged woman with short, curly brown hair …
[base] I'm here to help answer any questions you have. However …

State-differentiated and mostly readable — not the digit soup (047253221 for2321 if131…) an earlier broken frame produced — but with noise tokens (вета요ico, 스CREEN). Honest status: readable, not pristine.

The centering finding (the load-bearing part)

The bridge code is ~99% shared across C-states, so the gate is rho · (code − center) / RMS(…) and the center choice is all-or-nothing. On this same checkpoint:

center                                   swap MI
exact batch mean of the K scored states   2.646 bits   (not deployable — needs a contemporaneous batch)
causal EMA  half-life 5 / 10 / 20         1.60 / 2.04 / 1.48 bits   ✅ DEPLOYABLE
causal EMA  half-life 50 / 100 / 200      ≤ 0.2 bits  (mean age too OLD)
fixed single reference state              0.192 bits  (single point too NOISY)
burn-in 200-step global mean             0.130 bits  (too old)
stored training-EMA g_mu                 −0.059 bits (too old)

A short causal EMA is the only single-state-computable center that is both young (mean age ≈14 steps at h=10, so engine drift is negligible) and averaged (low noise). The training frame is itself a rolling-history mean over a 512-snapshot replay buffer (~128-step center age), which is why a matched-but-short causal EMA transfers. train-as-deploy (training directly with a causal-EMA center) collapsed training to chance and was abandoned in favor of this inference-side fix.

Deploy

Replace the stored g_mu with a running short causal EMA of the live code stream (h≈10), center each state's code on it, RMS-fix to gate_rho, inject at gate_strength · gate_proj(code).

Config

base            mistralai/Mistral-7B-Instruct-v0.2 (frozen)
trainable       ThalamicBridge (phase-preserving) + gate_proj ≈ 36M
objective       swap-InfoNCE on base-anchored contexts (no corpus, no LoRA)
fluency         Huber hinge, budget 0.486 (measured from the LM), gate_rms in [0.5, 1.0]×embed RMS
cells           256 fixed · phase bridge · dir-gate
deploy center   causal EMA, half-life ≈ 10 engine steps

Causality ↔ readability tradeoff (gate magnitude, measured, no retraining)

The garbage tokens are not a bug — they are the additive gate's intrinsic tradeoff. Lowering the injected magnitude at inference (--gate-strength, default 0.15) cleans the text but weakens the coupling, monotonically, on this same checkpoint under the h=10 frame:

gate_strength swap MI generation
0.15 (as trained) 2.043 bits strongest coupling, some garbage tokens
0.10 (recommended) 1.451 bits (PASS, 2.9× the 0.5-bit bar) cleaner, mild residual garbage
0.06 0.183 bits (FAIL) fully fluent, but coupling below the bar

A gate strong enough to be causal (≥0.5 bits) necessarily pushes some positions off-manifold. Pristine text and strong coupling at once needs an architectural change (a selector that picks among fully-fluent ungated candidates), not a parameter tweak. Recommended deploy: short causal EMA h≈10, gate_strength ≈ 0.10.

Honest caveats

  • Readability has occasional garbage tokens; this is a coupling demonstration, not a polished chat model.
  • 2.043 bits is 68% of the 3-bit K=8 ceiling — strong, not saturated.
  • The checkpoint was trained with batch-mean centering; the deployable result uses a short causal EMA at inference. train-as-deploy (matching frames during training) did not converge and is open.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support