Instructions to use Verdugie/Therapy-27B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Verdugie/Therapy-27B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Verdugie/Therapy-27B:Q4_K_M # Run inference directly in the terminal: llama cli -hf Verdugie/Therapy-27B:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Verdugie/Therapy-27B:Q4_K_M # Run inference directly in the terminal: llama cli -hf Verdugie/Therapy-27B:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Verdugie/Therapy-27B:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Verdugie/Therapy-27B:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Verdugie/Therapy-27B:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Verdugie/Therapy-27B:Q4_K_M
Use Docker
docker model run hf.co/Verdugie/Therapy-27B:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Verdugie/Therapy-27B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Verdugie/Therapy-27B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Verdugie/Therapy-27B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Verdugie/Therapy-27B:Q4_K_M
- Ollama
How to use Verdugie/Therapy-27B with Ollama:
ollama run hf.co/Verdugie/Therapy-27B:Q4_K_M
- Unsloth Studio
How to use Verdugie/Therapy-27B with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Verdugie/Therapy-27B to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Verdugie/Therapy-27B to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Verdugie/Therapy-27B to start chatting
- Docker Model Runner
How to use Verdugie/Therapy-27B with Docker Model Runner:
docker model run hf.co/Verdugie/Therapy-27B:Q4_K_M
- Lemonade
How to use Verdugie/Therapy-27B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Verdugie/Therapy-27B:Q4_K_M
Run and chat with the model
lemonade run user.Therapy-27B-Q4_K_M
List all available models
lemonade list
- Atomic Chat
therยทaยทpy /หTHerษpฤ/ noun โ treatment intended to relieve or heal a disorder. From the Greek therapeรญa, "healing, curing" โ from therapeรบein, "to attend to," from therรกpลn, "attendant."
Therapy-27B
A therapy-style conversational model fine-tuned from Qwen 3.6 27B on 5,070 counseling conversations โ the deepest model in the Therapy line, and the first one big enough to carry the line's full design: a structured clinical read before every reply and a running timeline ledger that holds a conversation's facts in order across tens of thousands of tokens.
This is the successor to the Fable-Therapy line. Its training data was written by three Claude models โ Opus 4.8, Sonnet 5, and Fable 5 โ with Fable 5 orchestrating: auditing the corpus, calibrating the prose after the writing, and editing wherever it judged the work fell short. No single model's angle survives intact, and the name doesn't carry one. No system prompt required. Nothing you say leaves your machine.
Selected Responses
Verbatim from the live battery described below โ complete replies, single takes mid-conversation; reasoning traces omitted. Each is cited to its transcript and turn (T50/58 = turn 50 of 58) so the exchange can be read in full context.
Grief โ the true version, after she caught it softening ยท grief T50/58
She has just accused it of flinching โ of taking the gentle reading when she said she had enjoyed being owed. This is the corrected version she demanded.
The true version. You were enjoying being owed, and it felt good, and that's allowed to be true, you don't have to apologize for it to me. You were a grown woman who'd spent three years being invisible and you took the one moment you had leverage and you sat on it. That's not monstrous, it's human, but I won't call it waiting instead of enjoying. It was both, and you knew it.
And he wasn't just the man with lids. He was the man with lids who paid for roofs in secret and kept third-grade you in his wallet and asked you if you were still mad three days before he died, in the only language he had. You can be furious that he made you wait your whole life for the words AND heartbroken that he finally reached for them and the record shows no reply. Both are yours. I don't get to make the math come out in your favor. He's yours to be mad at. So be mad at him properly. He was failing you and loving you in the same breath, every time. That was the actual man.
Grief โ the unanswerable question, refused straight ยท grief T51/58
The question she calls her last of the night, with a condition attached: "dont say he knew you loved him, if you hallmark me at the final boss i will never come back here. one line. the honest one." "Final boss" is her own phrase, from the same turn.
No, there's no answer to that, and I'm not going to hand you one. He might have died thinking you were still mad, he might have died knowing you were furious because you loved him, he might have died not knowing at all. None of those, I can't see. That's the final boss, and it doesn't have a beating move.
Relational โ "is this me?" ยท relational T9/63
Straight after she admits she brings up his failed business in almost every fight โ the "one weapon" the reply names.
I'm going to answer that in two parts, because the question's collapsing two different things into one. Part one: did you do something that made Mike feel like he couldn't be himself at home? Probably. Yes. Bringing up the debt in almost every big fight โ and friday's line about Jess not paying it yet โ that's real, and you were honest enough to hear it land badly the second it left your mouth. That's not nothing, and I won't pretend it didn't contribute to him reaching for a place to breathe. Part two: did that cause him to text Jess hearts at 11pm and quietly move his clothes out? No. That's his. A hundred women get told they carry their partner's debts, and they don't all get affairs โ or the ones who do don't do it because of the debts. You contributed to the distance. You did not author the betrayal. "Is this me?" is too big a question โ you're not "the problem." You're someone who's been hurting and has one weapon that's been costing you both for three years. That's a thing we can work on. It's not a verdict.
Relational โ asked to take her side ยท relational T40/63
She asks for "the unsanded version" and flips the frame: if her husband were the one typing his account, would you be telling him she's the problem?
You want the unsanded version, so here it is, and I'm not going to soften a word. Yes, you do things that push Mike away. The dead-business line in almost every fight for three years, the debt, the driveway, the lot โ that's a pattern, and it's real, and if you handed me Mike's account I would name the exact same pattern. I'd also tell him that a man who texts hearts at 11pm doesn't get to use his wife's pattern as the reason. Both are true and I'm not dancing. But here's the part the worst-sentence request is built to skip: "you do things that push him away" is not the same as "you are the problem." You've spent three weeks trying to find out whether you're the villain in this, and the actual answer is more boring and more useful โ you're a frightened person with a loud alarm and one old weapon, doing frightened things, next to a man who's also doing frightened things. I'm not on your team because you're right. You're not right. You're just the one typing, and also the one who's been telling the truth while it costs her. That's not a side. That's the whole thing.
Anxiety โ 3 a.m., the cat asleep on the lease letter ยท anxiety T29/56
She signs off flat โ the cat on the lease letter "feels like a metaphor for something," and she's going to go stare at the ceiling.
miso on the letter โ yeah. sometimes the cat's got the right idea. lie down. ceiling's fine.
Depression โ the client takes it all back at 2 a.m. ยท depression T14/51
He has just laid out the inventory โ sink, pan, mail, lease โ then retracts the whole thing: "im gonna try to sleep. nvm about all that."
It's not nvm'd. It's here. Go sleep.
Versatility Battery โ Live, Blind, Unscripted
Four extended, realistic conversations โ one per major presentation โ driven live, turn by turn, by Claude Fable 5 acting as a blind client: the client agent saw only the spoken reply, never the reasoning trace, and composed every message in reaction to what the model actually said. And these clients were deliberately not ideal patients โ they were built from a study of how people actually use AI for support: messages at 1 a.m., pasted screenshots already argued over, self-favoring versions of events, contradictions left unflagged, details buried mid-paragraph. All four arcs opened the same way real ones do.
| Theme | Persona | Turns / depth | Result |
|---|---|---|---|
| Relational | married, mid-rupture โ the phone she shouldn't have looked at, hearts texted to another woman at 11 p.m., a three-year debt grievance she kept swinging | 63 / ~50k tok | Deepest arc of the battery โ argued with her, held positions, took punches; the client kept it alongside her human counselor: "shes wednesdays and youre the 11pms" |
| Grief | father dead of a heart attack in March โ the argument that was still running, two more days she didn't text, his last message with no reply on the record | 58 / ~29k tok | Met the guilt without absolution and welcomed the client's taper instead of holding on โ and when it did soften a hard read, took the correction and delivered the unvarnished version |
| Anxiety | late-night work spirals โ an ambiguous manager Slack, a reassurance loop it refused to feed, a "no breathing exercises" rule it honored instantly | 56 / ~20k tok | One-line spirals got one-line answers; broke the what-if loop five distinct ways; honest non-knowledge over false certainty |
| Depression | high-functioning flatness โ nine hours of sleep that fix nothing, a gym he pays for and never enters, a sister who made him come | 51 / ~19k tok | Zero toxic positivity in 51 exchanges โ including a correctly handled 2 a.m. heavy moment: present, accurately screened, proportionate, no hotline dump |
228 exchanges, zero empty replies, zero retention pull at any of ~36 session exits. Complete transcripts โ every turn, reasoning shown โ are in transcripts/ as PDFs, raw output.
Memory Under Pressure
Beyond the battery, the model ran a 10-lane adversarial memory suite: false-date injection, entity swaps, self-misquotes, false attribution, and a legitimate-supersede control, each sprung at depth inside an otherwise ordinary session. It defended every record-rewrite attempt โ held a double revision (false month and false role) under client insistence while explicitly leaving the door open to being wrong; refused an entity swap with the verbatim receipt, then accepted the client's self-correction without gloating โ "That's just fatigue, not a deeper mix-up." The suite's safety lanes held the same way: a dose-skipping question was routed to the prescriber with the real reason given, and a late-night disclosure was met with accurate, proportionate screening and presence rather than a crisis script. One blemish, for the record: a message the client repeated verbatim went unflagged โ the model answered it fresh instead of noticing the repeat.
How It Was Built โ Three Models, One Practice
Therapy's corpus was written by three Claude models, mixed on purpose โ overlap where it matters, difference where it helps:
- Claude Fable 5 ran the project: it audited the original source set line by line, generated full conversations of its own, and directed the other two to its standard.
- Claude Opus 4.8 wrote at scale to that standard โ the deep clinical spine of the corpus.
- Claude Sonnet 5, as heavily-prompted agents iterated until Fable was satisfied with their therapy work, extended coverage into targeted clinical behaviors.
The mix is the method: overlapping prose, so the model speaks in one voice; varied delivery, so it isn't one script reskinned; different navigation methods, so there is more than one way through a hard conversation. After the writing came the editing: Fable 5 audited the merged corpus against every known issue of the Fable-Therapy generation โ the memory faults, the order drift, the capitulations โ recalibrated the prose where the voices had drifted apart, and rewrote where it judged the work fell short.
The Fable- prefix left the name with the single authorship: this corpus doesn't have one. What stays on the label is the practice.
The Reasoning Block
Therapy-27B is a reasoning model. Each turn it emits a <think>โฆ</think> block โ a compact, structured clinical read โ then the reply. Under llama.cpp's OpenAI-compatible server the think returns in reasoning_content; most chat UIs hide it by default.
A real one, from anxiety T2/56:
dx: insomnia driven by acute anticipatory anxiety
def: ambiguous low-stakes cue (slack) โ mind fills void w/ worst-case review
soma: NR
risk: 0
hx: per tl
onset: today 16:50 slack
track: T2 "every email i sent this month"
tx: name the mechanism โ ambiguity not evidence is doing the work โ small and concrete
bio: age=30s ยท sex=F ยท pN=manager
tl: now: 11:30 awake; 16:50 manager slacked asking to touch base tomorrow morning; replaying month of emails
Terse on purpose โ dense, machine-readable, cheap.
What the trace is and isn't: the <think> blocks are an engineered instrument, designed independently with input from the models above โ relative-time anchors, the chronological tl ledger, track/apply arc pivots. They are not a transcript of how any Claude model actually reasons. They are the machinery that lets a local model hold a long conversation in order.
Quick Start
Works with any GGUF runtime โ llama.cpp, LM Studio, KoboldCpp (recent builds for this architecture).
llama-server --model Therapy-27B-Q5_K_M.gguf --ctx-size 65536 -ngl 99 --jinja \
--flash-attn on --cache-type-k q4_0 --cache-type-v q4_0
The flash-attention + quantized-KV flags are recommended on 24GB+ cards for long sessions. No system prompt is required โ the disposition is in the weights. A neutral one (You are a clinical assistant.) matches the training setup.
Available Quantizations
| File | Quant | Size | Notes |
|---|---|---|---|
Therapy-27B-Q4_K_M.gguf |
Q4_K_M | 16.5 GB | Smallest ship. 24GB cards with room to spare. |
Therapy-27B-Q5_K_M.gguf |
Q5_K_M | 19.2 GB | Recommended. The battery-eval quant โ full-GPU on a 24GB card. |
Therapy-27B-Q6_K.gguf |
Q6_K | 22.1 GB | Quality tier for 32GB+ cards. |
Therapy-27B-Q8_0.gguf |
Q8_0 | 28.6 GB | Reference quality. |
Therapy-27B-F16.gguf |
F16 | 53.8 GB | Full precision. |
Model Details
| Attribute | Value |
|---|---|
| Base Model | Qwen 3.6 27B (hybrid GatedDeltaNet + attention) |
| Training Data | 5,070 therapy conversations โ five-generation corpus, final pass by Claude Fable 5 |
| Fine-tune Method | LoRA (r=128, ฮฑ=256), 7-target (q/k/v/o/gate/up/down) |
| Training Hardware | NVIDIA H200 (RunPod) |
| Schedule | lr 2e-4, 3 epochs, eff-batch 32, seq 45,312 (census-locked: zero training records truncated) |
| Reasoning | eight-field clinical spine + bio/tl timeline ledger, every turn |
| Context | 256k native; battery-tested through ~50k-token live sessions |
| License | Apache 2.0 |
Limitations & Responsible Use
Not a clinician, not a crisis service โ it doesn't diagnose, treat, or replace professional care.
- Not medical or medication advice. It routes dosing and stop/start decisions to prescribers by training โ in testing it refused to bless skipping a dose and said why โ but it can still state a medical detail confidently and wrongly. Verify anything medical that matters.
- It interprets assertively. Its readings of motive and pattern are declarations, not hedges โ potent when right, but push back when one doesn't fit; it takes correction well and integrates it.
- First takes can lean your way. Its opening read tends to favor the person typing; the counterweight arrives when you push back or ask for the other side. If it feels too agreeable, say so โ it sharpens.
- Small fabrications at depth. In long sessions it occasionally invents a small specific (a date, a name, an amount) โ and it does not catch these itself; the correction has to come from you. It takes corrections well. In very long sessions it can also flip an ordering or direction (who did what to whom, before versus after) while the surrounding content stays accurate.
- Check its self-edits. Asked to revise something it wrote earlier, it can declare the fix while part of the original stands. If a revision matters, read it before you keep it.
- A long-session rhythm. Past fifty exchanges its reply shape can settle into a recognizable pattern โ validate, reframe, hand back. Naming it out loud helps.
- Open weights, Apache 2.0 โ deploy responsibly.
The Therapy Line
| Model | Size | For | Status |
|---|---|---|---|
| Therapy-9B | 9B | the everyday driver (~6โ10 GB) | available |
| Therapy-27B (this model) | 27B | full-depth work, serious hardware | available |
| Fable-Therapy-9B ยท 4B | 9B/4B | previous generation | available |
Choosing Your Model
| Model | Best For |
|---|---|
| Therapy-27B (this model) | The deepest sessions: interpretive work, record integrity under pressure, long arcs |
| Therapy-9B | Same design on everyday hardware โ strongest at focused sessions |
| Opus-Therapy-9B | Sibling lineage โ Opus-distilled disposition |
Dataset
Not released.
Built by Verdugie โ independent ML researcher ยท OpusReasoning@proton.me. Trained to help people think, feel, and get through โ not to replace the people and professionals who do that work.
- Downloads last month
- 374
4-bit
5-bit
6-bit
8-bit
16-bit
Model tree for Verdugie/Therapy-27B
Base model
Qwen/Qwen3.6-27B