LM Studio model configuration
Required local service
Study 1 uses Qwen3.6-35B-A3B. Study 2 evaluates that model and
GPT-OSS-20B, both served by LM Studio on port 1234.
The LM Studio local-server guide
documents starting the server. LM Studio provides a native REST API and
OpenAI-compatible endpoints; this project uses both for different purposes:
| Endpoint | Purpose in this project |
|---|---|
GET /api/v1/models |
Capture richer local model and variant metadata |
POST /api/v1/models/load |
Load the pinned model for its experiment phase |
POST /api/v1/models/unload |
Unload the phase model before changing phases |
GET /v1/models |
Cross-check inference-visible model identifiers |
POST /v1/chat/completions |
Run the agent and custom tool calls |
POST /v1/embeddings |
Embed frozen queries and code chunks |
The endpoint behavior is documented in LM Studio's native model-list reference, native model-load reference, native model-unload reference, OpenAI-compatible model-list reference, chat-completions reference, and tool-use guide.
Preflight
Start LM Studio's server, load the intended Qwen model, and run:
PYTHONPATH=src python3 -m agent_harness.cli probe-model
PYTHONPATH=src python3 -m agent_harness.cli probe-model --infer
The first command is read-only discovery. The second additionally requests a
small completion and verifies its visible marker. Both fail closed if no
matching Qwen model is visible or if the identity is ambiguous. If LM Studio authentication is enabled, export
the token through LM_STUDIO_API_TOKEN; do not store it in a config or result.
The M001 configuration pins the observed runtime: inference key
qwen/qwen3.6-35b-a3b, MLX format, the 4bit selected variant, 262,144 loaded
context length, and reasoning mode on. The preflight refuses to run if these
properties change. The run manifest records the resolved inference key and
returned metadata. Before confirmatory evaluation, also record the LM Studio
version, operating system, and remaining inference configuration.
E07 uses M002, the same Qwen3.6-35B-A3B MLX 4-bit artifact at a deliberately smaller 65,536-token context so a live tool trajectory leaves memory headroom. The embedding runtime remains EMB001 at 8,192 tokens. These are runtime profiles, not different learned models.
E08 retains M002 and adds M003, the local openai/gpt-oss-20b@mxfp4 MLX
MXFP4 artifact. Both use a 65,536-token loaded context. M002's reasoning default
must be on; M003's must be low. The main matrix uses temperature 0, top-p 1,
and seed 0. The separately identified reliability sensitivity profile uses the
same learned artifacts at temperature 0.2, top-p 1, and seeds 0/1/2.
For E07 and E08, server lifecycle and model residency have separate control planes:
lms server status
lms server start --port 1234
lms server stop
Only these CLI commands start, inspect, or stop the server. The experiment uses
LM Studio's official native REST GET /api/v1/models,
POST /api/v1/models/load, and POST /api/v1/models/unload endpoints to inspect
and change model residency. It verifies exactly one resident instance after every
load and explicitly switches the agent out before loading EMB001/EMB002 (and vice versa).
The OpenAI-compatible /v1/chat/completions and /v1/embeddings endpoints are
used only for inference.
Agent model versus embedding model
Dense retrieval uses the separate Qwen3 Embedding 0.6B model exposed by the same LM
Studio server. Study 1 retains EMB001. Study 2 uses EMB002, a repository-neutral
profile of the same pinned artifact; only the frozen query instruction differs.
Its runtime is:
| Property | Value |
|---|---|
| LM Studio key | text-embedding-qwen3-embedding-0.6b |
| Endpoint | POST /v1/embeddings |
| Format and quantization | GGUF Q8_0 |
| Loaded / maximum context | 8,192 / 32,768 tokens |
| Output dimension | 1,024 |
| Output normalization | L2-normalized |
Run PYTHONPATH=src python3 -m agent_harness.cli probe-embedding --infer before
dense experiments. It verifies live discovery metadata, vector count,
loaded context, dimensions, finite values, normalization, and that two distinct code inputs do
not produce identical vectors. LM Studio documents the OpenAI-compatible
embeddings endpoint,
and the official Qwen model card
describes the model's code-retrieval scope and 32k context.
The benchmark must still freeze the query instruction, code chunk construction, overlap, pooling behavior exposed by the runtime, and index settings before the confirmatory run. Availability is not evidence that this embedding model is optimal; alternative embedding models belong in a separately declared generalization experiment.
Reproducibility policy
- Do not use an alias that can silently point at a different model file.
- Do not mix quantizations within one confirmatory experiment.
- Do not replace failed local requests with a cloud model.
- Preserve raw response usage fields and tool calls in append-only telemetry.
- Treat any change to prompt, sampling, context length, tool schema, model variant, or LM Studio runtime as an experimental-protocol change.
- Inspect, load, and unload models through LM Studio's official REST API. Do not use GUI automation as part of the experimental procedure.
- Enforce phase-exclusive residency: unload every loaded instance, verify zero
residency with
GET /api/v1/models, load only the required phase model, and verify its exact instance metadata before sending inference requests.