agent-harness / docs /LM_STUDIO.md
cuber12's picture
Publish agent harness research code and paper artifacts
d61821a verified
|
Raw
History Blame Contribute Delete
6.06 kB

LM Studio model configuration

Required local service

Study 1 uses Qwen3.6-35B-A3B. Study 2 evaluates that model and GPT-OSS-20B, both served by LM Studio on port 1234. The LM Studio local-server guide documents starting the server. LM Studio provides a native REST API and OpenAI-compatible endpoints; this project uses both for different purposes:

Endpoint Purpose in this project
GET /api/v1/models Capture richer local model and variant metadata
POST /api/v1/models/load Load the pinned model for its experiment phase
POST /api/v1/models/unload Unload the phase model before changing phases
GET /v1/models Cross-check inference-visible model identifiers
POST /v1/chat/completions Run the agent and custom tool calls
POST /v1/embeddings Embed frozen queries and code chunks

The endpoint behavior is documented in LM Studio's native model-list reference, native model-load reference, native model-unload reference, OpenAI-compatible model-list reference, chat-completions reference, and tool-use guide.

Preflight

Start LM Studio's server, load the intended Qwen model, and run:

PYTHONPATH=src python3 -m agent_harness.cli probe-model
PYTHONPATH=src python3 -m agent_harness.cli probe-model --infer

The first command is read-only discovery. The second additionally requests a small completion and verifies its visible marker. Both fail closed if no matching Qwen model is visible or if the identity is ambiguous. If LM Studio authentication is enabled, export the token through LM_STUDIO_API_TOKEN; do not store it in a config or result.

The M001 configuration pins the observed runtime: inference key qwen/qwen3.6-35b-a3b, MLX format, the 4bit selected variant, 262,144 loaded context length, and reasoning mode on. The preflight refuses to run if these properties change. The run manifest records the resolved inference key and returned metadata. Before confirmatory evaluation, also record the LM Studio version, operating system, and remaining inference configuration.

E07 uses M002, the same Qwen3.6-35B-A3B MLX 4-bit artifact at a deliberately smaller 65,536-token context so a live tool trajectory leaves memory headroom. The embedding runtime remains EMB001 at 8,192 tokens. These are runtime profiles, not different learned models.

E08 retains M002 and adds M003, the local openai/gpt-oss-20b@mxfp4 MLX MXFP4 artifact. Both use a 65,536-token loaded context. M002's reasoning default must be on; M003's must be low. The main matrix uses temperature 0, top-p 1, and seed 0. The separately identified reliability sensitivity profile uses the same learned artifacts at temperature 0.2, top-p 1, and seeds 0/1/2.

For E07 and E08, server lifecycle and model residency have separate control planes:

lms server status
lms server start --port 1234
lms server stop

Only these CLI commands start, inspect, or stop the server. The experiment uses LM Studio's official native REST GET /api/v1/models, POST /api/v1/models/load, and POST /api/v1/models/unload endpoints to inspect and change model residency. It verifies exactly one resident instance after every load and explicitly switches the agent out before loading EMB001/EMB002 (and vice versa). The OpenAI-compatible /v1/chat/completions and /v1/embeddings endpoints are used only for inference.

Agent model versus embedding model

Dense retrieval uses the separate Qwen3 Embedding 0.6B model exposed by the same LM Studio server. Study 1 retains EMB001. Study 2 uses EMB002, a repository-neutral profile of the same pinned artifact; only the frozen query instruction differs. Its runtime is:

Property Value
LM Studio key text-embedding-qwen3-embedding-0.6b
Endpoint POST /v1/embeddings
Format and quantization GGUF Q8_0
Loaded / maximum context 8,192 / 32,768 tokens
Output dimension 1,024
Output normalization L2-normalized

Run PYTHONPATH=src python3 -m agent_harness.cli probe-embedding --infer before dense experiments. It verifies live discovery metadata, vector count, loaded context, dimensions, finite values, normalization, and that two distinct code inputs do not produce identical vectors. LM Studio documents the OpenAI-compatible embeddings endpoint, and the official Qwen model card describes the model's code-retrieval scope and 32k context.

The benchmark must still freeze the query instruction, code chunk construction, overlap, pooling behavior exposed by the runtime, and index settings before the confirmatory run. Availability is not evidence that this embedding model is optimal; alternative embedding models belong in a separately declared generalization experiment.

Reproducibility policy

  • Do not use an alias that can silently point at a different model file.
  • Do not mix quantizations within one confirmatory experiment.
  • Do not replace failed local requests with a cloud model.
  • Preserve raw response usage fields and tool calls in append-only telemetry.
  • Treat any change to prompt, sampling, context length, tool schema, model variant, or LM Studio runtime as an experimental-protocol change.
  • Inspect, load, and unload models through LM Studio's official REST API. Do not use GUI automation as part of the experimental procedure.
  • Enforce phase-exclusive residency: unload every loaded instance, verify zero residency with GET /api/v1/models, load only the required phase model, and verify its exact instance metadata before sending inference requests.