LocallyAI

Home/The machines/Models

The models on
the current roster.

The Harness sorts everyday work into a Fast lane and heavier work into a Smart lane. This is the roster in the appliance manifest now: which model belongs to which lane, which tier carries it, and the memory budget each tier gives it.

Fast · Smart · vision · transcription · context limit tested on delivery hardware

What runs on your machine.

The model roster is pinned to a local snapshot, so what you signed off is what keeps running.

Model loading
Checked on delivery hardware: load, answer and idle release
Fast and Smart routing
Configured in the appliance profile, with the Compact carrying Fast only
Vision input
Included only when the quoted model and workflow pass the acceptance set
Transcription model
Whisper large v3 turbo, local; browser and phone recordings converted on the appliance
Text context
Quoted and tested per model, engine and hardware tier
Standard model download
Pinned to a local snapshot and installed with the appliance

A quote names the model set and platform being supplied. Harness updates need a temporary connection while they are applied.

The planning budget for each tier.

These are the model-memory budgets in the current manifest. They leave room for the operating system, the Harness and the work around the model instead of spending every last gigabyte on weights.

The Micro · 16 GB fitted
12 GB model budget · Fast lane, embedding and transcription
The Compact · 24 GB fitted
18 GB model budget · Fast lane and transcription
The Office · 48 GB fitted
36 GB model budget · Fast and Smart lanes
The Practice · 64 GB fitted
48 GB model budget · higher-precision Smart roster
The Firm · 96 GB fitted
72 GB model budget · the same standard roster at higher precision

Those budgets are the standard planning position, not a promise that every file below performs equally on every Mac, Strix Halo or Spark configuration. The platform check is part of the quote. Quants are chosen per tier: MLX preferred on Apple silicon, Q6 or Q5 GGUF elsewhere, and Unsloth dynamic quants where a smaller file lets a model fit its budget better.

Fast lane: all four tiers.

Qwen 3.5 9B is the current default. Ornith and Gemma are alternate roster choices for jobs where their behaviour is a better fit.

Qwen 3.5 9B

Q6_K · 65,536 configured context · companion vision file · Apache 2.0 recorded in the manifest.

Default Fast model on Compact, Office, Practice and Firm.

Ornith 1.5 9B

Q6_K · 65,536 configured context · companion vision file · MIT recorded in the manifest.

Fast-lane alternate on all four tiers.

Gemma 4 12B

Q6_K · 65,536 configured context · companion vision file · Google's Gemma Terms of Use recorded in the manifest.

Fast-lane alternate on all four tiers.

Nemotron 3 Nano 4B

Q6_K · 131,072 native context · NVIDIA Nemotron Open Model License recorded in the manifest.

Fast-lane alternate on every tier, and small enough to sit comfortably beside the Micro's transcription work.

Smart lane: Office, Practice and Firm.

The Compact has no Smart lane in the current manifest. The larger tiers carry the same main families at different compression levels, chosen to fit their memory budget.

Qwen 3.8 27B

The current Smart default: UD-Q4_K_M on The Office, Q8_0 on The Practice and The Firm. The roster includes its companion vision file.

65,536 configured context · Apache 2.0 recorded in the manifest.

Ornith 1.5 35B MoE

Q4_K_M on The Office, Q6_K on The Practice and The Firm. The roster includes its companion vision file.

65,536 configured context · MIT recorded in the manifest.

GPT-OSS 20B

MXFP4 on The Office, The Practice and The Firm. This roster entry is text-only rather than paired with a vision projector.

65,536 configured context · Apache 2.0 recorded in the manifest.

Gemma 4 26B-A4B

UD-Q6_K on The Office, UD-Q8_K_XL on The Practice and The Firm. The roster includes its companion vision file.

65,536 configured context · Google's Gemma Terms of Use recorded in the manifest.

Mistral Small 3.2 24B

UD-Q4_K_XL on The Office, Q6_K on The Practice, UD-Q8_K_XL on The Firm. The roster includes its companion vision file.

131,072 native context · Apache 2.0 recorded in the manifest.

Nemotron 3.5 Lightning 30B MoE

UD-Q4_K_XL on The Office, UD-Q5_K_XL on The Practice, UD-Q6_K_XL on The Firm. A reasoning MoE activating ~3B parameters per token.

131,072 native context · OpenMDW-1.1 recorded in the manifest.

DeepSeek R1 Distill Qwen 14B

Q6_K on The Office, The Practice and The Firm. The DeepSeek that actually fits an office machine; text-only, with reasoning output kept separate from the answer.

131,072 native context · MIT distill over an Apache 2.0 base, recorded in the manifest.

Images and speech are separate parts of the roster.

A model name alone does not prove the complete route from browser input to useful output. We track the companion files and the input path as well.

Vision-capable text models

Qwen 3.5 9B, Ornith 1.5 9B, Gemma 4 12B, Qwen 3.8 27B, Ornith 1.5 35B and Gemma 4 26B each have a companion vision file in the current manifest. The Harness now refuses an image on a text-only route instead of pretending it was handled.

Vision models read image attachments alongside text. Your quote names the exact model and projector files the platform carries.

Whisper large v3 turbo

F16 · transcription lane · every tier · MIT recorded in the manifest. Its dedicated speech server has passed a real WAV transcription.

Browser and phone recordings are converted on the appliance before transcription, so meetings and dictation work from any device on the office network.

Qwen3 Embedding 0.6B

Q8_0 · embedding lane · every tier, the Micro included · Apache 2.0 recorded in the manifest. Under a gigabyte, resident beside the lanes, and the model document search uses to rank your material against the question.

Embedding runs on the appliance like everything else: indexing a client file never sends its text anywhere.

The same roster does not mean the same result on every platform.

Mac
Our usual recommendation for local-model speed, quiet hardware and setup simplicity
Windows or Linux on Strix Halo
Useful unified-memory capacity; engine and driver support checked per workload
DGX Spark
NVIDIA-native Linux and CUDA; chosen when that software stack matters
RTX workstation
Discrete VRAM, CUDA and workstation power; custom-scoped around the exact job

See paired appliances and RTX workstations →

The licence travels with the model.

Models we assess and do not approve stay named rather than quietly absent: DeepSeek V4 Flash, for example, is a 284-billion-parameter release whose smallest usable quant outweighs every tier's budget, so it is a custom-build conversation rather than a roster entry — while the DeepSeek R1 distill that does fit is on the Smart roster above. The manifest records Apache 2.0, MIT or Google's Gemma Terms of Use for the current roster. We re-check the publisher's terms before a model is quoted because mirrors and summaries can be wrong.

What the check actually covers

That check is practical, not legal theatre: which model, which version, which files, where they came from and which terms apply. If you ask for another model, it does not become supported until its licence, provenance, memory and runtime have been checked.

Model terms can change. The quote and configuration record name the version assessed for your build.

Questions about the roster

Is 64k context available now?

The manifest target is 65,536 tokens for text models, but usable context depends on the exact model, engine and hardware. Your quote names the delivered limit, and that limit is tested on the final machine before handover.

Does every model run the same on Mac, Windows and Linux?

No. Memory decides whether a model can fit, but the operating system, model engine, file format and driver stack affect speed and compatibility. Mac is our usual recommendation because its local-model path is straightforward. Strix Halo and DGX Spark are main products too, but the quote names the exact models and engine checked for that platform.

Can we install a model that is not on this page?

Possibly. We check its licence, provenance, memory requirement, supported engine and safety implications first. A suitable model can be approved and installed as separately quoted work.

Unknown, guardrail-reduced or oversized models are not silently added to the supported roster. The agreed baseline and any customer changes are recorded so rollback and support stay clear.

What is the transcription status?

Whisper large v3 turbo is the standard transcription model. In the local configuration, browser recordings are converted and transcribed on the appliance. Any connected exception is named in the quote.

Recording access and retention behaviour are tested on the customer devices included in the acceptance set before handover.

Test the work, not the model name.

Bring a contract, a report, a scan or a research question. We can compare the Fast and Smart lanes on the kind of job you actually need, then size the machine around the result.

Bring us one job your team repeats every week. We will show you the system doing it →