LocallyAI

Home/Studio C

Making capable models
fit smaller machines.

Studio C is the name we use for early model-efficiency research. It is not a separate lab, a shipping compression product or a record of completed client work.

Research direction · low-bit quantisation · careful measurement before claims

The question we are testing.

How much memory can a model give back before the useful parts of its behaviour give way?

Why memory matters

Model weights have to fit in available memory with room left for context and the software doing the work. A smaller representation can put a more capable model on less expensive hardware, but only if the answers remain good enough for the intended job.

Why a file size tells you very little

A compressed model can load and still lose accuracy, instruction following, stability or speed. A useful result has to survive repeatable tasks and comparison with the uncompressed model and established formats.

Where the work stands.

The research direction includes 2-bit, 1.58-bit ternary and 1-bit representations. We have not published a reproducible benchmark, released a supported runtime format or added these formats to the machines we quote.

That status matters. A percentage without the exact model, dataset, task set, scoring method, runtime and comparison point sounds impressive but tells a buyer very little. We are not using internal experiments as a sales claim.

Machines supplied today use the standard, licensed model formats named in the written configuration. The models page explains the current choices.

What a result has to show.

Identity
Exact base model, source, version, licence, file hash and conversion settings.
Memory
Weight size and real peak memory, including context and runtime overhead.
Quality
Repeatable task results against full precision and a standard supported quantisation.
Speed
Load time, first response and sustained generation on named hardware.
Failures
Tasks that regress, unstable outputs and cases where the trade is not sensible.
Reproduction
Enough method and artefacts for another person to check the result.

What may come later

If a format proves useful, runs in software we can support and satisfies its licence, it may become a documented option. Until then it stays research. There is no promised release date.

Fine-tuning is separate

A future customer-specific fine-tuning job would be separately scoped and tested. The quote would name the data, location, licences, access, retention, deletion and deliverables before any material is accepted.

Research, stated plainly

Research stays useful when the status stays plain.

The system we quote is based on hardware, models and workflows available for that job now. Studio C is where we test what could improve the next one.

Bring us one job your team repeats every week. We will show you the system doing it →