LocallyAI

Home/Guides/Local AI hardware

Choose local AI hardware
from the work backwards.

The best machine is not the one with the largest number on a product page. It is the smallest supported configuration that passes your real workload with useful speed, enough capacity for the team and room for the system around the model.

By Harry Ycas and Matt Hicks, co-founders · 31 August 2026 · 9 minute read

Start with five business questions

Hardware shopping goes wrong when a buyer starts with processor names. Start with the work instead:

Job
Is this short drafting, long-document review, cited search, transcription, image work, software development or a mix?
People
How many staff may use it in the busiest hour, not merely how many have accounts?
Material
How large are the documents and recordings, and how quickly will the approved collection grow?
Boundary
Must core work continue offline, and which web, mail, storage or business-system connections are allowed?
Service
Who will patch, back up, restore, support and eventually replace the machine?

Turn the answers into three examples: an ordinary job, the heaviest realistic job and an awkward edge case. Those examples become the sizing test and later the acceptance test. A benchmark from somebody else’s machine cannot answer whether your scanned report, long meeting or shared workload is usable.

Memory decides what fits

For local language models, memory capacity is usually the first constraint. The machine needs space for the model files being used, the working context, the operating system and the services around the model. Document search, transcription, vision and several active users can add pressure. A model that technically loads while leaving no working room is not a sensible business specification.

Memory bandwidth affects how quickly many local models produce text. Model format, compression and runtime matter too, so two machines with the same capacity can feel different. That is why a quote should name the exact model, format, engine and tested context rather than promise performance from “128 GB” alone.

Leave headroom for the agreed workload and routine system activity. Do not buy speculative capacity for a model nobody has tested, but do not size to the last available gigabyte either. Memory on many compact systems cannot be expanded later.

Storage holds more than documents

Storage must cover the operating system, model library, document index, uploaded files, recordings, exports, logs and safe working space for updates. Backups normally belong on a separate target; a second copy on the same internal drive is not protection from drive failure or loss of the machine.

Ask how much data exists now, how quickly it grows and which source remains authoritative. If the appliance indexes files from a document system, decide whether it stores copies, an index or both, and what deletion from the source should do. Retention is a business and legal decision, not a number the hardware vendor should silently invent.

Three practical platform paths

LocallyAI currently quotes three main compact-hardware paths. They run the same Harness, but operating system and model-engine support differ:

Apple silicon

Our usual recommendation for an office that wants quiet hardware, high memory bandwidth and a straightforward local-model path. Memory and internal storage are chosen at purchase, so sizing and warranty planning matter.

AMD Strix Halo

A compact Windows or Linux path with useful unified-memory capacity. It suits organisations whose operating environment or other software points to a PC, provided the exact model runtime passes testing.

NVIDIA DGX Spark

A Linux and CUDA path for NVIDIA-native development or workloads that require that stack. CUDA compatibility is the buying reason; it is not automatically the fastest choice for every language model.

Product ranges and availability change. The current machines page records LocallyAI’s supported paths and links to official Apple, HP, Minisforum and NVIDIA specifications. The platform comparison covers the brand-level decision; this guide stays focused on sizing the business requirement.

A desktop appliance or a bigger server?

Many small offices do not need a rack server. A compact appliance can be quieter, easier to place and simpler to repair through a manufacturer’s normal channel. Bigger workstations or multiple appliances make sense when the tested job needs discrete GPU memory, more simultaneous work, separation between departments or a supported CUDA workflow.

More hardware also brings more heat, noise, power, network and support work. Before buying a high-powered workstation, confirm:

  • the office has a suitable circuit, outlet and power-protection plan;
  • the location has ventilation and is not accessible to visitors or casual staff;
  • the network path is wired and sized for the staff and files using it;
  • the manufacturer warranty and local repair route suit the business;
  • there is a plan for work while the machine is unavailable.

A UPS can help with short interruptions and safe shutdown, but it is not a backup or a guarantee of service continuity. Have the business’s IT provider or electrician check site-specific power and network requirements.

Capacity is about the busiest hour

Ten registered users do not necessarily require ten times the hardware. What matters is how many heavy jobs overlap. A receptionist transcribing a meeting, two staff searching long files and another person using a larger reasoning model can be more demanding than twenty people making occasional short requests.

Ask the supplier to show what happens under the expected peak: Does work queue? Can a smaller “Fast” model handle everyday jobs while a larger model handles selected work? Do speech and document indexing compete with chat? When would a second appliance be simpler than one larger machine? The answers should shape the quote and the handover instructions.

Local, hosted or a deliberate mix

Lowest upfront costusually hosted
Newest frontier modelsusually hosted
No local hardware administrationhosted
Core work during an internet outagelocal
Office-controlled normal processing pathlocal
Finite known capacitylocal
Selected internal work local, public work hostedmixed

A hosted product puts more infrastructure responsibility with its provider, but the buyer still owns account configuration, suitable use and data decisions. A local system removes the external model service from the normal path, but the buyer now needs a machine, backup, physical security and support plan. The Australian Cyber Security Centre’s cloud shared-responsibility guidance for small and medium businesses is a useful prompt when comparing what a hosted provider carries and what remains with the customer.

Price the whole ownership period

Compare the same useful life and workload on both sides. For local hardware include the installed purchase, GST treatment, electricity, backup target, warranty or hardware cover, support, expected update work, replacement planning and paid connected services. For hosted software include every required seat, plan tier, usage allowance, separate transcription or search product, expected headcount and exit costs.

A local model does not create a per-message model bill, and LocallyAI does not charge a Harness fee for every login. Capacity still has a ceiling. A growing team may need a larger or second appliance, and custom integrations remain separately scoped. Use the worked comparison method rather than accepting a universal break-even claim.

Buying for a regional or interstate office

The same workload test applies nationally. Outside LocallyAI’s listed New South Wales on-site regions, the normal route is a machine configured and tested before shipping, followed by remote setup with a staff member or IT provider at the office. The local person handles power, sign-in and the network point; configuration and acceptance can be completed over a scheduled screen share.

For a remote delivery, put freight, insurance, manufacturer repair, spare-device expectations and support hours in writing. No supplier can make distance disappear, but the job can be designed so a hardware fault follows a known path rather than starting a search for the receipt.

Hardware questions for the quote

Exact configuration
Manufacturer, model, processor, memory, storage, operating system and warranty.
Exact AI stack
Model names and versions, file formats, runtimes, context limits and licences.
Evidence
Results from the ordinary, heavy and awkward acceptance jobs on this configuration.
Peak use
Expected simultaneous jobs, queuing behaviour and the point where another appliance is needed.
Site needs
Power, wired network, physical placement, ventilation and staff-device requirements.
Recovery
Backup target, restore owner, keys, test schedule and the plan during hardware repair.
Change
What can be upgraded, what is fixed at purchase and what gets re-tested after change.
Exit
Data export, configuration handover, secure disposal and replacement route.

Next: read how an Australian installation should run and what support and maintenance should cover.

Questions

Before choosing hardware

How much memory does a business local AI machine need?

There is no honest universal number. It depends on the exact model files, working context, document and speech services, simultaneous users and required headroom. Ask for a representative demonstration before ordering, then require the heavy acceptance job to pass on the delivery configuration before handover.

Can we use hardware we already own?

Possibly. The machine needs an assessment of memory, storage, operating system, model-engine support, warranty status and performance on the intended workload. If it passes with enough headroom, reusing it can be the best buy.

Does a bigger AI computer always give better answers?

No. Hardware affects which model fits, how quickly it responds and how much work it can handle at once. Answer quality also depends on the model, prompt, source material and workflow. A larger machine running an unsuitable model is still a poor system.

Does local AI hardware need the internet?

Core chat, approved document search and transcription can run without it in the standard LocallyAI configuration. Web research, updates, external integrations, remote support and off-site services need a connection when enabled. The quote should name every exception.