Focus Max

- One strong GPU — the cheapest VRAM per dollar
- Serves local LLMs and runs LoRA fine-tunes
- Cores and RAM sized to keep the GPU fully utilised
- NVMe staging for datasets and checkpoints
The size of the graphics card’s memory decides which models you can run at all. We build machines with enough of it to serve models, fine-tune on your own data and run simulations — no job queue, no hourly billing, and data that never leaves your building. FOCUS MAX for serving a single model; FOCUS X and HYPERFOCUS when training is the wait.
Four machines on one ladder. FOCUS MAX is one strong card for models that fit its memory. FOCUS X adds a second when fine-tuning is the wait. HYPERFOCUS takes up to six cards when the machine is the cluster. CREATOR COMPACT is fourteen litres for the development seat. All four can be configured before we build them.




Focus and Creator run the same performance at each level — the difference is the chassis and how much of it you want on the desk. The full range shows every model and level side by side.
Training hits a wall at video memory. Simulation ignores it and asks for cores instead. Here is where each framework puts the load, what it looks like when a machine is short of it, and which model we would start you on.
| Program | What it depends on | How it shows up | Start with |
|---|---|---|---|
VRAM before anything else. The card's memory decides which models fit; compute decides how long they take. | An out-of-memory error rather than a slow run — the job does not start at all. | UltrafocusFine-tuning and parallel sweeps | |
VRAM, then the input pipeline. Cores and NVMe keep the card fed. | A GPU sitting idle between batches because data loading cannot keep up. | UltrafocusDistributed training in one box | |
Cores and memory, not VRAM. Only some toolboxes touch the graphics card at all. | A solve that scales with core count, and large matrices that exhaust the workspace. | Focus MaxCores, not video memory | |
Cores, memory and solver scratch. Licensing is often per core, so the spec and the licence have to be decided together. | FEA and CFD models that will not fit in memory, and solves that run overnight either way. | HyperfocusCheck your licensing first | |
| Running models locally | The card's memory is the ceiling. Quantised weights lower it; a longer context window raises it again. | A model that loads and runs on one machine and refuses to load on another with a faster card. | Size VRAM to the modelNot to the benchmark |
Where your work sits matters as much as which framework you open. Running published models locally is bounded by the card's memory alone, and a single strong GPU covers it.
Fine-tuning on your own data is where a second card stops the waiting. Multi-GPU training and sweeps is the level at which owning the machine starts costing less than renting it by the hour.
Five things decide how fast an edit feels. Here they are in the order that matters for post — what each one does, and what you notice when there is not enough of it.
VRAM first, speed second. Memory decides whether the job runs at all; the core count only decides how long it takes once it does.
Loading, decoding, tokenising and augmenting all happen on the CPU. With too few cores, the GPU sits partly idle waiting for the next batch. For solvers, the CPU is the whole job.
System RAM comfortably above total VRAM is the safe rule. It holds the dataset between epochs, and it's where layers land the moment anything is offloaded off the card.
Training reads millions of small files in random order, then writes a checkpoint over and over. That pattern is what NVMe is for — the working dataset stages locally, the rest archives.
A run is hours or days at 100%, not a benchmark burst. Throttling doesn't produce an error — it simply adds hours to every epoch.

Focus Max · Focus X · Hyperfocus · Creator Compact
Everything a run needs lives in the graphics card’s own memory — the model weights, the activations, and if you are training, the gradients and optimiser states on top. If it does not all fit, the job does not start. VRAM is the first number to settle and the last one to compromise on, and it is why FOCUS MAX, FOCUS X and HYPERFOCUS are separated by card memory before anything else.
When a model doesn't fit, there are three workarounds, each with a cost. Quantising to 8-bit or 4-bit gives up a little quality. Offloading layers to system RAM cuts throughput sharply, because data must cross the PCIe bus for every token. Reducing batch size and context length until the job fits slows training and shortens how much the model can see at once.
Every machine on this page is built to meet what TensorFlow, PyTorch and CUDA actually ask for, with components chosen to hold performance through a long run rather than throttle after a short burst. The most useful sizing detail is the largest model you plan to run, at the precision you plan to run it — with that, our corporate solutions team can size the card to the job rather than to a price bracket.
Batch size and context length are VRAM questions. Longer context and bigger batches both cost memory before they cost time.
Training costs far more memory than inference. A model you can serve comfortably can still be out of reach to fine-tune at full precision.
Two cards is more total VRAM, not one big pool. A single model only spans both if your framework shards it — otherwise each card holds its own copy.
Anyone can order the same parts. What you are really choosing is who builds the machine, who tests it before it ships, and who answers the phone when something goes wrong two years from now. That is where we are different.
4.9 on Google from over 4,200 customer reviews
Best Custom Desktop PC Brand 2020–2023
8 years of workstation builds — universities, studios and practices
Every system is checked as it is built, then stress-tested and fully updated before it leaves us — so problems are found here, not at your desk. It arrives ready to work, hand-foamed for safe transit anywhere in Australia. Four years of national awards say we get that right.
Full-system cover for three years, extendable to four — the whole machine, not a chain of separate component warranties you have to chase. If something happens in transit or years later, a dedicated team gets you running again. Support carries on for as long as you own it.
Add on-site servicing and we come to you. No packing it up, no courier, no waiting while it travels both ways — a technician diagnoses and repairs it where it sits. It is the difference between days of downtime and hours of it.
A dedicated workstation team — not a call centre — with eight years of these builds behind it, trusted by universities, studios and practices. Call or email and an Australian expert helps you match the right tool to the job — before you buy, and long after.
What we actually check before a machine is packed, how we specify one in the first place, what happens in year three, and an honest comparison against buying off the shelf or building it yourself.
Liquid particle simulation at the University of Melbourne, a museum installation running on brain-cell computers, and a security system pulling multiple 8K streams at once. Compute that runs unattended and cannot fall over.

A Hyperfocus workstation used for advanced liquid particle simulation — modelling how dirt dams respond to real-world conditions, to help predict and prevent catastrophic structural failures before they occur.

Real-time graphics processing alongside Cortical Labs’ brain-cell computers, for an installation that used biological neural cells to interpret and generate visual data. Our hardware ran the live Unreal Engine graphics.

A Hyperfocus workstation running PICA’s security infrastructure — a four-GPU system handling multiple simultaneous 8K streams, consistently and stably, around the clock.
Some of these relationships now run to several years and multiple deployments. Others began with a single machine.










Curated configurations held in stock and dispatched fast. Professionally built, tested and backed by Aftershock PC — no configuration required.
Start with the machine built for your software, then change what you like. We build, test and optimise it before it leaves us.
Tell us the camera, the codec and how many seats. We'll come back with a spec and a price — written enquiries answered within one business day.