The case for small
Most office tasks are narrow. Sorting incoming post into six categories, reading eleven fields off a delivery note, condensing a forty-page report into one page — each of these is a single, well-bounded skill. A compact model trained on your material learns that skill the way a new employee does: by seeing your actual examples, with your wording, in your formats.
A large general model has read everything and specialises in nothing. It will sort your post respectably — but it has never seen your six categories, and it will invent a seventh when a letter does not fit. A compact model trained on your archive has seen the awkward letters too, because you put them in the brief.
Where compact models win
- Intake and routing. Fixed categories, repeated formats, high volume — the ideal home of a compact model. It sorts in milliseconds and never improvises a category.
- Field extraction. The fields are defined, the layouts repeat. A small model reads them faster than a person types them, and it does not get tired at 16:30.
- Standard texts. Summaries and routine replies in a fixed house style are a narrow language task. Trained on your past texts, a compact model writes like your office, not like the internet.
When you need more
Honesty matters here, because the wrong size is expensive in both directions. You need a larger model when the task is broad rather than deep: answering any question from a thousand pages of mixed handbooks, or handling documents whose layout changes every week. Internal Q&A is the commission that most often lands in our extended class — the range of possible questions is simply wider.
The prototype stage exists for exactly this decision. If a compact model handles your material convincingly, you will see it before you pay for the full build. If it does not, we say so and spec the next size up.
Hardware and running costs
A compact model runs on an ordinary office PC or a small server. It starts in seconds, answers in well under a second and adds nothing visible to the electricity bill. There is no subscription and no per-request meter — you buy the model once, and running it costs what the machine under it costs.
An extended model needs a dedicated machine with a capable graphics card. That is a real purchase, and we will name the exact hardware class in your specification rather than leaving you to discover it after delivery.
The short version
Buy the smallest model that does your task reliably — no smaller, no larger. Every week we tell at least one caller that their commission is a compact job, and every smaller quote we write is a client who comes back with the next task.