There’s a one-page artifact we ask for before any sizing conversation, and I’d rather show it than describe it. A load profile, filled in for an imaginary but typical institution:

  • Peak concurrent tasks: 6, observed between 9:10 and 10:30 on filing days.
  • Representative input: 40 to 300 pages; output: 2-page cited draft.
  • Response obligation, interactive class: under 30 seconds to first useful text.
  • Response obligation, batch class: complete by 6:00 a.m. review.
  • Batch window: 11:00 p.m. to 5:00 a.m., roughly 400 documents.
  • Degraded mode: interactive work continues at reduced speed; batch pauses.
  • Expansion trigger: interactive queue exceeds 90 seconds for five consecutive business days.

Every line of that profile is doing work that “we have fifty users” can’t do.

The first line kills the headcount myth. Fifty people with occasional short requests can generate less simultaneous demand than four people processing long records at the same moment, so we count overlapping tasks, never accounts.

The input and output lines change the arithmetic more than most buyers expect. A benchmark built on short prompts can’t speak for a workflow that ships three hundred pages in and wants citations back. Context length, output length, retrieval, and tool calls all move the requirement.

The two response-obligation lines exist because “fast” isn’t testable. A lookup while someone waits on the phone needs seconds. An overnight comparison needs to beat the morning meeting. Both can share one system if the queue’s designed on purpose, priorities included.

Degraded mode gets a line because reserve capacity with a stated job is resilience, while spare capacity without one is a question mark wearing a price tag. And the expansion trigger turns future growth from a sales argument into a measurement, which is where it’s always belonged.

The facility signs off last; power and cooling can cap sustained load before the compute budget does, which is why this profile travels with the survey.

After launch, the profile’s the scorecard. Queue’s growing? Find out whether that’s adoption, changed workload, or bad scheduling before it becomes a purchase order. Hardware’s occasionally the answer. I’ve rarely seen it be the first one.