One prompt. One power plan. Thirteen models.
We handed the same sheet and the same instruction to our fine-tuned takeoff model, four frontier LLMs and eight commercial takeoff platforms. Pilars counted and classified duplex receptacles at 97.5% accuracy. Everything else landed between 65% and 88%.
The prompt every system received
Zero-shot, no examples, no retries, no follow-up clarification. Exactly this text and exactly this image.
It reads like a simple count. It isn't. The plan uses one base symbol with four kinds of modifier that change what the device actually is: a mounting-height note (8", 30", 4' AFF), a GF tag for ground-fault protection, a WP tag for weatherproof exterior devices, and AC for dedicated equipment circuits. A correct answer has to find the symbol, read the two-character text sitting beside it, and not confuse it with the junction boxes, switches, exhaust-fan connections and data outlets drawn in the same weight on the same wall.
What was tested
- SheetE101 · Power Plan and Power Plan – Bid Alternate #1
- ProjectFEMA 361 safe-room addition to an elementary school, Camden County, Missouri
- Scale1/8" = 1'-0", two plan views on one sheet
- Spaces12 classrooms, MP room, lobby, offices, corridors, restrooms, exterior
- Modifiers8" / 30" / 4' mounting heights, GF, WP, AC, NEMA 14-30R, keynotes E.3–E.7
- DistractorsJunction boxes, switches, EF and WH connections, data outlets, panel and transformer symbols
- Ground truthVerified by hand by two estimators, count reconciled to panel schedules
Accuracy on the same sheet, same prompt.
Score is the share of ground-truth receptacles located with the correct type and modifier, after subtracting false positives. One number per system, one attempt each.
| System | Type | Accuracy | Where it broke |
|---|
The annotated result
Every device Pilars found is boxed on the plan, coloured by the class it assigned. This is the output an estimator reviews, not a bare number.




Six failure modes, one sheet.
The same mistakes recurred across frontier LLMs and takeoff platforms alike. The pattern is the point: these are symbol-reading problems, not reasoning problems.
Counted the symbol, ignored the text
Most systems found the receptacle and skipped the 8", 30", GF or WP beside it, returning a flat count where the prompt asked for classification. A plain duplex and a WP GFCI on the exterior wall price very differently.
Junction boxes and switches counted as receptacles
Circle-with-line symbols, EF connections and single-pole switches drawn in the same line weight were folded into the receptacle count, inflating totals in every restroom cluster.
Dropped devices where symbols stack
The work-room counter and the corridor 111 area hold receptacles a few pixels apart. General vision models merged neighbours into one detection or stopped counting partway through the cluster.
WP devices on the building line skipped
Weatherproof receptacles sit outside the wall line, often at the edge of the drawing. Several systems never looked there.
Mishandled the bid alternate
The sheet carries the base plan and Bid Alternate #1 side by side. Some systems counted one view, some counted both without saying so, and none separated them unprompted. Pilars reports them as two scopes.
A number with no location
LLMs returned a total and a paragraph. Without a boxed device on the plan, an estimator has no way to check the count except by redoing it, which removes the reason to use AI at all.
The same test, recorded.
Screen recordings of the identical prompt and sheet on every system. Videos are being added as they are finished.
How the test was run
- One input. Sheet E101 exported once as a single raster image and used unchanged for every system. No cropping, no upscaling, no per-tool preprocessing.
- One prompt. The instruction above, verbatim. For takeoff platforms with a fixed counting UI, the receptacle symbol was selected and the platform's own AI count was run with its default settings.
- One attempt. First response only. No retries, no follow-ups, no prompt engineering per model.
- Ground truth by hand. Two estimators independently counted and classified every receptacle on both plan views, resolved disagreements against the panel schedules and circuit numbers, and froze the answer before any AI run was scored.
- Scoring. A device counts as correct only when it is located and its type and modifier text match ground truth. Missed devices, extra devices and wrong modifiers all reduce the score.
Pilars is used for takeoffs across dozens of trades and sheet types; this page tests one prompt on one sheet and should be read as such. We publish the input so anyone can rerun it.
Frequently asked questions
What exactly was tested?
A single electrical power plan sheet (E101) from a real school addition bid set was given to thirteen systems with the identical prompt: count all duplex receptacles including GFCI and WP variants and read the modifier text on each. Output was scored against a human-verified ground truth.
How is accuracy calculated?
Each receptacle in the ground truth is a scoring unit. A system earns credit only when it locates the device and reports the correct type and modifier text. Missed devices, false positives and wrong modifiers all count against the score.
Why do frontier LLMs score lower than a fine-tuned model?
General vision models were not trained on electrical symbol conventions. They confuse receptacles with junction boxes and switches, drop devices in dense areas and ignore the small modifier text that changes what a device is.
Were the competitors given any advantage?
Every system received the same image at the same resolution and the same prompt, zero-shot, with no retries. Takeoff platforms were used through their normal AI counting workflow.
Can I run the same test on my own plans?
Yes. Upload a power plan to the Pilars demo, ask for the count, and compare the annotated output against your own takeoff.