← Blog

How to test AI quantity takeoff: a checklist for your demo

How to really evaluate AI-assisted quantity takeoff: a blind test, 14 checkpoints for the demo, five questions that separate a nice interface from a reliable system, and eight questions for every vendor.

27 August 2026 · 9 min read

Quantity takeoff with calculation, source and confidence class for each line item

AI-assisted quantity takeoff promises a lot: upload a permit plan, get back a bill of quantities with quantities. Whether that holds up for your projects you will not learn from a product page, nor from a demo with a prepared example. You find out when you confront the system with a plan of your own and ask hard questions. This article is the test protocol for that: vendor-neutral and applicable to any software, including ours.

Ask the question that matters economically

Most evaluations fail because they start from the wrong question: “Can the AI produce a tender 100 % on its own?” The answer is no for every vendor, and anyone who claims otherwise is selling. The question that counts economically is:

“How much of an estimator’s working time does the system save, and how much review time is left?”

An estimator needs one to three days for the takeoff and bill of quantities of a single-family house. If that turns into under an hour of processing plus one to two hours of professional review, there is a business case. If three days turn into three hours of processing plus two days of rework, there is not. The test therefore has to measure how checkable the result is, not just whether the total looks plausible.

Stage 1: run a blind test on your own plan

Do not let yourself be shown an interface. Give the vendor a real plan that you have already estimated and say something like: “We want to see what the system gets out of this plan by itself. Please start without our help.” Keep your own bill of quantities to yourself until the evaluation.

Do not pick a neat, simple single-family house. Pick one with rough edges: demolition, sloping ground, a basement below groundwater, different floor build-ups, a roof with dormers, several plan revisions. That is exactly where a takeoff system separates from an AI demo.

Test 1: simple quantities

Measure these by hand beforehand and compare line by line afterwards:

  • m² of floor per build-up
  • m² of screed
  • m² of interior plaster or skim coat
  • m² of exterior render / ETICS
  • m³ of concrete per element
  • m² of formwork
  • running metres of wall per wall type

Document the deviation per line item. Not the total: that can be right by chance while individual quantities are completely wrong.

Test 2: the difficult elements

  • Roof (pitch, dormers, roof windows)
  • Demolition (without a demolition plan: what is assumed, and is that stated?)
  • Ground and excavation (depth from the section, working space, shoring)
  • Foundations and base slab in several thicknesses
  • different wall build-ups as separate line items
  • slabs of different thicknesses, sloped slabs
  • openings and their deduction
  • level differences between storeys
  • several plan revisions

Test 3: the one that decides

Once the bill of quantities has been generated, go through it item by item and do not ask “Is this right?” but: “Please show me how this quantity was derived.”

If it says “plaster interior walls – 428.5 m²”, the system should be able to show: ground floor X m², upper floor Y m², minus doors Z m², minus windows … = 428.5 m². With a reference to the plan sheet the dimensions come from. If it cannot, you will have to check every figure by hand when it matters, and the time saving was an illusion.

Stage 2: 14 checkpoints for the demo

Print this list and tick it off during the demo. Each point can be shown live with a plan of your own, or it cannot.

#CheckpointWhat to look forOK
1Recognising roomsAre all rooms captured with name and area? Do the room areas match the room stamps?☐
2Correct areasGross floor area per storey, floor areas per build-up, wall areas: can they be recalculated from the dimension chains?☐
3Wall lengthsInternal and external walls in running metres, separated by build-up? Or estimated per storey as a lump, and if so, is that flagged?☐
4Recognising windows/doorsCounts and sizes from floor plan, elevation or window schedule? Where from, if no window schedule is supplied?☐
5Deducting openingsAre openings deducted from plaster, facade and tile areas? Is the deduction visible in the derivation?☐
6Heights from sectionsStorey heights, excavation depth, groundwater level, parapet taken from the sections?☐
7Roof areasPitch taken into account? Dormers, roof windows and superstructures counted?☐
8Different wall typesRC 30 cm, RC 18 cm, drywall 10 cm, service wall: each its own line item with its own quantity?☐
9Traceable derivationIs the calculation shown for every quantity? Example: “124.06 m × 3.40 m × 0.30 m = 126.54 m³”.☐
10Automatic item assignmentDoes every line item reference a number in the standard specification? What happens to work the specification does not know?☐
11ÖNORM / ONLVDoes the system export a data file according to ÖNORM A 2063 that your tendering software imports without rework?☐
12Assumptions flaggedCan you tell which quantities are stated in the plan, which are derived and which are assumed?☐
13Errors/uncertainty flaggedDoes the system say what is missing from the plan, and what that may cost? Or does it output a number everywhere?☐
14Plan revisionsWhat happens when you later upload plan revision B? New calculation, a difference report, or start from scratch?☐

Stage 3: five questions that expose a demo

1. “Where does this quantity come from?”

The most important question of all, for every line item. If it says 43.7 m² of interior plaster, the system must be able to name the calculation and the plan sheet. A figure without a derivation is an estimate with decimal places.

2. “How does the system know that interior plaster has to be included here?”

The plan says “external wall 25 cm brick”, but not “interior plaster 15 mm”. If the answer is “We have build-up and work rules on file for that”, good. If it is “The AI assumes, based on its experience, that …”, be careful. For a tender, a plausible assumption is not yet a correct quantity takeoff.

3. “Do you deduct windows and doors? From what size? Under which rule?”

Take a wall with several windows and doors. This is where a good takeoff system separates from a nice interface. A flat-rate deduction is acceptable too, as long as it is stated openly in the derivation and you can check it against your measurement rule. A hidden deduction that nobody can explain is not acceptable.

4. “Can the system tell these four wall types apart?”

25 cm brick, 12 cm brick, reinforced-concrete wall, drywall: four build-ups, four line items, four quantities. If everything ends up under “walls m²”, the bill of quantities is not fit for tender.

5. “What exactly happens during the processing time?”

If the vendor quotes under an hour for a single-family house and two to three hours for a larger project, ask: does the project run through fully automatically, or does that include manual checking or reworking by the vendor’s staff?

Both can be a good product. But “upload → 3 hours → finished result” and “upload → 3 hours → staff rework it for another two hours” are two completely different things economically once you estimate 20 projects a year. It also decides whether the result is reproducible.

Eight questions for every vendor

Independent of the demo, these questions belong in every first meeting. The last one is the one that counts.

  1. How many real projects have been estimated with the system so far?
  2. How many line items have actually been cross-checked by people?
  3. What is the average quantity deviation per trade?
  4. On average, how many line items are recognised correctly and automatically?
  5. How does scale and geometry recognition work with scanned PDFs?
  6. How is missing information handled?
  7. Which parts are done by AI and which are rule-based?
  8. Can I have one of my own projects estimated as a blind test and compare the result with my existing bill of quantities?

If the vendor answers question 8 with “Yes, of course, give us a real plan”, that is a very good sign. If they dodge it, that tells you something too.

The evaluation criterion that counts

Do not look only at the hit rate. A system that is right on 95 % of line items and clearly marks the other 5 % as uncertain is far more valuable than one that spits out some number for 100 % of line items. You can check the first in an hour. The second you have to recalculate completely.

In practice: check whether the system distinguishes between quantities stated in the plan, derived and assumed, and whether, for assumptions, it says what is missing from the plan and what that may cost. Missing structural calculation, missing soil report, missing demolition plan, missing window schedule: these are the points that eat into the margin later. A system that names them is a tool. One that calculates straight past them is a risk.

And BauKIT?

We did not write this list because it is convenient for us. We wrote it because we want to pass it, point by point, on your plan. Every quantity in a BauKIT quantity takeoff carries its calculation, its source in the plan and one of three confidence classes; missing information starts with “NICHT IM PLAN:” (not in the plan) and a reason. The processing runs fully automatically, with no rework by us.

Whether that is enough for your projects is not for us to decide. Bring this list and let us estimate a plan you have already calculated: request a blind test.

FAQ

Frequently asked questions

What is the best way to test AI quantity takeoff?

With a blind test: have the system estimate a plan you have already estimated yourself, without showing the vendor your own bill of quantities, then compare line by line, not just the total. 10 to 20 line items measured by hand beforehand are enough to document the deviation per line item.

Why is the total not enough as a test criterion?

Because errors can cancel each other out. A total can be right by chance while individual quantities are clearly wrong. What matters is the deviation per line item and per trade.

What matters more: hit rate or flagged uncertainty?

Flagged uncertainty. A system that is right on 95 % of line items and clearly marks the remaining 5 % as uncertain is more useful in practice than one that outputs some number for 100 % of line items.

Which question should you definitely ask the vendor?

“What exactly happens during the processing time: does the project run fully automatically, or does someone at the vendor rework it manually?” Both can be a good product, but economically they are completely different once you estimate 20 projects a year.

How can you tell whether quantities are rule-based or guessed?

From the answer to “How does the system know that interior plaster has to be included here?”. If the answer is “from build-up and work rules on file”, that is reliable. If it is “the AI assumes so based on its experience”, be careful: a plausible assumption is not yet a correct quantity takeoff.

See BauKIT work on one of your own plans.

Book a no-obligation demo. We run BauKIT on one of your permit plans and discuss how it fits into your workflow.