The failed prompt myth and hidden limits
The presenter opens with a familiar fantasy: drop a set of drawings into a chatbot, ask for a flawless estimate, and collect the price. He warns that a contractor who prices work that way may soon be out of business. The blunt questions follow immediately: how would a general model know your crew rates, your usual inclusions and exclusions, or how to handle gaps and contradictions in the bid set? A bare prompt has no memory of your business, so it invents what it does not know.
Beyond the obvious gaps sit quieter failures the presenter measured himself. The best model he tested, GPT6 Astra, scored only about 48 percent on simple takeoff plans, and drawings filed in the wrong folder or given misleading labels often went unread. Even flagship reasoning models such as Opus and Fable stumbled when asked to connect contract language with estimate lines: told the site car park sat far from the workface, they rarely raised labour rates to cover lost walking time. The lesson is that fine-grained professional judgment, not arithmetic, is where raw models fall short.
Outside checks point the same way. A study shared via arxiv reports Handoff-H1 reaching an 81.6 composite on TakeoffBench against a 77.6 human reference, while frontier models land between 35 and 61 and GPT-4.1 scores 75.4 percent on CEQuest. A provision trial in which Claude priced a job at 443K against a 907K human figure, 51 percent low, found concrete within 2 percent and steel counts exact, yet secondary beams 45 to 96 percent light. And a Togal AI study kept in the exa library found area quantities systematically low, counted items exact, and ceilings the most consistent trade.
The presenter distills this into a thesis: an estimator is not a calculator but the person who secures profitable work through judgment, risk control, and commercial awareness. The answer is therefore an end-to-end system that pairs human strengths with machine strengths. Software can sweep 500 pages of bid documents in minutes and flag what a busy team missed, a job that might take people weeks. Deciding the delivery methodology, carrying the risk, and standing behind the margin stay firmly with the estimator.
System map and business context
He breaks the workflow into repeating phases rather than one giant prompt. Standing information about the company forms the business context , while job-specific files form the project context . Each bid then moves through go-no-go screening with a concept estimate, measured takeoff, line-item pricing, mapping to the client schedule, submission, and finally a learning loop that feeds results back. Two background routines run throughout: a correspondence tracker that logs client instructions, and a refresh step that keeps the project files current. Every phase hands a checked output to the next.
The cost library is the heart of the business context. Zebel-style parametric concept budgets built from historical project data give early direction, while an ediphi-style benchmarked project database keeps comparisons apples to apples by aligning scope, units, and tax treatment. For storage he suggests starting with a plain spreadsheet and graduating, once files swell past roughly 10 to 20 megabytes, to a structured base such as a claribase Airtable store that both humans and models can query. The payoff he stresses is a single source of truth where estimated and actual costs sit side by side.
Linked to the library is a register of assemblies that turns one measured item into a full resource list. A single light fitting, the primary quantity , expands into secondary quantities : the fitting itself plus wastage, a support bracket, metres of cable, electrician hours, and scissor-lift hours. Keeping these standard recipes in one template means every estimator prices the same activity the same way. The presenter argues this discipline matters more than clever software, because inconsistent recipes silently corrupt every later comparison.
The glue is a set of small reusable skills plus a short agents file that maps the folders. The claude official Agent Skills documentation defines a skill as packaged instructions that activate when their description matches the task at hand, which is exactly how he uses them: one skill per step, each with a narrow job and a verifiable output. An agents file at the project root tells the model where drawings, costs, and correspondence live and which rules to obey. Because each skill is checked before the next runs, errors compound far less than in one long autonomous chain.
Project context and triage
Every new bid starts by cloning a template folder and running a project indexer skill that catalogues what arrived, flags gaps, and drafts the agents file. The presenter insists on reading the scope and key drawings personally for a few hours before asking the model anything, judging whether the client pack is coherent or a pile of contradictions. Connectors then link the working folder to business systems such as SharePoint, mail, and takeoff tools through MCP links, so the model reads live data instead of stale copies. Each artefact gets one defined home to stop spreadsheet and software versions drifting apart.
Drawings get special treatment because misfiled sheets are a leading cause of model error. A drawing QA pass checks scale, orientation, and whether a file labelled electrical actually holds electrical content, followed by a query tool for targeted questions. His favourite aid is a drawing atlas view that stitches separate sheets, such as a road corridor or an architectural set with section links, into one navigable visual. Seeing the whole job at once makes missing sheets and mismatched sections obvious before any quantities are lifted.
The go-no-go screen then weighs the job against capacity, cash, workload, and risk appetite. Its key input is a parametric concept estimate: the model drafts a work breakdown, describes each activity, and matches it to the closest historic cost, guessing only where no record exists. His rule of thumb is blunt, a contractor turning over 10 million a year should think hard before chasing a single 20 million bid. The output is an order-of-magnitude budget plus a recommended bid or decline, so weak pursuits die before they consume estimating weeks.
Correspondence becomes a searchable database rather than scattered inboxes: client emails, minutes, and addenda flow into the project folder either through a mailbox connector or a scheduled sweep, so new instructions and conflicts surface automatically. From the same foundation a bid plan skill extracts every deliverable, internal and client-facing, into a bid register with owners, milestones, and review dates, and re-runs the pass whenever an addendum lands. Internal management reviews sit in that register as first-class deliverables, with an estimate review sheet and sign-off gates that force a second pair of eyes onto method, margin, and risk before submission.
Takeoff and pricing
On takeoff his philosophy is crisp: the model does not do the counting, it prepares everything around the counting. It builds the item list, checks the pack, and assembles context, while the estimator still verifies the measures. He is sceptical of expensive dedicated takeoff platforms for smaller firms, noting the same structured lists can live in connected sheets and files when budgets are tight. What matters is not the badge on the tool but a clean handover where every measured line traces back to a drawing and a recipe.
The primary-secondary distinction carries the estimating logic. Primary quantities are what you count on the sheets, fittings, metres of pipe, cubic metres of concrete. Secondary quantities are everything each primary drags with it: fixings, cable, wastage, labour hours, plant hours, and testing time. Because ratios such as cable metres per fitting or hours per cubic metre encode real crew experience, getting them wrong moves the whole price. He treats these ratios as company knowledge to be guarded, reviewed, and updated from job feedback.
A construction takeoff skill turns that logic into action. It reads the project documents and the standard assembly templates, then generates the job-specific takeoff list, for example seven light-fitting variants B1, B2, C1, and C2, each pre-loaded with its labour, plant, material, and subcontract lines. The estimator then works through the sheets doing the counts item by item. Copying and tweaking dozens of near-identical assemblies by hand is exactly the tedious setup work he is happy to delegate, provided a human still checks the resulting list.
Pricing converts those quantities into money through a workbook where each line sums labour, plant, materials, and subcontract cost, with quantities joined to a resource library by simple lookups. A line-item pricing skill drafts this first pass from the bill of quantities, the project scope, and the agreed methodology. He is candid that many firms may prefer to price large or risky lines themselves and let the model draft only the long tail of small items under an eighty-twenty rule. Either way the draft is a starting point for review, never a finished price.
Rates then get tuned for reality. A productivity factor converts pure task time into paid hours by adding breaks, inductions, access delays, and other non-wrench time the raw measure ignores. Project specifics adjust further: quoted subcontract figures replace generic allowances, local hire rates apply, and all-in labour loadings for overtime and on-costs are verified per hour. The presenter stresses reconciling units and tax treatment line by line, since mixed gross and net figures are a classic silent killer of apparent competitiveness.
Before numbers harden, a requirements register and its partner clarification skill hunt for trouble. The register extracts every product and process obligation, from the concrete slab on a named drawing to the monthly programme submission, into one checkable list. The clarification skill then scans for contradictions, omissions, and clashes and drafts concise requests for information stating the source conflict and the question. He cautions against firing thirty queries at a client unread: each draft must be read and owned first, or the sender looks careless on the first phone call.
Submission and learning
Submission starts by mapping the internal build-up onto the client pricing schedule, the document that defines how the contractor gets paid. If the client breakdown omits scope, the presenter adds lines rather than absorbing the loss, since anything missing from the structure goes missing from the price. Firms with repeat work may keep their own standard breakdown and let the model map it across instead. Either route ends with the same sanity check he repeats like a mantra: does the total reconcile, and would you sign a cheque sized like this one, down to the last two hundred thousand dollars?
Commercial review runs in parallel. A contract skill compares proposed terms against the company standards and logs every departure in a register with its pricing consequence, while payment terms feed a cash-flow forecast built from the labour, plant, material, and subcontract burn against the schedule of values. The forecast reveals the maximum deficit and whether the business can fund it, which can send the team back to renegotiate terms or reshape the pricing schedule. He calls this an ideal machine task: everyone knows cash matters, yet bid timetables rarely leave room to model it.
The non-price deliverables get equal rigour: a methodology skill drafts the construction approach, organisation chart, and client-requested statements from approved templates. A compliance check then confirms every requested document is present, correctly named, and free of obvious errors, while an estimate review compares the detailed price against the early concept at both project and trade level. A concrete rate of five hundred per cubic metre against a historic fifteen hundred, for instance, must explain itself. Only after these gates pass does the letter of offer go out, aligning exactly what the price includes and excludes.
After submission the system must learn or the next bid starts cold. Outcomes, client feedback, market price versus tendered price, and any missed scope found on site are recorded as lessons that update libraries, assemblies, and skills. The presenter notes that his packaged implementation of this whole framework, Contractor OS, collects the guides, templates, and workflows in one place for firms that want a head start. Whether bought or built, the principle stands: every bid should leave the database smarter than it found it, which is the whole point of the learning loop .
He closes by restating the failure mode he opened with: without structured context and a staged process, even brilliant models produce confident, uncheckable numbers. Drawings in and price out feels fast until the missing car-park walk, the misfiled drawing, or the untaxed rate quietly removes the margin. Structure is what converts a clever demo into a repeatable commercial function. Firms that skip it are not saving time; they are borrowing trouble against the final account.
For readers wondering where to begin, his advice is deliberately modest: pilot the system on one small live job, not the whole pipeline. Start today by assembling the cost library from past estimates, since consistent historic estimates beat messy actuals for a first version. Add one skill at a time, verify each output with a human in the loop , and only then connect the next. Small, checked steps compound into the end-to-end machine the guide describes, and the first pilot pays for the second.
Key moments
- Drawings in, price out: the risky fantasy
- Hidden limits: 48 percent takeoffs
- Estimator as judge, not calculator
- Five phases and human checkpoints
- Business context and cost library
- Assemblies: fittings to full recipes
- Skills and the project map file
- Drawing atlas: seeing the whole job
- Go no-go and the concept estimate
- Correspondence and the bid register
- Takeoff prep and line-item pricing
- Submission gates and the learning loop
AI commentary
"The guide is valuable because it treats AI as junior staff inside a strict process rather than an oracle. Readers should copy the verification gates and the libraries, not just the prompts."
AI assessment
The fair counter-view is that staged checking slows bids down and leans on benchmarks run over small, simple plan sets. The arxiv-shared Handoff-H1 result of an 81.6 TakeoffBench composite against a 77.6 human reference, with frontier models at 35 to 61 and GPT-4.1 at 75.4 percent on CEQuest, is encouraging but narrow, and independent replication on messier commercial packs is still needed before treating any model as takeoff-safe.
Two gaps stand out. Earthworks and other high-variability civil quantities go largely untested in the cited studies, which cluster around building trades, concrete, steel, and finishes. And the guide offers no measured dollar outcome: no win-rate lift, margin change, or estimating-hours saved, so the two-hundred-thousand-dollar reconciliation story reads as craft wisdom rather than evidence.
Readers should also weigh the presenter's commercial interest. The full skill pack, templates, and step-by-step implementation live inside a paid product, Contractor OS, so demonstrations naturally show the system at its best. That does not invalidate the method, but it means claims about effort saved and errors caught deserve the same verify-before-trust treatment the guide demands of AI output.
The practical takeaway survives those caveats: build the cost and assembly libraries first, run one small pilot with human gates at every step, and record outcomes so each bid improves the next. Adopt the discipline of structured context and staged verification even if you never buy a template. That operating habit is the durable asset; any single model will be replaced within months.
Sources
8 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — AI for Construction Estimating guide
- @arxiv.org arXiv — Handoff-H1 TakeoffBench report
- @provision.com Provision — Claude estimate accuracy test
- @exa.ai Exa — Togal AI quantity takeoff study
- @claude.com Claude — Agent Skills overview
- @zebel.io Zebel — Historical construction database
- @claribase.com Clarifbase — Airtable construction management
- @ediphi.com Ediphi — Conceptual cost modeling
construction estimating · artificial intelligence · quantity takeoff · claude skills · contractor business