Case study · Original research

What Counts as a Change Order?

A self-directed white paper built entirely from public federal data, along with the full content program planned around it.

  • Original Research
  • Content Strategy
  • Editorial
  • Thought Leadership
A field condition flows into a photo, RFI, contract action, invoice, and portfolio metric.
The cover figure: one physical condition on a jobsite becomes five separate administrative records, each with a different label.

Start with the source material

Read the ten-page brief, or take the dataset and code now, then come back for the story behind them.

The artifact is a ten-page original research brief, built from public data. It analyzes federal construction contract actions to ask a plain question and defend a harder answer. I wrote the copy, ran the analysis, made the figures, and designed the document.

130,782
federal contract actions analyzed, from eleven fiscal years of public GSA construction data
1,481
completed $1M+ contracts in the headline cohort
30.6%
carried an official change-order label
48.8%
of audited descriptions (39 of 80) clearly changed the work

The deliverables index has every file. The rest of this page is how the brief came together.

The brief I set myself

I wanted a piece that would survive a skeptical read, so I made the rules deliberately hard. Nothing in the finished document could fall apart if a reader stopped to ask where a number came from.

  • Public sources only. No proprietary or internal company data.
  • No new interviews. The argument had to stand on records anyone can pull.
  • Ten pages, tight enough to respect a busy reader.
  • At least one original visual on every page.
  • Distributable as real work.

Choosing the question

The most difficult decision when preparing for a research project is choosing which question to ask. I developed six candidate topics, each paired with a public dataset that could support it, and ran each through three filters.

Filter 1

Is there a surprise?

Does a counterintuitive finding plausibly exist in the data, or would the work only confirm what people already assume?

Filter 2

Can one person do it?

Can a single researcher run the analysis with public tools, without a team or a data budget behind them?

Filter 3

Would someone publish it?

Is there an identifiable organization that would put its name behind the finding, so the piece reads as a real property?

Change orders in construction cleared all three. The topic promised a surprise, public federal data could reach it, and it speaks to a real audience of contractors and the software companies that serve them.

Following one change from the field to the report

At 9:17 on a Tuesday morning, a foreman opens a wall and finds conduit where the drawings show open space. He photographs it, calls the project manager, and moves the crew. By lunch the condition is a text thread. The project manager turns it into a formal question about the drawings. The contracting office later records a modification. Finance sees a cost adjustment. An executive sees one line in a quarterly report.

The physical condition happened once. Its label changed at every handoff. That relabeling is what the paper is really about, and it is the reason numbers in this industry can be so misleading.

A six-stage chain shows a field condition moving through evidence, direction, contract action, billing, and portfolio reporting, with three points where context can be lost.
A single field condition loses context at each handoff between the crew, the contract, and the report.

The turn

The thesis changed twice as I worked, and each version was smaller and more focused than the last. This is where the editorial work actually happened: deciding what the evidence could support, then publishing that narrower claim instead of the bigger, louder one.

  1. Public data can reveal what change orders cost.

    Where I started. The federal contract files looked like a clean path to an industry benchmark.

  2. The most-repeated industry statistic will not hold up.

    First contraction. The familiar claim that change orders run 8 to 14 percent of contract value could not be traced to a primary source, and the best available research pointed much lower.

  3. The category itself may not mean one consistent thing.

    Second contraction. Reading the records showed the official labels mix substantive scope changes with schedule extensions, options, and administrative actions. The question moved from what change orders cost to whether the label measures anything stable.

Each step led to a more defensible position. The final claim is narrow, and anyone can check it against the same public records.

The method spine

Credible research needs a backbone that ties every claim to its evidence. Four working tools governed this paper, and each one kept a different part of the work honest to the same underlying records.

01

Source ledger

Every source catalogued with its type, owner, date, and planned use, so no citation entered the draft unexamined.

02

Claim ledger

Every factual assertion carried an ID, its source references, a verification status, and a required qualifier. A claim without a source ID stayed out of the draft.

03

Frozen analysis

Results were locked with a manifest of hashes, a random seed, and a version, so drafting referenced stable outputs instead of a moving calculation.

04

Manual audit

200 records hand-classified against a documented codebook, with random and high-dollar samples kept separate so one could not contaminate the other.

The analysis narrows from 130,782 transactions to 1,481 headline awards, while the coding lens widens from D and L to include B and A.
The cohort funnel, and the three coding lenses used to stress-test the official taxonomy.

What the research found

Four findings carried the paper. Each one is stated here with the qualification it has in the brief, because none of these numbers means much on its own.

What the 3.24% actually counts

Across the 1,481 completed, firm-fixed-price GSA contracts that began with at least $1 million committed, actions carrying an official change-order code added $212.6 million net, or 3.24% of the amount originally committed. The math is reproducible. The label is less dependable. In a random sample of 80 descriptions, 39 clearly changed the required work, 21 changed only the schedule, and the rest were paperwork, options, or too vague to classify. One $10.6 million action mostly exercised construction options while carrying a change-order code.

Three panels show 30.6 percent coded incidence, 212.6 million dollars net, and 39 of 80 audited descriptions with substantive scope.
The executive summary figure: coded incidence, net obligation, and audit composition in one view.

The qualification

The 3.24% is a reliable total for records carrying the “change order” label. That label covers scope changes, schedule extensions, options, and paperwork together, so the figure does not indicate what change orders cost.

What the descriptions actually said

Reading the records by hand showed what the codes conceal. In the random D/L sample, 31 records clearly changed physical or professional work and eight more changed both work and schedule, for 39 total. Twenty-one changed the schedule only. The high-dollar review was starker. A single $10.6 million option exercise sat under a change-order code, alongside millions in administrative and correcting records.

A bar chart shows 31 substantive scope records, 8 scope-plus-time, 21 time-only, 12 indeterminate, and 8 other records; a callout identifies high-dollar administrative, option, and indeterminate actions.
The description audit: what 80 randomly sampled change-order records actually contained, with the high-dollar exceptions called out.

The qualification

The 80-record random sample carries an approximate 95% margin of error of plus or minus 11 percentage points. The high-dollar records are diagnostic, and are reported on their own rather than blended into the rate.

Incidence rises sharply with contract size

Official change-order labels appeared far more often as contracts grew, from 5.1% on contracts below $100,000 up through the size bands to 59.3% on contracts of $5 million or more. The gradient is real, and it is where record quality stops being a clerical concern and starts affecting what a company-wide report can honestly show.

Five bars rise from 5.1 percent for awards below 100,000 dollars to 59.3 percent for awards of at least 5 million dollars, with award counts shown.
Coded change-order incidence by original obligation band.

The qualification

This does not prove that larger contracts cause more changes. Bigger awards involve more time, more people, and more chances for a contract action. The data does not separate those factors or identify who initiated a change.

The answer depends on which codes you count

The same contracts produce very different numbers depending on which codes count. The narrow definition yields 3.24%. Adding within-scope supplemental agreements raises the apparent figure to 17.07%. Adding additional-work agreements moves it only slightly, to 17.36%. Reading the largest of those added records showed why the jump is a warning rather than a range. Nine of them established or exercised options worth $658.1 million, which are not changes at all.

Three definitions yield 3.24, 17.07, and 17.36 percent. Four cards show that options and administrative records dominate the high-dollar review.
The definition gap: three ways of counting the same records, and what sits inside the largest additions.

The qualification

Widening the definition to include options and paperwork produces that jump. No additional change orders turned up, which is why 17.07% cannot serve as an upper estimate of cost.

A practical response

A classification problem needs an operating answer, so the brief closes with one. The Change Evidence Standard is a set of seven fields that stay linked from the first field evidence through closeout: Condition, Evidence, Direction, Impact, Classification, Resolution, and Closure. It keeps schedule-only actions, options, paperwork, and true scope changes apart at the point of capture, which is the cheapest place to keep them separate.

Seven linked stages run from condition through closure and feed the project team, finance and risk, and executive portfolio reporting.
The seven-field Change Evidence Standard, and the three downstream readers it serves.

Fact-checking every claim

A research draft is not finished until someone has tried to break it. After the fourth editorial draft, I read the paper the way a hostile reviewer would and pulled out every statement that needs evidence, starting with the number in the title and ending with the contact details on the back cover. That produced 104 separate claims. The method spine’s claim ledger governed what went into the draft. This one tests what came out.

Each claim is a row carrying its risk level, the kind of check it needs, where it sits in the paper, every other place the same figure repeats, the source that should support it, and the steps to reproduce or confirm it. 78 are rated high risk, which covers the headline numbers, the quotations, the named organizations, and anything a reader is likely to repeat somewhere else. 30 are calculations that have to be rerun against the frozen analysis files rather than read off an earlier draft. Ten rest on outside publications, catalogued on a second tab with a publisher, a URL, and the specific passage the checker has to open.

A ledger view lists five sample rows with a claim number, risk level, check type, item to verify, and source ID, above four tallies: 104 claims extracted, 78 rated high risk, 30 calculations to reproduce, and 10 outside sources catalogued.
The claim-checking ledger: five sample rows and the shape of the whole review.

Finishing a row and believing a claim are tracked in separate columns, which matters more than it sounds. A row closes only after the evidence has been opened and a result has been recorded as verified, verified with qualification, revised and needing a recheck, unsupported, or unresolved. The presence of a citation never counts as a check, and no AI review is allowed to stand in for the original record.

Two pieces of this are portable. The phase-by-phase checklist works on any high-stakes document, and a pair of AI system prompts handles the mechanical part of extracting and testing claims while leaving the judgment with a person. Both are in the deliverables index, alongside the ledger itself with every row still open. The process expects an independent challenge pass on high-risk claims and a named approver for anything qualified or removed, and a solo project supplies neither (though all the claims in this work have been solo fact-checked). What the file shows is the instrument, not a finished verdict.

Errors, and how they were caught

Research writeups rarely list their own mistakes. This one documents three, because catching them is part of what makes the analysis trustworthy. Several AI systems reviewed the work at different stages, and when they disagreed, the frozen analysis files settled it.

  • A wrong byline

    An early draft carried an incorrect author attribution. A review pass caught it before the version was shared.

  • A double-counted audit statistic

    The narrative summary presented eight scope-and-time records as additional to a total that already contained them. Corrected to 31 scope-or-cost records plus eight scope-and-time records, 39 in total.

  • Two miscounted review claims

    A review itself introduced two miscounted style claims. Those were checked against the frozen outputs and corrected in turn.

The artifact

Though the research and work are real, the paper is published under a fictional imprint, Northline Field Research, so it reads as a finished property. Northline is a device for making the artifact feel complete. It is not a real client, and every publisher and call-to-action element in the document is labeled as sample material.

The brief’s cover: the title What Counts as a Change Order, published as a sample under Northline Field Research, with the field-to-report chain along the bottom.
Title
What Counts as a Change Order?
Format
Ten-page original research brief
Built from
Public USAspending GSA construction data, FY2015 through FY2025
Author
Colin Wright
Status
Fourth editorial draft, in fact-checking

What the analysis cannot show

The panel covers GSA awards only, so it is a transparent federal case study rather than a sample of the whole construction industry. Federal obligations are amounts committed, which do not reveal contractor margin, job cost, or final project cost. Transaction descriptions summarize rather than document, so an indeterminate record simply means the public text was insufficient. And a coded action is not proof of a unique event, since an initial direction and a later pricing action can describe two stages of one change.

A two-column matrix distinguishes coded-action measures from unavailable measures such as true change-order cost, contractor margin, unique events, and a private-industry benchmark.
The limitations matrix: what the data can support, and what it cannot.

From flagship to system

A paper like this works best as the seed of a larger content program. Before writing a single derivative, I mapped what the ecosystem could hold: 30 assets across six tiers, pulled from an inventory of the findings, frameworks, and narrative the brief already contains.

Five built, the rest mapped

Thirty assets were planned. Five are finished. Each one below is downloadable.

A one-page Change Evidence Standard reference listing the seven field stages from Condition to Closure.

Change Evidence Standard

The seven-field one-pager the brief promises on its back cover.

Download (SVG)
A social carousel slide showing the three definitions that yield 3.24, 17.07, and 17.36 percent.

Definition-gap carousel

The ladder finding recut from landscape to a portrait social slide.

Download (SVG)
An article page titled Why I stopped trying to calculate what change orders cost.

Lead article

“Why I stopped trying to calculate what change orders cost,” the series opener.

Read (Markdown)
A code repository preview showing the analysis script, methodology, and dataset files.

Reproducibility package

Script, methodology, derived panel, and manifest, packaged to reproduce.

Download (ZIP)
A six-slide webinar deck skeleton, from the cold-open story through the audit, the definition gap, and the standard.

Webinar deck skeleton

A forty-minute run of show and the opening slides, ready to record.

Download (SVG)

All five together

The built derivatives in one archive: each preview as SVG, plus the full text of the lead article.

The mapped ecosystem

The other twenty-five sit in the plan, organized by tier. All nine of the paper’s figures already exist and can be recut for each channel without redrawing.

Tier What it holds Assets
Companion downloads The one-page standard, a self-audit worksheet, an executive summary, a classification codebook, and a reproducibility package. 5
Article series Five standalone posts, each pointing back to the paper, led by the thesis-shift story rather than the number. 5
Social Carousels, single-anecdote posts, a size-gradient chart, a process thread, and a short screen-recorded video. 8
Live and spoken A webinar with a guest, a conference talk, a podcast guest kit, and the cut-downs from the recording. 4
Sales enablement Role-specific one-pagers, a discovery question set, an objection handler, and a benchmark conversation starter. 4
Earned and syndicated A trade-press pitch, a bylined article, an association feature, and the dataset as a citable resource. 4

The full breakdown, atom inventory, specimens, and launch sequence are in the derivative content plan, linked below.

The refresh path

Adding a second agency to the panel, or rerunning with a trailing year of data, turns a one-off into an annual property. That refresh step is what makes the research a repeatable asset rather than a single publication.

What proprietary data would change

This brief used public records only, with no internal data and no interviews, and it still produced a documented 30-asset program. Point the same method at a company’s own platform data and its subject-matter experts, and the result is research a competitor cannot reproduce.

Deliverables index

Everything the project produced, in one place.

  • The research brief PDF

    The ten-page paper, What Counts as a Change Order?

  • Dataset and analysis code ZIP

    The extraction and analysis script, methodology, and the derived award panel. The reproducibility package.

  • Atomized content, all five ZIP

    Every built derivative in one archive, plus the lead article in full. The individual files are listed below.

  • Change Evidence Standard SVG

    The one-page seven-field reference.

  • Definition-gap carousel SVG

    The ladder finding as a portrait social slide.

  • Lead article Markdown

    “Why I stopped trying to calculate what change orders cost.”

  • Webinar deck skeleton SVG

    The forty-minute run of show and opening slides.

  • Claim-checking ledger XLSX

    All 104 claims with risk level, check type, source IDs, and reproduction steps, plus the source library and the recommended order of work. Every row is still open.

  • Fact-checking checklist DOCX

    The reusable phase-by-phase process, written for any high-stakes document rather than this one.

  • AI fact-checking prompts DOCX

    Two copy-ready system prompts: one builds the claim ledger, the other verifies against the evidence package.

  • Process record Markdown

    A first-person account of how the brief was researched, checked, and corrected.

  • Derivative content plan Markdown

    The full 30-asset ecosystem, atom inventory, specimens, and launch sequence.