# How I Built an Original-Research Brief on Construction Change Orders

## Project overview

I created *What Counts as a Change Order?* as a ten-page example research brief for a construction-technology audience. The goal was to see whether public federal contracting data could support a useful, credible piece about the cost and management of change orders.

The project combined quantitative analysis, manual record review, source verification, editorial development, document design, and quality control. I used public sources only, without access to proprietary company data or subject-matter experts.

I also used two AI systems during the work: Claude Fable 5 and OpenAI’s GPT-5.6 Sol. I assigned them research, analysis, drafting, production, and checking tasks at different stages. I remained responsible for the research question, the instructions, the interpretation of the evidence, and every editorial decision. When an AI-produced claim, count, or sentence did not hold up, I corrected it or sent it back for another pass.

## Finding the real story in the data

I began with a straightforward benchmark question: What do change orders typically cost?

USAspending transaction files appeared to offer a way to answer it. They include contract action codes, dates, descriptions, and federal obligations. Once I examined the records more closely, however, I found that the official labels could not support a clean industry benchmark. Actions coded as change orders did not always describe substantive changes, while adjacent codes sometimes included work that looked change-related.

That changed the premise of the brief. Instead of forcing the data to produce a simple benchmark, I focused on what the classification problem revealed:

> Public data can measure coded contract actions precisely, but inconsistent classification limits what those figures mean. Contractors need a change record that preserves evidence and commercial status across every handoff.

The weakness in the dataset became the main finding. It also connected the analysis to a practical operating problem: the same field event may be described differently in a photo log, RFI, contract action, invoice, and portfolio report.

## Building a reproducible dataset

I used USAspending contract transaction files for the General Services Administration from fiscal years 2015 through 2025. I limited the analysis to construction NAICS prefixes 236, 237, and 238.

The source panel contained:

- 130,782 transactions
- 40,212 identifiable base awards
- 30,855 mature firm-fixed-price awards in the full size-band cohort
- 1,481 mature firm-fixed-price awards with at least $1 million in original obligations
- $6.57 billion in original obligations in the headline cohort

Several methodological choices mattered. I treated each row as a transaction within an award, not as a complete project. I identified base awards through modification number 0, joined later actions to the base award using the award’s unique key, and determined completion from the latest recorded period-of-performance end date across all transactions. I retained both obligations and deobligations.

I also treated the action codes as records of contract actions, not as proof that each row represented a unique underlying event. That distinction became important when I began reading the transaction descriptions.

Once the cohort and calculations were stable, I froze the analysis outputs and recorded them in a reproducibility manifest. This gave me a fixed source of truth for drafting, fact-checking, and revisions.

## Stress-testing the official taxonomy

I calculated the results using three nested definitions:

| Definition | Included action codes | Share of original obligations |
|---|---|---:|
| Narrow | D: Change Order; L: Definitize Change Order | 3.24% |
| Moderate | D/L plus B: Supplemental Agreement for work within scope | 17.07% |
| Broad | D/L/B plus A: Additional Work | 17.36% |

I did not present these as low, middle, and high estimates of the true cost of change orders. The spread is too large, and the categories are too noisy, for that interpretation. I used them as a sensitivity test. The measured result changed dramatically when adjacent action codes were added, which showed how strongly the answer depended on classification.

## Reading the records instead of trusting the labels

The coded analysis told me there was a definition problem. To understand its character, I created a 200-record editorial audit:

- 80 randomly selected D/L records
- 20 high-dollar D/L records
- 80 randomly selected B records
- 20 high-dollar B records

I kept random and high-dollar samples separate. The random records provided a representative cross-section of the descriptions. The high-dollar records served as a diagnostic check for individual transactions that could materially affect the aggregate.

Within the random D/L sample, I classified:

- 31 records as substantive scope or cost changes
- 8 as scope-and-time changes
- 21 as time-only changes
- 12 as indeterminate
- 8 as administrative, option, closeout, or apparently unrelated actions

That left 39 of 80 descriptions with clear substantive scope content. For context, an 80-record random sample has an approximate 95% margin of error of plus or minus 11 percentage points.

The high-dollar review exposed the limits of the taxonomy more clearly. One $10.6 million option exercise carried a D change-order code. The sample also contained $36.9 million net in administrative or correcting records and $7.5 million net in records whose public descriptions were too thin to classify.

The B-coded sample was even less suitable for a change-order estimate. Nine option records accounted for $658.1 million, while five administrative or clause records accounted for another $299.9 million. That evidence ruled out the B-inclusive 17.07% figure as a defensible upper estimate of change-order cost.

## Controlling the evidence before drafting

Before writing the report, I organized the research into three working tools:

1. A source ledger covering government, practitioner, insurance, consolidation, white-paper, and product sources.
2. A claim ledger listing each important assertion, the sources behind it, its verification status, any required qualification, and the planned backnote.
3. An audit workbook combining the headline metrics, definition test, manual classifications, claim status, source status, and page-by-page evidence outline.

This helped me keep calculation, interpretation, and explanation separate. A number can be computed correctly and still be used to support a claim that goes beyond the evidence. The ledgers made that gap easier to see.

They also gave me a practical way to direct the AI systems. I could assign a bounded task, such as checking one cohort definition or tracing one sentence to its source, then compare the result with the frozen analysis rather than accepting a broad response at face value.

## Shaping the argument for a short report

I planned the brief page by page before drafting:

1. Cover and premise
2. Executive brief and central finding
3. One event, several identities
4. Method and cohort construction
5. Noise inside the D/L codes
6. Coded-action incidence by award size
7. The definition gap
8. A practical response
9. Limitations and backnotes
10. Sample call to action

The opening follows a field condition as it becomes a photo, an RFI, a contract action, an invoice, and a portfolio-level report. That gave readers a concrete way into the classification problem before they reached the methodology.

For the practical section, I synthesized a seven-stage Change Evidence Standard:

1. Condition
2. Evidence
3. Direction
4. Impact
5. Classification
6. Resolution
7. Closure

The recommendation grew out of the research finding. If classifications become unreliable as information passes between systems and teams, the useful response is a record that keeps the evidence, direction, impact, and commercial status connected.

## Producing and checking the document

I developed the brief as a designed Word document with ten original visual assets, including the cohort funnel, definition ladder, audit charts, field-to-portfolio handoff, limitations matrix, and seven-stage standard. I kept the visual system restrained so the charts clarified the argument instead of turning the paper into a slide deck.

The report used claim-level citation markers and backnotes. Any publisher and call-to-action elements were clearly labeled as sample material.

After each significant revision, I had Codex render the document to page images and inspect it for pagination, clipping, overflow, spacing, citation density, and small labels. I also checked the document structure, headings, and image alternative text.

The checking process caught substantive issues as well as layout problems, including a byline error and a double-count in the narrative summary of the audit. I corrected the latter to 31 scope-or-cost records plus eight scope-and-time records, 39 total. I applied the same standard to editorial feedback: verify the premise first, then decide whether the change improves the piece.

## Result

The finished piece is a ten-page research brief built from public data, with a documented cohort, a sensitivity analysis, a manual description audit, visible limitations, source-level backnotes, and a practical operating framework.

The strongest part of the process was the decision to let the evidence change the story. I started by looking for a benchmark. When the classifications could not support one, I investigated it, quantified it, and built the brief around what the data could actually show.