AI Market Sizing: A Defensible Bottom-Up Workflow

A market estimate can look polished and still collapse under one client question: “Where did that number come from?” A duplicated customer segment, an unsupported price point, or an unclear market boundary can turn a neat slide into a credibility problem.
That is why AI market sizing needs more than a clever prompt. Bottom-up analysis starts with observable units, then builds toward revenue or volume through explicit assumptions.
AI can speed up research, extraction, and model documentation. However, the consultant still owns the market definition, the evidence, and the final judgment.
Set the Rules for AI Market Sizing Before Research Begins
A bottom-up estimate usually follows a simple logic: count eligible buyers or units, apply adoption or purchase frequency, then multiply by a relevant price or spend level. The hard part is deciding what belongs in each variable.

Define one market and one unit of analysis
Write a one-paragraph market definition before asking AI to gather facts. State the product or service, customer type, geography, period, currency, and metric. For example, distinguish software revenue from total IT spending, or annual recurring revenue from first-year contract value.
A useful market equation might be:
Market revenue = eligible accounts x penetration rate x average annual spend
The bottom-up market-sizing methodology from Umbrex follows the same discipline: build from customers, units, adoption, and economics rather than starting with a broad industry total.
Lock down exclusions, too. If the estimate covers independent clinics, don’t let hospital systems reappear in another segment. If pricing excludes implementation fees, keep them out of every calculation.
Build an assumption register before the model grows
Create one row for every assumption, even obvious ones. Include its value, unit, source, source publication date, retrieval date, rationale, model owner, and confidence rating.
This register prevents a common consulting problem: a figure migrates through slides and spreadsheets until nobody remembers its origin. It also gives reviewers a fast way to challenge inputs without rebuilding the model.
A source URL alone isn’t an audit trail. The model needs the exact table, page, data year, definition, and date retrieved.
Give AI Structured Research Tasks, Not Open-Ended Questions
AI is useful when the task is narrow and the required output is clear. It can scan a source pack, extract candidate figures, normalize units, identify gaps, and draft source notes. It should never become the final source of a market input.
Feed the model a controlled source pack
Start with primary and near-primary materials where possible. Government datasets, company filings, regulatory registries, trade associations, procurement records, and customer data usually carry more weight than a broad market report summary.
For each source, capture the publisher, title, URL, date, geography, population covered, metric definition, and relevant page or table. AI can turn this material into a research log quickly, but a human should open each cited source and verify the claim.
This matters because AI hallucinations can include fabricated facts and citations. A plausible link, a confident statistic, or a familiar company name isn’t evidence.
Ask for extraction, comparison, and gaps
Use prompts that ask AI to separate facts from interpretation. For example:
“Extract every reported count of eligible establishments from the attached sources. Return a table with value, unit, geography, reference period, source title, page or table, exact supporting quote, and conflicts across sources. Do not infer missing values.”
Then ask a second question:
“Compare these establishment counts against the market definition. Flag overlaps, exclusions, stale data, incompatible geographies, and any input that cannot support a bottom-up calculation.”
This workflow keeps AI market sizing grounded in visible evidence rather than broad web summaries.
Convert Research Into a Reviewable Spreadsheet Model
A client should be able to trace every headline number back through formulas to a documented input. Build the spreadsheet before you write the story.
Use columns that expose the calculation
Keep raw inputs separate from calculations, and calculations separate from slide-ready outputs. A compact model structure can look like this:
| Model field | What it records |
|---|---|
| Segment | A mutually exclusive buyer or unit group |
| Eligible base | Count of customers, sites, users, or units |
| Adoption | Share expected to buy within the defined period |
| Spend | Price, annual spend, or units per buyer |
| Formula | The cell logic that creates segment value |
| Source ID | Link to the assumption register |
| Low, base, high | Sensitivity values for uncertain inputs |
Each segment should have a unique inclusion rule. If enterprise accounts include subsidiaries, don’t also count subsidiaries as separate mid-market accounts. Double counting often hides in segment definitions, not in arithmetic.
Make AI outputs fit the workbook
When a workflow is API-based, OpenAI’s Structured Outputs documentation describes how a supplied JSON Schema can require consistent fields. That is helpful for turning extracted data into repeatable rows such as source ID, metric, geography, date, unit, and confidence.
Still, a structured response is not a validated response. The model can place an incorrect value into every required field. Reviewers must test the source, units, formula references, and market fit.
Keep narrative text outside calculation cells. The spreadsheet should calculate the number; the slide should explain why the number is reasonable.
Triangulate the Estimate and Show the Range
A bottom-up build gains credibility when independent evidence points in the same direction. It loses credibility when one fragile assumption drives the entire result.

Test the model against independent evidence
Compare the estimate with a top-down industry total, public-company revenue, channel capacity, procurement volumes, or known customer counts. The totals won’t match perfectly because definitions differ. The comparison should reveal why they differ.
For example, a bottom-up estimate may cover a narrow product category while an industry report includes services and adjacent products. Document that definition gap instead of forcing a false match.
IBM’s explanation of AI hallucinations is a useful reminder that fluent output can still be wrong. Treat AI-generated benchmarks as leads to verify, not evidence to cite.
Use sensitivity ranges instead of false precision
Choose the two or three assumptions that move the result most. Adoption rate, eligible accounts, and annual spend often matter more than minor cost inputs. Set low, base, and high cases with a written rationale for each value.
A client-ready slide should show the base case, range, and key drivers. Avoid a headline such as “$842.7 million” when source quality supports only a rounded range.
If a 10 percent change in one assumption moves the market by 40 percent, the assumption belongs on the slide.
Prompts That Produce Audit-Friendly Work
Good prompts tell the model what it may use, what it must return, and where it must stop. They also make uncertainty visible.
Use a source-first research prompt
Paste the following prompt after attaching approved files or a source log:
“Act as a research assistant for a bottom-up market model. Use only the supplied sources. For every candidate input, provide the value, unit, date, geography, source ID, exact citation location, and a short definition. Label unsupported claims as ‘not verified.’ Do not estimate missing values.”
This prompt blocks a frequent failure mode: AI filling a blank cell with a believable number. It also makes peer review faster because each input arrives with a traceable record.
Ask AI to audit the model logic
Once the spreadsheet is populated, provide a clean export of assumptions and formulas:
“Review this market-sizing table for double counting, inconsistent segment definitions, unit mismatches, date mismatches, circular logic, and false precision. List each issue, affected rows, why it matters, and the human decision required. Do not change values.”
Ask a separate reviewer, either a colleague or a second model run, to challenge the same assumptions. Independent review is more useful than asking the original conversation to approve its own work.
Package the Analysis for the Client Team
A defensible deliverable lets a client update the estimate after the project ends. Keep the executive story concise, but attach enough detail for a finance, strategy, or research team to inspect it.
Use a simple five-part package:
- Present the market definition, headline range, base case, and decision implication on the first slide.
- Show the bottom-up logic with segment counts, adoption assumptions, and spend inputs.
- Include sensitivity cases and name the assumptions that drive the range.
- Attach the source register with citations, publication dates, retrieval dates, and confidence ratings.
- Deliver an unlocked spreadsheet with clear input cells, formulas, version date, and change log.
AI market sizing works best when the model supports the analyst’s reasoning rather than replacing it. The final deliverable should make that reasoning easy to test.
Final Thoughts
A reliable bottom-up estimate is built from defined units, sourced assumptions, transparent formulas, and honest ranges. AI can reduce the time spent extracting, organizing, and checking material, but it can’t certify a market claim.
Keep every important number traceable to a dated source and a visible calculation. That discipline turns AI market sizing into work a client can review, challenge, and use.