← Back to Projects Industry Research

Regional Synthetic Biology Industry Sizing Representative-Company Extrapolation and Multi-Source Calibration

Reconstructed the estimation scope and built a regional emerging-industry sizing framework using representative-company analysis, concentration-assumption sensitivity testing and multi-source calibration.

Independently completed the research, method design and estimation iterations for the industry-sizing component.

Industry SizingSynthetic BiologyCompany ResearchData Calibration

Background and Core Question

Emerging industries often lack a dedicated and stable official statistical classification. When a regional planning project required an interpretable industry baseline, the available evidence was dispersed across broader industry statistics, park-level records and enterprise data.

The target city was the principal estimation scope, with park-level information used as a nested calibration reference and broader regional information serving only as context or reasonableness comparison. The central question was how to estimate the size of an emerging regional industry under incomplete statistical coverage, and to select an estimation method appropriate to the data conditions.

Work Completed

Work Completed

  • Company long-list construction through keyword search and public sources.
  • Technology-relevance screening and local-attribution assessment for candidate companies.
  • Initial segment-level model testing across value-chain stages and application areas.
  • Scope and definition reconstruction after identifying geographic, industry-boundary, planned-capacity and double-counting issues.
  • Representative-company output estimation and concentration-assumption sensitivity testing.
  • Multi-source calibration and neutral-estimate selection.

Scope Boundary

  • The wider planning report was a team deliverable; this case covers only the independently completed industry-sizing component. No claim is made to authorship of the full plan, spatial layout, fund structure, investment-promotion design or safety assessment.

Why the Initial Method Was Unsuitable

The initial approach was not a failure of bottom-up estimation as a principle. A complete segment-by-segment model of the full value chain was tested and found unsuitable for an early-stage local industry with a limited number of enterprises and incomplete data coverage.

The real issue was that the initial model mixed provincial-level projects with local activity, broad and narrow industry definitions, planned capacity with current output, and upstream inputs with downstream products — while also creating double-counting risks. Bottom-up analysis was retained at the representative-company level; only the estimation unit changed.

Method Evolution

The estimation proceeded through four phases. The initial segment-level model was found unsuitable for the data conditions; scope reconstruction and a shift to representative-company analysis produced the final framework.

  1. Data Foundation and Company Sample

    Collected multi-level industry, statistical and enterprise evidence. Built a candidate-company sample and screened for technology relevance and local attribution.

  2. Initial Segment-Level Model

    Tested a bottom-up estimate across value-chain stages and application areas using output, price, capacity and enterprise data where available.

  3. Method Diagnosis and Scope Reconstruction

    Identified geographic-scope mixing, broad/narrow industry-boundary mixing, planned-capacity vs. current-output mixing, incomplete data and double-counting risk. Reconstructed the estimation scope.

  4. Representative-Company Extrapolation and Calibration

    Estimated locally attributable output at company level, applied a concentration assumption with sensitivity testing, and calibrated the resulting range against broader official and aggregate reference points.

Conceptual diagram based on reconstructed methodology. No original report content is reproduced.

Method Selection

Before continuing the estimation, the industry's maturity and concentration were assessed. The local synthetic-biology industry was at an early stage, with a relatively concentrated company structure and incomplete segment-level statistics. Under these conditions, representative-company extrapolation — rather than full segment-by-segment aggregation — was the more appropriate approach.

This judgement is the central analytical value of the case. It demonstrates that method selection should follow an assessment of data conditions, industry structure and estimation purpose — not convention or precedent.

Method-Selection Logic

The choice of estimation method was driven by an assessment of industry maturity and data conditions — not by convention. In this case, the local industry was at an early stage with a concentrated company structure and incomplete segment-level statistics, making representative-company extrapolation the more appropriate approach.

Assess industry maturity and data conditions
Stable segment statisticsDispersed company structureRich historical data
Segment-level estimation
Incomplete segment statisticsConcentrated company structureLimited historical data
Representative-company extrapolation

Conceptual diagram. The method is not recommended as a universal template for all early-stage industries.

Company-Sample Construction

The company sample was built from a broad search universe using keyword searches and public sources. Companies were then screened by actual technology and business relevance — not by broad industry labels — to exclude firms with only tangential connections to synthetic biology. A further local-attribution assessment identified companies whose economic activity could be reasonably tied to the target geography.

The final representative sample consisted of companies for which sufficient public information was available to support individual output estimation and concentration extrapolation.

Company-Sample Screening Funnel

The sample was built from a broad search universe and narrowed through technology-relevance screening and local-attribution assessment to a representative set of companies.

Company Search Universe Companies identified through keyword search and public sources across the target region.
Technology-relevant firms Companies whose actual processes, products and revenue sources indicated genuine synthetic-biology or biomanufacturing activity.
Locally attributable activity Companies with economic activity that could be reasonably attributed to the target geography.
Representative sample The core set of companies used for individual output estimation and concentration extrapolation.

Illustrative funnel. Company counts are not displayed.

Estimation and Calibration

Locally attributable output was estimated company by company. A concentration assumption was applied with sensitivity testing under alternative conditions to extrapolate to total regional industry size. The resulting range was then compared with official statistics, public industry evidence and aggregate reference information available during the project.

The calibration step did not mechanically add or average the different sources. It used external reference points to test the plausibility of the company-level estimate, to identify scope mismatches and to establish reasonable boundaries for the final range.

Final Estimation Workflow

The estimation followed a structured eight-step workflow from scope definition to the final range and neutral estimate.

  1. Scope definition

    Define geography, time period, technical boundary and output measure.

  2. Company long-list construction

    Build a candidate list through keyword search and public sources.

  3. Technology-relevance screening

    Screen by actual process, product and revenue relevance, not broad industry labels.

  4. Local-attribution assessment

    Identify companies with economic activity attributable to the target geography.

  5. Company-level output estimation

    Estimate locally attributable output for each representative company.

  6. Concentration-assumption sensitivity testing

    Apply a concentration assumption and test sensitivity under alternative conditions.

  7. Official and aggregate-reference calibration

    Compare the resulting range with official statistics and aggregate reference information.

  8. Range and neutral estimate

    Produce a reasonable range and neutral estimate as a quantitative reference.

Conceptual workflow diagram. Does not display actual data or estimates.

Concentration Sensitivity and Calibration

Concentration is a core assumption, not a confirmed fact. Rather than selecting a single ratio — which would create false precision — alternative conditions were tested to understand how the estimate changed. The sensitivity testing made explicit how the result depended on the concentration assumption.

Broader aggregate reference points were then used to assess which parts of the resulting range were plausible, and which were likely distorted by scope mismatch or data limitations. The final output is a reasonable range, not a precise point estimate.

Illustrative Sensitivity Matrix

The matrix demonstrates how the estimate changes under different concentration assumptions, and how broader aggregate reference points are used to test plausibility. With representative-company output held constant, a lower assumed concentration produces a larger extrapolated industry-size estimate. All values are illustrative indices. The base scenario is indexed to 100.

Illustrative Sensitivity Matrix
Concentration assumption Indexed industry-size estimate SCENARIO ROLE
Illustrative low concentration 128 Illustrative variant
Illustrative moderate-low 112 Illustrative variant
Illustrative base assumption 100 Illustrative base
Illustrative moderate-high 93 Illustrative variant
Illustrative high concentration 85 Illustrative variant

Calibration reference points

  • Official statistics for larger industrial enterprises
  • Public industry association data
  • Aggregate reference information available during the project

The estimate is not a single number but a reasonable range. Concentration assumptions are the primary driver of range width; calibration against external reference points establishes plausibility boundaries.

Illustrative indices only. All values are fictional. No real concentration ratios, company outputs or final estimate values are displayed.

Key Analytical Judgements

Method selection follows industry maturity and data conditions

The choice between segment-level estimation and representative-company extrapolation depends on industry maturity, company concentration and the completeness of available statistics. In this case, the early-stage, concentrated structure of the local industry made representative-company extrapolation the more appropriate approach.

Industry labels alone are insufficient for company inclusion

A company's self-declared industry classification may not reflect its actual technology and business activity. Screening required examining processes, products and revenue sources to distinguish genuine synthetic-biology activity from tangential connections.

Company and macro evidence must be calibrated, not mechanically added

Company-level estimates and broader statistical data serve different purposes at different aggregation levels. Calibration respects each source's distinct role: company estimates as the primary signal, macro data as a plausibility check.

Planned capacity cannot be treated as current output

Publicly announced project capacity often differs materially from current production. Mixing them produces estimates that are simultaneously too high and unverifiable.

Uncertainty should be expressed as a reasonable range

When multiple key inputs carry material uncertainty, a single-point estimate is less decision-useful than a range. The final output presents a reasonable range with a neutral estimate, making uncertainty explicit and avoiding false precision.

Project Output

Produced a reasonable range and neutral estimate as a quantitative reference for subsequent planning analysis and parameter discussion. The project also produced a complete industry-size estimation framework, a structured company sample, concentration-sensitivity documentation and a multi-source calibration record.

Publication Note

This case presents only a reconstructed and anonymised analytical framework. Company identities, real operating data, model parameters, estimation results, original report pages and non-public aggregate information are not disclosed. The wider planning report was a team deliverable; this case covers only the independently completed industry-sizing component.