Horizon 3 Innovation Scorecard (Excel Template)
As an Amazon Associate I earn from qualifying purchases. Product links on this page are affiliate links — they cost you nothing extra.
⏱ 23 min read
How to Measure Pre-Revenue Horizon 3 Progress Without ROI
Pre-revenue Horizon 3 innovation accounting measures project progress through evidence velocity: the speed and net cost at which an internal venture team validates or invalidates critical business assumptions across desirability, feasibility, and viability. Instead of forecasting discounted cash flows on non-existent markets, leaders evaluate projects using hypothesis completion ratios, test strength scores, and capital-to-learning efficiency. When immediate revenue is zero dollars, progress is measured by the systematic elimination of strategic risk per dollar spent.
Horizon 3 projects are long-term corporate bets designed to create entirely new business models or capabilities for markets that do not yet exist.
Forcing conventional discounted cash flow models onto early bets creates perverse corporate incentives. In The Lean Startup, Eric Ries observed that when finance committees demand multi-year revenue projections for unvalidated concepts, project teams simply invent optimistic five-year hockey-stick forecasts to secure budget approvals. A team pitching a brand-new platform might project $45M in year-four revenue simply to hit the corporate hurdle rate of 15% internal rate of return. Everyone in the boardroom knows the spreadsheet is fiction, but standard capital allocation processes demand the ritual. The team then focuses on building the five-year plan on schedule rather than discovering whether any customer wants the product.
This dynamic leads straight to corporate failure: a team delivers software on time and within its $1.5M budget, only to discover at launch that zero customers are willing to buy it.
To fix this, finance leaders must shift away from tracking delivery milestones—such as meeting software development deadlines or completing engineering sprints—and instead measure risk-retirement velocity across venture cycles. Traditional project management asks, "Did you ship the feature on time?" Agile Innovation Accounting asks, "Did that experiment reduce our market uncertainty, and how much cash did we burn to find out?" In David Bland and Alexander Osterwalder’s framework from Testing Business Ideas, every venture carries lethal assumptions. If an assumption proves false, the venture dies. Measuring the speed at which you retire those lethal assumptions gives you an objective metric of progress.
The structural dilemma is obvious: how does a chief financial officer defend spending $250,000 across a 90-day discovery cycle when zero dollars enter the top line?
The justification lies in real-options theory, an economic valuation method established by Stewart Myers at the MIT Sloan School of Management. In real options, funding an early-stage project does not buy future cash flows; it buys the right, but not the obligation, to make a larger investment later once uncertainty drops. A $50,000 experiment that disproves customer willingness to pay saves the firm from approving a $3M commercial rollout that would have failed. In innovation accounting, rapid invalidation is a positive return on capital because it preserves capital for viable bets. You can evaluate these tradeoffs alongside your core portfolio using a Pivot vs Persevere Matrix: 5-Part Scorecard (With Template).
How to Calculate Evidence Velocity on Pre-Revenue Bets
-
Extract Lethal Assumptions Across Three Pillars
Map every operational belief into desirability (do customers care?), feasibility (can we build it?), and viability (can we make money sustainably?). Rank them by lethality: if this specific belief is false, does the venture collapse immediately? -
Assign Minimum Test Strength Scores
Score every planned experiment from 1 to 5 based on behavioral rigor. A customer survey or interview scores a 1 because spoken opinion is weak evidence; a live currency deposit, letter of intent, or prepaid pre-order scores a 5 because customer action requires skin in the game. -
Establish Capital-to-Learning Efficiency
Divide the cash burned in the cycle by the number of validated or invalidated high-risk hypotheses. A 6-week cycle spending $30,000 that resolves three critical assumptions yields a capital-to-learning efficiency of $10,000 per resolved risk factor. -
Audit the Hypothesis Completion Ratio
Track the percentage of critical assumptions tested against the scheduled sprint plan. Calculate the ratio by dividing resolved hypotheses by total identified critical hypotheses; healthy discovery teams maintain a ratio above 70% per 60-day review cycle. -
Release Staged Funding Based on Risk Retirement
Release the next tranche of capital only when the team clears its required evidence threshold. If the evidence velocity stalls or test scores remain below 3, freeze subsequent capital allocation or shutter the initiative.
To balance your overall pipeline against lower-risk operational bets, link these Horizon 3 thresholds to a broader framework to Manage Innovation Budgets: 70-20-10 (Excel Template).
The real hurdle is standardizing these velocity scores so non-technical finance committees can compare completely different bets side by side on one sheet. Up next, you will see the exact formulas used to calculate the composite Risk-Retirement Score inside the downloadable template below.
Key Takeaways
- Traditional ROI metrics kill pre-revenue projects prematurely by demanding revenue data that does not exist yet.
- Horizon 3 innovation accounting tracks risk reduction velocity across 3 pillars: desirability, feasibility, and viability.
- A sprint score above 70% unlocks metered funding stages, replacing annual budget allocations with tranche gates.
- Kill rates of 60% to 80% at early gates preserve capital for validated, high-conviction opportunities.
Table of Contents
- How to Measure Pre-Revenue Horizon 3 Progress Without ROI
- The Three Evidence Pillars That Replace Financial Projections
- Setting Up Metered Funding Gates and Tranche Thresholds
- Scoring Methodology Across Five Levels of Evidence Strength
- The Ready-to-Build Excel Scorecard Architecture and Formula Guide
- Sources & Further Reading
The Three Evidence Pillars That Replace Financial Projections
Horizon 3 projects explore unproven business models or nascent technologies where historical run-rates do not exist. Discounted cash flow models and five-year net present value forecasts fail for these early bets because financial projections reward teams for inventing fictional revenue numbers in a spreadsheet.
Innovation accounting is an evidence-based management system that tracks a team’s progress in testing assumptions and reducing uncertainty rather than measuring sales revenue or delivery milestones. Instead of fictional financial returns, teams manage speculative initiatives using the three evidence pillars established by design firm IDEO: desirability, feasibility, and viability.
Recommended gear
Innovation Accounting: A Practical Guide For Measuring Your Innovation Ecosystem's Performance
Innovate confidently with Innovation Accounting: A Practical Guide for Measuring Your Innovation Ecosystem's Performance, ensuring your company's growth through robust innovation metrics.
Affiliate link
1. Desirability Metrics: Quantifying Skin in the Game
A survey asking "Would you buy this?" produces polite lies, not validation. True desirability measures skin-in-the-game commitment, which occurs when a prospect sacrifices scarce resources like non-refundable cash, proprietary data, or time with their chief executive to test your solution. In Alberto Savoia’s pretotyping methodology published in The Right It, customer pull is measured through irreversible commitments rather than verbal interest.
Track desirability across four distinct commitment tiers:
- Attention Investment: A prospect gives you a business email address or completes a 15-minute qualification interview.
- Time and Access Investment: A target stakeholder introduces your team to their procurement lead or signs a mutual non-disclosure agreement to share internal operational logs.
- Data and Workflow Investment: A prospective enterprise customer provisions access to an isolated staging environment or exports a 10,000-row sample dataset for workflow testing.
- Financial Investment: A buyer signs a formal non-binding Letter of Intent (LOI) containing specific volume requirements, pays a $2,500 deposit for a prototype run, or signs a pilot contract.
When tracking early user traction, integrate your data into a Seed-Stage Innovation Scorecard (Spreadsheet Template) to map behavioral proof against project milestones.
2. Feasibility Scoring: De-risking Beyond Code
Feasibility proves that an engineering, regulatory, or operational architecture can function at minimum production thresholds. In Clayton Christensen’s research on tech commercialization at Harvard Business School, technical breakthroughs frequently stall because internal delivery teams confuse laboratory success with industrial deployment. Feasibility scoring separates this work into three categories:
- Core Technical Feasibility: The core mechanism achieves the baseline physics, throughput, or latency requirement. A corporate data platform, for example, must prove it can process 50,000 queries per second with sub-5-millisecond latency.
- Regulatory Clearance: The solution satisfies statutory and compliance guardrails. For regulated industries, teams must show clearance paths such as FDA 510(k) premarket notifications or SOC 2 Type II controls. Compare your metrics against a Healthcare Innovation Scorecard: 10 Metrics (Template) when measuring clinical validation milestones.
- Operational Supply Chain Readiness: Third-party partners can supply components, run compute, or assemble physical hardware at scale. The metric here is production lead time under a fixed bill of materials, not laboratory availability.
Score each feasibility milestone on a binary gate system: 0 (unproven hypothesis), 0.5 (demonstrated in an isolated lab environment), or 1.0 (validated in production conditions with real data).
3. Viability Indicators: Unit Economics Under Uncertainty
Viability evaluates whether your project can sustain itself financially if the technology works and customers buy it. Do not forecast aggregate five-year revenues. Instead, model the baseline unit economics and margin requirements using early operational data.
In Alexander Osterwalder’s Testing Business Ideas, viability testing evaluates whether the business captures sufficient economic value to cover its structural delivery costs. Focus on three core indicators:
- Gross Margin Floor: The estimated price minus direct delivery costs must clear your business unit’s minimum gross margin target (typically 70% for enterprise software or 35% for physical hardware).
- Ecosystem Take-Rates: The revenue share claimed by distribution partners, value-added resellers, or app marketplaces. If an app ecosystem takes 30% of gross transaction value, your direct margins must absorb that toll without exceeding customer willingness-to-pay thresholds.
- Addressable Market Density: The absolute number of potential accounts that match your exact early-adopter profile. A theoretical $10B total addressable market is useless if only 40 domestic enterprise companies match your specific technical prerequisites.
When unit economics shift during early customer validation, teams should deploy a Pivot vs Persevere Matrix: 5-Part Scorecard (With Template) to evaluate whether product changes compromise their baseline margins.
Weighting the Scorecard: Tech-First vs. Market-First
Do not apply identical pillar weightings across every portfolio initiative. Projects face different primary points of failure. Weight the scorecard to penalize the venture’s primary source of risk, as outlined in frameworks for Agile Innovation Accounting.
A tech-first venture carries fundamental scientific or engineering uncertainty (for example, a solid-state battery chemistry or a new machine learning model). If the science works, the commercial market exists. Feasibility must dominate the scorecard.
A market-first venture uses proven, off-the-shelf technology to serve an unverified customer behavior or unproven business model (such as an on-demand freight brokerage platform). The engineering works on day one; the risk is that nobody cares or will pay for it. Desirability must dominate the scorecard.
| Evaluation Dimension | Tech-First Allocation | Market-First Allocation | Example Milestone Trigger |
|---|---|---|---|
| Desirability Weight | 20% | 50% | Signed Letter of Intent (LOI) with contract terms |
| Feasibility Weight | 60% | 20% | Functional proof of concept meeting latency gates |
| Viability Weight | 20% | 30% | Unit contribution margin exceeding 40% floor |
| Review Cadence | Monthly technical gates | Bi-weekly sprint reviews | Evidence board review of customer discovery |
| Primary Failure Metric | Performance below specs | Customer drop-off rate | High churn or non-renewal of paid pilot tests |
Once you set these pillar weights in your scorecard, you can map the resulting evidence scores directly into stage-gate funding formulas that release capital in tranches.
Setting Up Metered Funding Gates and Tranche Thresholds
Metered funding releases capital in small, predetermined installments tied to validated learning milestones rather than calendar quarters or traditional delivery schedules.
Horizon 3 projects—bets aimed at creating entirely new businesses or markets—fail when corporate leaders fund them like established business units. Traditional annual budgeting hands a project team $1.2M upfront, checks back twelve months later, and finds a polished prototype that nobody wants. To protect capital, finance committees must replace bulk allocations with three disciplined tranches: Discovery, Validation, and Incubation.
[Discovery Tranche]
$25k - $50k | 6-8 Weeks
Proof of Customer Pain
|
v
[Validation Tranche]
$100k - $250k | 12-16 Weeks
Proof of Solution & Demand
|
v
[Incubation Tranche]
$500k - $1.5M | 6-9 Months
Proof of Scalable Engine
The Discovery tranche runs on $25,000 to $50,000 over six to eight weeks. The team must prove that the targeted customer problem exists, causes real pain, and commands an urgent desire for a solution. Teams that fail to confirm clear demand drop out immediately, having spent less than the price of a corporate offsite.
Teams that clear Discovery enter Validation, backed by an allocation between $100,000 and $250,000 across three to four months. Here, the squad tests early solution concepts, price sensitivity, and initial willingness to pay using lightweight smoke tests or interactive mockups. A team working on digital concepts might evaluate early user interest with Wireframing for UI/UX Innovation before writing production code.
Incubation represents the final pre-revenue gate, releasing $500,000 to $1.5M over six to nine months. At this level, the venture must build a functioning delivery engine, run live pilots, and identify clear unit economics. To implement this three-tier flow without creating custom accounting rules for every project, teams use Agile Innovation Accounting to structure governance around verified evidence rather than projected revenue.
A hurdle rate is an objective benchmark of retired uncertainty that an innovation team must achieve to qualify for their next round of funding.
Never evaluate Horizon 3 gates using return on investment, net present value, or projected five-year market capture. In The Corporate Startup, authors Tendayi Viki, Danoma Toma, and Esther Gons establish that early innovation hurdles must measure evidence gathered per dollar spent. At the Validation gate, require teams to retire 70% of their critical, high-impact assumptions before requesting follow-on funds.
In practice, a team lists their core assumptions across customer desirability, technical viability, and financial viability. If they identify ten lethal assumptions, they must systematically falsify or validate seven of them with auditable customer actions—such as executed letters of intent or verified waitlist deposits—before unlocking the next tranche. Teams can track these dynamic risk retirements directly inside a Seed-Stage Innovation Scorecard (Spreadsheet Template).
⚠️ Anti-Pattern: The Zombie Project Life Support Trap
What it looks like: Committees approve modest 10% budget extensions every quarter because a venture team hit software build deadlines, even though no customers have signed up.
Why it’s tempting: Killing a project forces leadership to write off sunk capital and admit a strategic bet failed.
What it costs: Chronic capital drain that starves breakthrough concepts while anchoring top engineering talent to commercially dead initiatives.
Do instead: Enforce strict stage gates where lack of validated customer signal triggers an immediate, blameless project shutdown regardless of development progress.
To produce reliable returns, enterprise venture portfolios must enforce an intentional 60% to 80% project kill rate at the Discovery stage. Venture capital performance data published by Correlation Ventures shows that 65% of professional early-stage startup investments fail to return their invested capital. Established enterprises cannot bypass those underlying venture odds.
When an organization lets 90% of early concepts slide through into high-cost development, it creates portfolio congestion. Ten speculative teams end up consuming millions of dollars building unverified software, diluting the resources needed to scale the rare concept that shows genuine traction. By pruning 60% to 80% of concepts during Discovery, leadership reclaims capital and protects the organization’s financial return, aligning with best practices in Manage Innovation Budgets: 70-20-10 (Excel Template).
Clear shutdown protocols are formal operating rules that define how an organization winds down unvalidated projects while protecting the teams that ran them.
If killing an initiative damages a team member’s career prospects or performance review, employees will hide invalidating data to protect their standing. Management literature from the Harvard Business Review repeatedly demonstrates that psychological safety is the bedrock of corporate risk management.
Reward teams for finding negative evidence rapidly. If a team exhausts a $30,000 Discovery tranche in three weeks and proves customers will not pay for the concept, reward that team with priority placement on the next approved venture. They saved the enterprise $500,000 in downstream engineering costs.
To formalize these pivotal decisions, introduce a systematic Pivot vs Persevere Matrix: 5-Part Scorecard (With Template) that turns shutdown reviews into routine data checks rather than personal interrogations.
Once you set up these funding gates, you must determine how to calculate the quantitative risk-retirement velocity scores inside your spreadsheet model.
Scoring Methodology Across Five Levels of Evidence Strength
The evidence strength score in an innovation scorecard discounts unvalidated assumptions by multiplying raw test results against an objective reliability weight between 0.1 and 1.0.
Horizon 3 projects—exploratory initiatives aimed at creating entirely new markets or capabilities—cannot rely on traditional revenue or return-on-investment metrics during early discovery.
An evidence weight multiplier is a numerical factor that reduces the value of weak data to prevent early teams from treating customer polite interest as commercial proof.
Without this discount, teams report false progress. In their book Testing Business Ideas, David J. Bland and Alexander Osterwalder point out that teams routinely mistake customer conversation for customer validation. A founder hears ten potential buyers say an idea sounds valuable, marks the hypothesis as proven, and requests another budget tranche. Evidence weighting corrects for that cognitive trap.
Levels 1 and 2: Low-Fidelity Signals (0.1 to 0.3 Multiplier)
Level 1 and Level 2 evidence captures what people say, not what they do.
Level 1 evidence consists of internal opinions, secondary industry reports, and executive intuition. Assign these data points an evidence weight of 0.1. A 40-page market forecast from Gartner or Forrester shows that a macro opportunity exists, but it provides zero evidence that anyone wants your specific solution.
Level 2 evidence covers self-reported interest gathered through surveys, focus groups, and unstructured discovery interviews. Weight this data at 0.2 to 0.3. When you run customer interviews, interviewees often provide encouraging feedback because politeness costs them nothing. Rob Fitzpatrick documents this dynamic in The Mom Test, noting that asking people whether they would buy a hypothetical product yields consistently misleading positive responses. If 40 out of 50 survey respondents state they would pay $100 per month for your proposed diagnostic tool, a 0.2 multiplier values that finding at only 8 units of validated evidence.
Level 3: Observable Action (0.5 Multiplier)
Level 3 evidence measures observed behavior and passive commitment.
Assign an evidence weight of 0.5 to actions where users spend time, surrender contact information, or alter their workflow to evaluate an early concept. Examples include landing page click-through rates, opt-in email conversions from cold digital campaigns, and time spent reviewing technical white papers.
Behavioral data strips out conversation bias. If 200 prospective enterprise users visit a gated feature page and 30 submit their corporate email addresses to join a private beta, you have a 15% conversion rate backed by real attention. You can pair this stage with an established VOC Translation Matrix (With 5-Step Template) to translate those behavioral cues into verifiable operational requirements. However, because no legal or financial assets changed hands, the evidence remains halfway to confirmation.
Levels 4 and 5: Skin in the Game (0.8 to 1.0 Multiplier)
Levels 4 and 5 represent definitive validation through financial commitment, legal exposure, or deep operational integration.
Level 4 evidence carries a 0.8 multiplier and requires non-trivial skin in the game. This tier includes signed non-binding letters of intent (LOIs), refundable pilot deposits, and access granted to internal enterprise databases for sandbox testing. When a enterprise buyer executes an LOI, they spend internal political capital and legal review hours. That investment indicates real organizational urgency.
Level 5 evidence earns a full 1.0 multiplier. This is conclusive validation: non-refundable deposits, binding pilot contracts, or functioning technical prototypes deployed inside the customer’s live operational architecture. If an industrial logistics lead installs your sensor prototype on five active delivery vehicles for a 30-day trial, feasibility and desirability risks drop sharply. Integrating these high-conviction metrics directly into your Seed-Stage Innovation Scorecard (Spreadsheet Template) gives governance committees defensible data for follow-on funding decisions.
Evidence Progression
|
+-- L1: Desk Research (0.1)
| - Opinions, analyst reports
|
+-- L2: Stated Intent (0.3)
| - Customer interviews, surveys
|
+-- L3: Observed Action (0.5)
| - Landing page clicks, beta signups
|
+-- L4: Skin in the Game (0.8)
| - LOIs, refundable deposits
|
v
+-- L5: Committed Usage (1.0)
- Live pilot deployment, paid upfront
The Composite Sprint Score Rubric
To calculate a composite sprint score in your scorecard, score both the criticality of the underlying hypothesis and the reliability of the test method.
Hypothesis Criticalness (\(H\)) is scored on a 1-to-5 scale:
- 1: Minor interface or positioning detail.
- 2: Non-essential feature preference.
- 3: Usability or channel assumption.
- 4: Core unit economic or pricing assumption.
- 5: Fundamental problem or technical feasibility assumption (if false, the venture dies).
Test Reliability (\(R\)) matches the five evidence levels (0.1, 0.3, 0.5, 0.8, 1.0).
\(\text{Composite Score} = \text{Hypothesis Criticalness } (H) \times \text{Evidence Multiplier } (R)\)
| Sprint Hypothesis | Criticalness (\(H\): 1-5) | Test Executed | Evidence Level | Multiplier (\(R\)) | Sprint Score (\(H \times R\)) | Status |
|---|---|---|---|---|---|---|
| Cloud brokers will cut latency by 40% | 5 | Architect desk estimate | Level 1 | 0.1 | 0.50 | High Risk |
| Plant managers want continuous monitoring | 4 | 20 discovery calls | Level 2 | 0.3 | 1.20 | Unconfirmed |
| Safety teams will review weekly dashboard | 2 | Click-track on mockup | Level 3 | 0.5 | 1.00 | Moderate |
| Fleet owners will pay $200/mo/vehicle | 5 | Signed pilot LOI with deposit | Level 4 | 0.8 | 4.00 | Validated |
| Telemetry integrates with legacy ERP | 5 | On-premise API live test | Level 5 | 1.0 | 5.00 | Confirmed |
A team that runs 10 interviews on a critical hypothesis scores only \(5 \times 0.3 = 1.5\) out of 5.0. A team that runs an on-site technical sandbox test scores \(5 \times 1.0 = 5.0\). This arithmetic stops vanity testing immediately. Using this mechanism inside your Agile Innovation Accounting system helps teams measure real certainty instead of sprint velocity. Teams can also apply the Pivot vs Persevere Matrix: 5-Part Scorecard (With Template) when composite scores remain under 2.0 across two consecutive validation cycles.
You decide: allocating sprint budget under conflicting evidence
Imagine you lead an internal venture team building an automated cold-chain monitoring system for pharmaceutical distributors. You have a two-week sprint and a $10,000 discovery budget remaining before the quarterly stage-gate review.
Decision point: Choose which validation track receives the sprint budget.
Option A — Commission an expanded survey of 200 regional logistics directors
You secure survey responses from 214 logistics leads within 8 days. 78% indicate that temperature deviations cause severe quarterly inventory write-downs, but no respondents provide direct contact information for follow-up testing.
Present the findings to the stage-gate committee
The committee records a high-volume sample that remains capped at Level 2 evidence (0.3 multiplier). The resulting sprint score fails to cross the threshold required to unlock seed capital, leaving the critical operational risks unanswered.
Option B — Offer three regional distributors a manual, on-site temperature audit
Two facilities decline due to compliance rules. One facility manager agrees to a supervised 48-hour pilot, granting your engineer temporary site access to monitor three active storage units.
Present the pilot audit data to the stage-gate committee
The real deployment yields Level 4 evidence (0.8 multiplier) on data accessibility. Even with a sample size of one facility, the direct physical commitment provides verifiable proof of technical integration and customer access.
Structuring evidence this way changes how your stage-gate committees review projects during funding reviews.
The next step is wiring these multipliers directly into the Excel template’s automated aggregation tabs so your team can calculate pipeline velocity in real time.
The Ready-to-Build Excel Scorecard Architecture and Formula Guide
A reliable innovation accounting scorecard for Horizon 3 ventures requires exactly four linked spreadsheet tabs: an Assumption Log, an Evidence Scoring Matrix, a Portfolio Overview Dashboard, and a Gate Review Summary. Horizon 3 ventures explore transformational opportunities that operate without historical sales data or defined target markets. Standard financial metrics like net present value will kill these early concepts prematurely. In his book The Lean Startup, Eric Ries introduced agile innovation accounting to substitute vanity revenue projections with empirical milestones of validated learning.
You can build this entire calculation engine in Microsoft Excel or Google Sheets in 15 minutes. It replaces guesswork with numeric risk calculations, giving corporate venture boards an auditable paper trail before they release tranche funding.
The Four-Tab Architecture
Open a new workbook and set up four tabs. Configure each sheet with the specific column headers and data types detailed below.
+------------------------------------------+
| WORKBOOK ARCHITECTURE |
+------------------------------------------+
|
v
+------------------------------------+
| 1. ASSUMPTION LOG |
| Raw hypotheses & risk ratings |
+------------------------------------+
|
v
+------------------------------------+
| 2. EVIDENCE SCORING MATRIX |
| Experiment rigor & signals |
+------------------------------------+
|
v
+------------------------------------+
| 3. GATE REVIEW SUMMARY |
| Pass / Pivot / Kill triggers |
+------------------------------------+
|
v
+------------------------------------+
| 4. PORTFOLIO DASHBOARD |
| Multi-venture governance roll-up|
+------------------------------------+
Tab 1: Assumption Log
This tab decomposes business plans into falsifiable hypotheses across desirability, feasibility, and viability.
- Column A (Assumption_ID): Format as plain text (e.g.,
A-001,A-002). - Column B (Category): Data validation drop-down menu with three values:
Desirability,Feasibility,Viability. - Column C (Statement): Text field stating the hypothesis clearly (e.g., "Field technicians will adopt voice data-entry within 14 days").
- Column D (Impact): Whole integer from 1 (marginal impact) to 5 (critical business model failure if wrong).
- Column E (Confidence): Whole integer from 1 (no data, complete guess) to 5 (confirmed by multi-cohort quantitative behavior).
- Column F (Weighted_Risk): Calculated column using
=D2*(6-E2). This weights high-impact assumptions that have low confidence.
Tab 2: Evidence Scoring Matrix
This sheet records the experimental tests designed to validate assumptions from Tab 1. In Testing Business Ideas, David J. Bland and Alexander Osterwalder demonstrate that an evidence score must reflect the strength of user commitment rather than conversational intent.
- Column A (Test_ID): Text format (e.g.,
T-001). - Column B (Assumption_ID): Data validation linking to Tab 1, Column A.
- Column C (Evidence_Type): Drop-down list:
Discovery Interview(Weight: 10),Click-Through Test(Weight: 30),Concierge MVP(Weight: 60),Financial Deposit(Weight: 100). - Column D (Sample_Size): Numeric count (e.g.,
45). - Column E (Success_Threshold): Percentage target (e.g.,
20%). - Column F (Observed_Result): Percentage measured (e.g.,
8%). - Column G (Evidence_Score): Calculated column:
=(F2/E2)*VLOOKUP(C2, Admin_Weights, 2, FALSE).
Tab 3: Gate Review Summary
This sheet synthesizes the data for governance decisions, showing whether to release the next tranche of capital. If a project runs into structural barriers, pair this workflow with a pivot vs persevere matrix to evaluate reallocation options.
- Column A (Metric): Static rows:
Aggregated Assumption Risk,Evidence Velocity,Capital Consumed,Gate Recommendation. - Column B (Current_Value): Formulas aggregating Tab 1 and Tab 2.
- Column C (Threshold_Target): Governance rules agreed on by the innovation committee.
Tab 4: Portfolio Overview Dashboard
This view aggregates scores across multiple pre-revenue Horizon 3 ventures to help leadership prioritize R&D projects across business units.
- Columns A–E:
Project Name,Current Stage,Weighted Risk Score,Evidence Velocity (Tests/Month),Gate Status Indicator.
Core Formulas and Automation Engine
To calculate overall venture exposure across assumptions, write the weighted assumption risk formula into cell B1 of the Gate Review Summary tab:
=SUMPRODUCT(Assumption_Log!D2:D50, 6-Assumption_Log!E2:E50)/COUNT(Assumption_Log!D2:D50)
Where:
Assumption_Log!D2:D50is the named rangeImpact_Range.6-Assumption_Log!E2:E50transforms the 1-to-5 confidence scale so that lower confidence yields higher risk weight.COUNT(Assumption_Log!D2:D50)is the named rangeTotal_Risks.
This calculation scores portfolio risk on a clean 1.0 to 25.0 scale.
Next, track testing speed. In innovation accounting, rapid learning cycles drive capital efficiency. Track your team’s evidence velocity in cell B2 using:
=COUNTIFS(Evidence_Matrix!G2:G50, ">0", Evidence_Matrix!H2:H50, ">="&EDATE(TODAY(), -1))
This counts the tests producing verified data completed within the past 30 days.
Conditional Formatting Rules for Gate Reviews
Automate governance flags on the Portfolio Overview Dashboard tab (Column E) to keep reviews objective. Link Column E to this conditional logic formula:
=IF(AND(B1<=6.0, B2>=3), "Ready to Pitch", IF(B1>12.0, "Candidate for Deprioritization", "More Evidence Needed"))
Apply three automated formatting rules to the status cells:
- Green (Ready to Pitch): Set formula rule to
=E2="Ready to Pitch". Fill hex#D4EDDA, font color hex#155724. The venture has reduced critical risks below 6.0 and completed at least 3 valid experiments in the last 30 days. It is eligible for the next round of capital. - Amber (More Evidence Needed): Set formula rule to
=E2="More Evidence Needed". Fill hex#FFF3CD, font color hex#856404. Residual risk remains elevated (between 6.1 and 12.0). The venture keeps its current runway but receives no budget increases until confidence scores rise. - Red (Candidate for Deprioritization): Set formula rule to
=E2="Candidate for Deprioritization". Fill hex#F8D7DA, font color hex#721C24. Total risk sits above 12.0 with weak evidence velocity. Flag the project for structured decommissioning or return to basic research. If you run a specialized setup, check our seed-stage innovation scorecard for stage-specific alternatives.
📋 Pocket Cheat Sheet: H3 Innovation Scorecard
Reference rules for H3 venture milestone tracking.
CORE SCORING FORMULAS Risk Score: =D2*(6-E2) Portfolio Risk: =SUMPRODUCT(D2:D50, 6-E2:E50)/COUNT(D2:D50) Velocity: =COUNTIFS(G2:G50, ">0", H2:H50, ">="&EDATE(TODAY(),-1)) CONFIDENCE TIERS (Col E) 1: Zero data / opinion 2: Anecdotal interviews (n<10) 3: Broad discovery surveys (n>100) 4: Behavioral currency (clicks, signups) 5: Financial commitment (LOI, deposit) GATE THRESHOLDS (Col E Flag) GREEN: Risk <= 6.0 AND Velocity >= 3/mo AMBER: Risk 6.1-12.0 (Runway preserved) RED: Risk > 12.0 OR Velocity = 0 (De-fund)
Copy this into your notes app.
According to corporate venture performance data tracked by McKinsey & Company, over 70% of early-stage corporate ventures fail to scale because teams scale operational spending before confirming market validation. By wiring these specific formulas into your tracking sheet, you replace executive intuition with transparent data.
Take your most uncertain project today, list its top five unproven assumptions into Tab 1, and run the calculation engine. If the overall risk score lands in the red, halt engineering spend immediately and reallocate those resources to empirical tests.
Sources & Further Reading
Rigorous innovation accounting rests on published portfolio theory, venture capital capital-allocation mechanics, and formal validation frameworks rather than standard enterprise discounted cash flow formulas.
Innovation accounting is a structured framework of non-financial leading indicators designed to measure an early-stage venture’s progress in reducing business model risk before commercial revenue exists. Standard accounting practices fail here because they track historic cash flows, which forces experimental teams to present fabricated revenue forecasts to survive executive stage-gates.
To govern bets that take years to mature, organizations rely on the Three Horizons model developed by Mehrdad Baghai, Stephen Coley, and David White in The Alchemy of Growth (2000). Their research across 30 global companies established that Horizon 3 initiatives require exploratory milestone funding rather than Horizon 1 operating margins. When you judge a pre-revenue discovery team by current operating profit, you inadvertently incentivize them to halt experimental discovery in favor of safe, low-margin feature extensions.
In The Lean Startup (2011), Eric Ries defined the foundational baseline for this methodology by demonstrating that early validation velocity beats premature scaling. Ries documented how teams that track validated learning loops through actionable metrics avoid spending 100% of their initial capital on products nobody wants. Following this logic, Dan Toma, Esther Gons, and Cristian Mitreanu formalized corporate governance metrics in Innovation Accounting (2021), proving that teams should track evidence strength and assumption burn-down rates across 12-week tranches instead of quarterly revenue targets.
Institutions that govern large portfolios treat Horizon 3 exploration like early-stage equity investing rather than budget consumption. Research published by the Harvard Business Review illustrates that venture-style staging protects capital by releasing funds incrementally against clear risk reduction. For example, Bansi Nagji and Geoff Tuff demonstrated in their 2012 Harvard Business Review study, "Managing Your Innovation Portfolio," that top-performing enterprises allocate 10% of their innovation resources to transformational bets, yet those investments generate roughly 70% of long-term cumulative returns.
- Baghai, M., Coley, S., and White, D. (2000), The Alchemy of Growth, Texere — Establishes the Three Horizons framework separating mature operational management from exploratory, pre-revenue initiatives.
- Ries, E. (2011), The Lean Startup, Crown Business — Introduces innovation accounting, defining actionable metrics and systematic hypothesis validation over vanity metrics.
- Toma, D., Gons, E., and Mitreanu, C. (2021), Innovation Accounting: A Practical Guide for Measuring Your Innovation Ecosystem, BIS Publishers — Outlines the formal indicators, scorecard rubrics, and governance tiers needed to measure corporate exploration.
- Nagji, B. and Tuff, G. (2012), "Managing Your Innovation Portfolio", Harvard Business Review — Quantifies the 70-20-10 resource allocation model and demonstrates how Horizon 3 exploration drives outsized financial returns.
- Blank, S. (2006), The Four Steps to the Epiphany, K&S Ranch — Details the Customer Discovery process that generates the evidence criteria used across milestone scorecards.
Featured image by Nika Benedictova on Pexels