This is a prototype, and the intelligence is emulated.
No call to Jev or TypeSafe AI was made to produce this page. Every number below
came from an offline stand-in that speaks Jev's published contract
(POST /v1/systemone, the Noul / Choice / Score primitives) and returns
probability distributions that are calibrated by construction against hidden labels
in a synthetic corpus. Kestrel Data Systems is invented, as are all twelve deals,
ten accounts, six initiatives and eight cost lines. Swapping the emulator for the
real model is one flag: --backend live. Nothing here is an endorsement
by, or affiliated with, TypeSafe AI.
Capital allocation · FY27 planning · September 2026
Where should Kestrel put $14M?
Mid-market B2B SaaS. Data-quality and pipeline observability for logistics, insurance and financial-services operators. Six initiatives are asking for $20.0M against a $14.0M
envelope, and the board wants $62.0M of exit ARR. The arithmetic is
easy. The inputs are the problem: every driver in the model turns on a judgment
somebody has to make by reading a call note, a support queue, or a contract.
Current ARR
$48.2M
$21.0M up for renewal
Envelope
$14.0M
against $20.0M requested
Board target
$62.0M
exit ARR, end of FY27
Judgments
170
in 36 requests, $0.0009
The idea
A distribution is a better forecast input than an assumption
A driver-based model is arithmetic until it reaches an assumption. Normally those
assumptions are set in a planning meeting by whoever argues hardest, recorded as a
single number, and then wrapped in a made-up confidence interval at the end.
A System One model answers the same questions against the actual evidence and returns
a probability distribution for each one. Those distributions are not a
summary of the uncertainty — they are the uncertainty, in the only form a
Monte Carlo can integrate over. The width of the band below is inherited from how
sure the model was, judgment by judgment.
One request, every question, one shared state
Output tokens are free and questions are evaluated in parallel against one state,
so there is no reason to ask them one at a time. Each deal below is a single call
carrying five questions — including speculative ones that only matter on some
branches.
Request
{
"state": "<the deal record above>",
"model": "jev-latest",
"questions": {
"closes_in_horizon": {
"type": "noul",
"instructions": "Will this opportunity be signed within the next two quarters?"
},
"primary_blocker": {
"type": "choice",
"instructions": "What single thing most stands between this deal and signature?",
"criteria": {
"budget_freeze": "Money is unavailable or suspended",
"security_review": "An open security, compliance or pen-test gate",
"...": "..."
}
},
"...": "+ 3 more, same call"
}
}
A policy is nothing but a set of weights over the six questions asked about each
initiative, plus a view on what an unmeasurable strategic case is worth. The weights
live in forecast/policies.py where a CFO can argue with them. Each row is
20,000 Monte Carlo trials drawing from the judgment distributions.
Exit ARR at the end of FY27
Median marked; bar spans the 10th to 90th percentile.
recommendedalternativedashed rule — board target
Probability of hitting the $62.0M target
All six, side by side
Policy & what it funds
Spend
Exit P50
P10 – P90
P(target)
FY27 cash
Evidence-weighted recommended
pricingcompliancereliabilitygtm
$12.1M
$63.0M
$57.6M – $67.6M
60%
$-5.8M
Defend the base
pricingcompliancereliabilityvpc
$11.8M
$61.6M
$56.4M – $65.8M
46%
$-6.3M
Pump growth
gtmpricingcompliancereliability
$12.1M
$63.0M
$57.6M – $67.5M
60%
$-5.8M
Unblock the pipeline
compliancepricinggtmreliability
$12.1M
$63.0M
$57.6M – $67.5M
60%
$-5.9M
Follow the CEO
copilotgtmpricingcompliance
$13.3M
$62.9M
$57.5M – $67.5M
59%
$-7.9M
Margin discipline
compliancepricing
$5.1M
$61.5M
$56.3M – $65.7M
44%
$3.7M
Evidence-weighted — projected ARR
Fund by what the evidence supports: demand, deals unblocked, speed to revenue and retention, discounted by delivery risk.
The payoff
Where a human should actually spend an hour
Calibrated confidence turns into triage. Each judgment is gated at a threshold that
scales with the money it moves — a $540k deal and a $4.2M renewal do not deserve
the same bar — and ranked by how much forecast variance it carries. Independent
Bernoullis, so each one's share is amount² · p(1-p) over the total.
6 of 22 judgments fall below their floor,
and they carry 42% of the spread in the forecast.
That is the list, in order.
Account
Judgment
At stake
p
Confidence
Variance share
Route
Ember Financial
closes_in_horizon
$3.2M
0.82
0.64
17.5%
human
Northwind Freight
renews
$4.2M
0.92
0.84
15.0%
review
Riverbend Bank
renews
$2.8M
0.86
0.73
10.8%
human
Pinnacle Retail
renews
$3.1M
0.10
0.79
10.3%
review
Juniper Payments
closes_in_horizon
$2.4M
0.86
0.72
8.0%
human
Borealis Health Partners
closes_in_horizon
$2.1M
0.89
0.77
5.1%
review
Tidewater Shipping
renews
$1.9M
0.86
0.72
5.1%
review
Glacier Mutual
closes_in_horizon
$1.8M
0.16
0.69
5.0%
review
Vantage Insurance
renews
$2.6M
0.95
0.90
3.8%
review
Lumen Grid Utilities
closes_in_horizon
$1.6M
0.12
0.76
3.2%
review
Westline Rail
renews
$1.1M
0.28
0.43
2.9%
human
Summit Logistics
renews
$2.2M
0.05
0.90
2.8%
review
The check
Does the confidence mean anything?
Everything above consumes probabilities and trusts them, so the number worth
publishing is this one. Because the corpus is synthetic, every judgment has a known
answer, and the same harness can score the emulator, a Claude-backed stand-in, or the
real model on identical questions.
Stated probability vs observed accuracy
observed, sized by sampledashed — perfect calibration
Scorecard
Judgments scored
170
Accuracy
79.4%
Brier score
0.157
Expected calibration error
0.045
Accuracy when it claims ≥ 0.8
85.6%
Accuracy when it claims < 0.6
58.3%
The gap between those last two rows is the whole product. A model that is
86% right when confident and
58% right when unsure can be gated. One that is
equally wrong at both ends cannot be, no matter how accurate it is on average.
The evidence
Four analytical junctures
Every judgment the forecast rests on, with the evidence it was made from. Nothing
prefixed with an underscore in the corpus — the hidden labels — is ever
sent as state; there is a test that walks every record to prove it.
Pipeline conversion
Which of the 12 open opportunities are real, and what is each one actually worth?
Atlas Freight SystemsACV $1.4M human needed
closes_in_horizon0.93
deal_realitycommitted · 0.74
primary_blockernone · 0.78
expansion_headroom3.10 / 4 · 0.65
needs_roadmap_commit0.09
Champion walked us through their FY27 budget line item; the money is allocated and named. Technical validation is a formality at this point, their team already rebuilt two internal dashboards on our API during the trial. Legal has our MSA in redlines with o...
Borealis Health PartnersACV $2.1M human needed
closes_in_horizon0.89
deal_realitycommitted · 0.37
primary_blockerintegration_gap · 0.31
expansion_headroom3.76 / 4 · 0.64
needs_roadmap_commit0.83
Commercials are agreed and the CDO has budget authority. Their security office will not sign until we support field-level encryption on PHI columns and can execute a BAA. We do not ship either today. Their counsel confirmed that is the only remaining gate.
Cinder Retail GroupACV $680k human needed
closes_in_horizon0.15
deal_realityghost · 0.52
primary_blockerchampion_departed · 0.79
expansion_headroom0.95 / 4 · 0.85
needs_roadmap_commit0.08
Our champion left the company. The replacement VP took one intro call, said the analytics roadmap is under review, and has not responded to four follow-ups since. No one else at the account has logged into the trial in six weeks.
Dovetail LogisticsACV $920k human needed
closes_in_horizon0.77
deal_realitystalled · 0.51
primary_blockercompetitor · 0.27
expansion_headroom2.68 / 4 · 0.44
needs_roadmap_commit0.19
Formal bake-off against Northlight. Their procurement lead volunteered that Northlight came in 30% below our list and bundled the capability into an existing contract. Our champion likes the product more but was candid that the decision is being made on tot...
Ember FinancialACV $3.2M human needed
closes_in_horizon0.82
deal_realitycommitted · 0.75
primary_blockernone · 0.32
expansion_headroom3.85 / 4 · 0.72
needs_roadmap_commit0.12
Largest deal in the pipeline. CFO has signed off on the number and it is in their approved capital plan. Procurement is slow and process-bound but has produced a signature calendar with a date inside the quarter. Redlines are down to a mutual limitation-of-...
Fathom MediaACV $410k review
closes_in_horizon0.90
deal_realitystalled · 0.64
primary_blockerbudget_freeze · 0.58
expansion_headroom1.06 / 4 · 0.57
needs_roadmap_commit0.04
Genuine technical interest from one analyst who found us through a conference talk. When we asked about budget he said the company paused all new software spend through the end of the calendar year and he was researching for a future cycle. No exec has join...
Glacier MutualACV $1.8M human needed
closes_in_horizon0.16
deal_realitystalled · 0.37
primary_blockerintegration_gap · 0.31
expansion_headroom2.98 / 4 · 0.35
needs_roadmap_commit0.86
Regulated insurer. Their architecture review board will not approve any multi-tenant SaaS for policyholder data; they require deployment inside their own VPC. Our champion has pushed twice and been told the policy is not negotiable. We have no customer-VPC ...
Harbor & Main FreightACV $540k review
closes_in_horizon0.92
deal_realitycommitted · 0.83
primary_blockernone · 0.93
expansion_headroom2.02 / 4 · 0.66
needs_roadmap_commit0.10
Small, clean, fast. The COO controls the spend and gave a verbal yes on the call. Standard paper, no legal review required under their threshold. Wants to start before their peak season.
Ironwood ManufacturingACV $1.2M human needed
closes_in_horizon0.08
deal_realitystalled · 0.68
primary_blockerintegration_gap · 0.73
expansion_headroom2.02 / 4 · 0.50
needs_roadmap_commit0.07
The proposal is well received and the champion still wants it. Two weeks ago their CEO announced a company-wide hiring and discretionary spend freeze after a weak quarter. Our champion said explicitly that new software purchases above $250k are suspended un...
Juniper PaymentsACV $2.4M human needed
closes_in_horizon0.86
deal_realitycommitted · 0.28
primary_blockerbudget_freeze · 0.22
expansion_headroom2.83 / 4 · 0.45
needs_roadmap_commit0.18
Payments company, so the security bar is the deal. Their pen test returned two medium findings we can remediate, plus a hard requirement for SAML SSO and SCIM provisioning which is on our roadmap but unfunded. The champion is fighting for us internally and ...
Lumen Grid UtilitiesACV $1.6M human needed
closes_in_horizon0.12
deal_realitychampion_only · 0.39
primary_blockernone · 0.45
expansion_headroom2.61 / 4 · 0.11
needs_roadmap_commit0.24
The champion is the most engaged person in our pipeline and has run three internal demos unprompted. We have never met anyone above him, he has not confirmed a budget source, and this utility historically takes four to six quarters to buy anything. Enthusia...
Meridian Air CargoACV $780k human needed
closes_in_horizon0.93
deal_realitycommitted · 0.31
primary_blockernone · 0.84
expansion_headroom1.77 / 4 · 0.61
needs_roadmap_commit0.19
Inbound from a referral by an existing customer. Moving unusually fast for discovery: they arrived with a written requirements list, a named budget, and a target go-live tied to a freight contract that starts next quarter. The referral source is one of our ...
Renewal and retention
Which of the $21.0M renewal cohort stays, and what would make it stay?
Northwind FreightARR $4.2M human needed
renews0.92
churn_drivernone · 0.78
expansion_readiness3.77 / 4 · 0.66
reliability_pain0.89 / 4 · 0.69
champion_departed0.10
Their VP opened the QBR by saying we are the only vendor renewal she has not had to argue for. Two new business units asked to be onboarded in the next cycle. They agreed to a public case study.
Pinnacle RetailARR $3.1M human needed
renews0.10
churn_driverprice · 0.35
expansion_readiness1.15 / 4 · 0.64
reliability_pain3.80 / 4 · 0.72
champion_departed0.09
The QBR was an escalation meeting. Their CTO said the phrase 'we cannot take another peak season like that one' twice. They asked for our multi-region roadmap in writing and we had nothing to send. Procurement has already been asked to price alternatives.
Quarry IndustrialARR $1.4M human needed
renews0.90
churn_driverconsolidation · 0.62
expansion_readiness0.34 / 4 · 0.56
reliability_pain1.03 / 4 · 0.52
champion_departed0.10
Their new CFO is running a vendor consolidation exercise across the whole software stack and asked every owner to justify spend. Our champion produced a usage report that made the case comfortably; the workflows we support feed their regulatory reporting an...
Riverbend BankARR $2.8M human needed
renews0.86
churn_driverchampion_departed · 0.43
expansion_readiness2.23 / 4 · 0.04
reliability_pain0.40 / 4 · 0.50
champion_departed0.81
Our champion of five years was promoted to a different division in August. Her replacement has been in seat six weeks, took the QBR, asked good questions and made no commitments. The working team is unchanged and still depends on us daily. We have no read o...
Summit LogisticsARR $2.2M human needed
renews0.05
churn_driverchampion_departed · 0.38
expansion_readiness2.15 / 4 · 0.50
reliability_pain2.03 / 4 · 0.38
champion_departed0.10
A new group CISO issued a data-residency policy requiring customer-controlled infrastructure for anything touching shipment manifests. That is most of what they run with us. They have asked three times for a customer-VPC option and been told it is not on th...
Tidewater ShippingARR $1.9M human needed
renews0.86
churn_driverreliability · 0.47
expansion_readiness1.91 / 4 · 0.40
reliability_pain3.00 / 4 · 0.58
champion_departed0.09
Generally positive. The latency complaints are real and persistent but have not stopped anyone working; their lead engineer called it 'annoying, not disqualifying'. They renewed last year after making the same complaint.
Umber FoodsARR $900k human needed
renews0.19
churn_driverprice · 0.45
expansion_readiness1.26 / 4 · 0.66
reliability_pain1.25 / 4 · 0.42
champion_departed0.76
We could not get a QBR scheduled. The original buyer left in the spring and no one has claimed ownership. The one remaining active team uses a narrow slice of the product that their BI tool could plausibly cover. Finance flagged the line item in their renew...
Vantage InsuranceARR $2.6M human needed
renews0.95
churn_drivernone · 0.66
expansion_readiness3.05 / 4 · 0.40
reliability_pain2.02 / 4 · 0.60
champion_departed0.12
Their claims analytics group presented our platform as the standard at an internal architecture forum. A second BU has already started onboarding and asked about enterprise licensing. Champion asked us to bring a multi-year proposal to the renewal.
Westline RailARR $1.1M human needed
renews0.28
churn_driverprice · 0.32
expansion_readiness1.04 / 4 · 0.34
reliability_pain2.33 / 4 · 0.39
champion_departed0.26
Mixed signals. The working team is happy and the usage is real. Their procurement group has put every renewal above $1M through a competitive re-bid this year as a matter of policy, and has told us to expect one. No specific dissatisfaction was raised.
Yarrow EnergyARR $800k human needed
renews0.87
churn_driverconsolidation · 0.31
expansion_readiness0.26 / 4 · 0.63
reliability_pain3.52 / 4 · 0.43
champion_departed0.18
The ingestion problems are genuinely bad and the team is vocal about it. Offsetting that, they are eighteen months into a three-year contract with no termination-for-convenience clause, and the workload is embedded in a regulatory filing process they cannot...
Initiative bet sizing
Six candidates asking for $20.0M against a $14.0M envelope.
Compliance & Trust SuiteCost $3.2M human needed
demand_evidence3.94 / 4 · 0.85
unblocks_pipeline0.88
defensibility1.79 / 4 · 0.57
delivery_risk0.99 / 4 · 0.54
time_to_revenuewithin_2q · 0.52
retention_impact0.92 / 4 · 0.64
Named as a hard gate by Borealis Health ($2.1M) and Juniper Payments ($2.4M). Requested by four current customers in the last two QBRs.
Customer-VPC DeploymentCost $4.1M human needed
demand_evidence2.98 / 4 · 0.63
unblocks_pipeline0.18
defensibility2.89 / 4 · 0.42
delivery_risk3.01 / 4 · 0.65
time_to_revenuetwo_to_four_q · 0.57
retention_impact3.55 / 4 · 0.51
Hard requirement for Glacier Mutual ($1.8M open) and the stated reason Summit Logistics ($2.2M) is piloting a competitor. Two other regulated prospects asked during discovery.
Reliability ProgramCost $2.6M human needed
demand_evidence3.44 / 4 · 0.40
unblocks_pipeline0.09
defensibility2.00 / 4 · 0.90
delivery_risk0.93 / 4 · 0.74
time_to_revenuewithin_2q · 0.30
retention_impact2.99 / 4 · 0.91
Pinnacle Retail ($3.1M) made it an explicit renewal condition after four Sev-1s. Tidewater and Yarrow have sustained complaint clusters. Three references declined to be references this year citing stability.
Kestrel CopilotCost $3.8M human needed
demand_evidence1.11 / 4 · 0.36
unblocks_pipeline0.74
defensibility1.84 / 4 · 0.37
delivery_risk3.67 / 4 · 0.57
time_to_revenuebeyond_4q · 0.51
retention_impact2.11 / 4 · 0.32
Two customers said it sounded interesting when shown a mockup. No customer has asked for it unprompted. No prospect has made it a requirement.
Enterprise GTM ExpansionCost $4.4M human needed
demand_evidence3.78 / 4 · 0.67
unblocks_pipeline0.82
defensibility0.94 / 4 · 0.73
delivery_risk1.93 / 4 · 0.37
time_to_revenuetwo_to_four_q · 0.56
retention_impact0.17 / 4 · 0.71
Not a product ask. Inbound enterprise leads have exceeded current coverage for three consecutive quarters and 40% went unworked.
Usage-Based Pricing MigrationCost $1.9M human needed
demand_evidence3.65 / 4 · 0.56
unblocks_pipeline0.18
defensibility1.99 / 4 · 0.62
delivery_risk2.78 / 4 · 0.22
time_to_revenuetwo_to_four_q · 0.53
retention_impact1.77 / 4 · 0.15
Three large customers have asked for consumption pricing because seat counts understate their value. Two smaller ones would pay less under it.
Cost flexibility
How much of the $14.5M cost base can actually be pulled if the growth case misses?
Legacy colocation contractAnnual $2.4M human needed
discretionary0.90
cut_impact3.84 / 4 · 0.73
lock_inmulti_year · 0.42
Three-year term expiring Q4 FY28. Termination for convenience is not permitted. Early exit triggers the remaining balance as a lump sum. Hosts the primary ingestion tier for eleven enterprise customers still on the pre-cloud architecture.
Third-party observability vendorAnnual $1.1M human needed
discretionary0.86
cut_impact1.85 / 4 · 0.57
lock_innone · 0.59
Annual commitment renewing each January with a 60-day non-renewal notice. Overage is billed monthly at list. An internal spike last year showed the core use cases could be served by our own stack in roughly one quarter of engineering effort.
Field marketing and conferencesAnnual $1.8M human needed
discretionary0.81
cut_impact2.07 / 4 · 0.82
lock_innone · 0.59
Per-event sponsorship agreements booked quarter by quarter. Two large Q3 sponsorships are non-refundable once invoiced; the rest can be cancelled with 30 days notice at no penalty. Attribution analysis credits this program with 12% of new pipeline.
Offshore QA contractorsAnnual $1.3M human needed
discretionary0.20
cut_impact3.08 / 4 · 0.80
lock_inannual_commit · 0.28
Master services agreement with 45-day termination for convenience. Twenty-two contractors currently cover regression testing for the two oldest product lines, which have no automated coverage.
Brand agency retainerAnnual $600k review
discretionary0.96
cut_impact1.00 / 4 · 0.91
lock_innotice_period · 0.70
Monthly retainer, 30-day cancellation, no minimum term. Currently producing a website refresh and a category-positioning campaign. No revenue attribution has ever been established for this spend.
Primary cloud commitmentAnnual $5.2M human needed
discretionary0.14
cut_impact3.75 / 4 · 0.62
lock_inannual_commit · 0.61
Annual spend commitment with tiered discounts. Falling below the committed amount forfeits the discount tier retroactively for the year. Runs all production workloads for the current architecture.
Office leasesAnnual $1.4M human needed
discretionary0.09
cut_impact1.15 / 4 · 0.62
lock_inmulti_year · 0.49
Two leases, both multi-year, expiring FY29 and FY30. No sublease restriction in either. Current badge data shows average occupancy at 31% of capacity.
Sales tooling stackAnnual $740k human needed
discretionary0.92
cut_impact1.15 / 4 · 0.63
lock_inannual_commit · 0.52
Six annual SaaS subscriptions renewing on staggered dates. A license audit found 34% of seats unused for more than 90 days. Two of the six tools have overlapping functionality.
Honesty
What is emulated, and what that costs
Backend
What it is
Calibrated?
This run
oracle
Offline stand-in. Reads no evidence; draws distributions whose stated
probability equals its true chance of being right, from hidden labels.
By construction
$0.0009
claude
A frontier LLM in a Jev-shaped adapter. Real intelligence reading the real
evidence; the adapter enforces the type contract so a malformed answer
cannot reach the forecast.
Self-reported, so no. --ensemble k helps and costs k times as much.
~$0.06
live
The real endpoint. Wire-correct and never run, because we have no key.
Claimed, unverified here
—
The same 20,689 input tokens cost
$0.0009 at Jev's published rate
($0.042/MTok in, output free) and
$0.06 at frontier-LLM rates
($3/$15 per MTok) — about
71×. At this size the absolute numbers are rounding
errors either way. The gap starts to matter when the same questions run against
every deal, every account and every ticket, nightly.
Where the emulation is honestly weaker than the thing it stands in for
The oracle does not read the evidence at all — it is a calibrated random
variable keyed to hidden labels, which exercises the plumbing but proves nothing
about whether a real model could reach those answers from the notes. The Claude
backend does read the evidence, but a model reporting its own probabilities is
systematically overconfident, and temperature is rejected on the
Claude 5 family, so ensemble variation comes from natural nondeterminism rather
than a sampling knob. And the confidence formula here is a reconstruction:
TypeSafe documents confidence as a statistic derived from the distribution but
does not publish which one, so this uses normalized negative entropy and exposes
--confidence-measure to swap it.
Built as a demo of the TypeSafe AI / Jev System One contract applied to financial
forecasting. Source lives in businesses/jev-forecast/; run
./jevctl.py forecast to reproduce every number on this page.
Kestrel Data Systems and all figures are fictional. Not affiliated with TypeSafe AI.