Every other prototype here invents hidden labels so the emulator can be calibrated. This one does not have to. In bridge the ground truth is simply the layout: the cards you cannot see are a fact, the contract either makes or it does not, and every judgment along the way can be marked right or wrong afterwards. That makes it the honest test case.
4♠ by South · N-S vulnerable · West 3♥ (weak) – North Pass – East Pass – South ?
4♠ makes exactly: two heart losers and one diamond, ten tricks.
A bridge engine needs thousands of these per deal and has seconds to produce them. The possible answers are known before the question is asked — which call, which card, which line, who holds the jack — so there is nothing for a language model to add by writing them out in prose, and a great deal to lose in latency. Output tokens being free matters here more than anywhere: asking about the trumps, the side suits and the contract all at once costs the same as asking about one.
The trump suit breaks 4-0 and the diamond finesse loses, so a declarer following the textbook goes down in a contract that makes. Each answer below is scored against the actual cards.
| Judgments scored | 19 |
| Accuracy | 78.9% |
| Brier score | 0.162 |
| Expected calibration error | 0.090 |
Nineteen judgments is far too few to say anything statistically about calibration, and this page does not pretend otherwise — the number is here because the deal makes the scoring honest, not because the sample is adequate. The reliability curves worth reading are on the other three prototypes, where the corpora are larger.
Four prototypes of the TypeSafe AI / Jev System One contract applied to real decisions.
Source in businesses/jev-forecast/; ./jevctl.py dashboard --all
rebuilds every page. All companies, figures and hands are fictional.
Not affiliated with TypeSafe AI. Built September 2026.