Pat Simmons put OpenAI's new Decisions API into real apps instead of stopping at a playground example. His video and written breakdown cover seven everyday use cases, a game experiment, full build prompts, measured costs, and a direct comparison with Jev. This guide puts the demonstrations, source code, and important limits in one place.
The Decisions API is a public-beta API for small, fast, bounded decisions on text or images. It can judge a yes/no proposition, choose from supplied options, or score an ordered scale. Pat's strongest examples use it to choose the next step while ordinary code handles timers, playback, editing, or UI. His demos are promising prototypes, not proof of production reliability or universal superiority to Jev.
Watch Pat's Full Test
Credit: The experiments, prompts, measurements, and demo repository are Pat Simmons' work at AI for Mortals. Numbers below are creator-reported results for his setup, not independent benchmarks. Safety and deployment cautions are editorial additions.
What the Decisions API Actually Does
According to OpenAI's API guide, a request to POST /v1/decisions uses gpt-6-luna in this beta. The three output shapes are a probability that a statement is true, a distribution over choices you supply, and a score across ordered levels. The service accepts text and images; it is not a general chat response or a tool that automatically clicks, blurs, or edits anything. Your application must perform those actions.
OpenAI lists $0.10 per million input tokens and no output-token charge for this endpoint. That is only the decision call: screenshots, transcription, model preparation, storage, infrastructure, retries, and a human reviewer still count in the cost of a finished job. OpenAI describes it as faster than calling its Responses API for these narrow tasks; actual end-to-end latency depends on your entire pipeline.
Seven Use Cases, With Their Results and Limits
The video opens with a separate AssaultCube experiment. An initial screenshot-only approach failed. Pat then exposed game state, used plain code for movement and aiming, and let the Decisions API choose bounded actions. His reported 32 kills to 5 deaths against three easiest offline bots is a hybrid software result, not proof that a vision model learned to play unaided. The game demo code is public.
| Demo | What the decision model chose | Pat's reported result | Important limit and source |
|---|---|---|---|
| 1. Focus orb | Whether a screen capture matches the user's focus task. | Eight of eight test screens classified as expected; about $0.11 per hour at his check interval. | Screen capture can disclose private material. Code. |
| 2. Mac voice control | Which permitted desktop or browser command follows a transcription. | 17 of 17 synthetic test commands; 362 ms median end-to-end in his test. | Not a live-microphone reliability test; OS permissions matter. Code. |
| 3. X feed cleaner | Whether a post looks like an ad or engagement bait. | 16 of 16 synthetic feed posts categorized as expected; about $0.0063 per 100 posts. | Live feeds and changing page layouts need further testing. It folds posts, not publishes on your behalf. Code. |
| 4. Auto-blur | Whether a sampled video frame contains private information, and roughly where. | 84 of 144 frames flagged in a 72-second synthetic clip; about $0.06 for the scan. | The original frames are sent for analysis before the resulting blur. Do not assume this protects confidential footage. Review the report and code. |
| 5. Wikipedia race | Which allowed outgoing link is most promising toward a target article. | All six routes reached the target; OpenAI was faster in four, Jev cheaper in six. | A six-route comparison, not a general benchmark. See the interactive demo or local code. |
| 6. Calorie tracker | Food category, portion estimate, then a matching food entry. | A working iPhone prototype combining image decisions with USDA food data and plain-code arithmetic. | Portion and nutrition accuracy are unvalidated. The written prompt is available; this app's code was not in the linked repo when checked. Never bundle a production API key in a mobile app. |
| 7. Bad-take finder | Whether a transcript segment is a flub, restart, or keeper. | Seven takes cut across 4.75 minutes of raw video; about $0.0091 in Decisions API calls. | Transcription dominated runtime and adds cost. It does not repair a stumble inside a kept line. See the report and code. |
Pat's original article contains the long, copyable build prompts for each example. The MIT-licensed repo includes READMEs and setup notes for the public code. Some demos assume macOS or a particular browser. Inspect permissions and dependencies before running a repo locally.
What the Jev Comparison Shows
TypeSafe's Jev is another small decision-oriented model. Pat's six Wikipedia races are useful for one question: which model helps this link-selection loop finish faster at lower cost? OpenAI finished faster in four races; Jev cost less in all six. His app also performs retrieval, link collection, and retries, so the outcome is an application-level result, not a clean inference benchmark.
At the published endpoint rates, Jev's provider lists about $0.042 per million input tokens through OpenRouter, while OpenAI lists $0.10 per million for Decisions. Prices, routing, modality, and model revisions can change. The practical answer is to run your own labeled cases and compare task accuracy, abstentions, latency, total cost, and privacy requirements, not just token price.
How to Start Without Shipping an Unsafe Prototype
- Pick a bounded decision. Start with three to five unambiguous choices, such as
needs_review,safe_to_queue, andirrelevant. If you need paragraphs or arbitrary JSON, use a generation endpoint instead. - Try the Decisions playground. Compare predicate, choice, and score outputs against the official examples. Access and billing may depend on your API account.
- Build a small labeled test set. Include common cases, ambiguous cases, and expensive mistakes. Evaluate the decision separately from the software action it triggers.
- Keep the key server-side. OpenAI's API security guidance says not to expose secret keys in browser or mobile clients. A prototype prompt that injects a key into an iPhone build is not a production pattern.
- Add an abstain path and approval. Low-confidence or high-impact outputs should wait for a person. For recordings and screenshots, decide whether uploading the unredacted original is permissible before sending it to any API.
- Measure a complete job. Track correct decisions, false positives, false negatives, wall-clock latency, non-decision model costs, review time, and actual user outcome. Recheck data controls for your region and workload.
Jump to the Part You Need
| Time | Chapter |
|---|---|
| 00:00 | Game opener and introduction |
| 00:44 | Start building the Pomodoro timer |
| 01:47 | How the API works |
| 03:43 | Decisions API vs Jev |
| 04:59 | Getting started |
| 05:30 | Focus orb |
| 06:54 | Mac voice control |
| 10:47 | X feed cleaner |
| 12:37 | Auto-blur private information |
| 14:00 | Wikipedia race |
| 17:13 | iPhone calorie tracker |
| 19:35 | Editing out bad takes |
| 24:43 | Final verdict |
Turn a Small Decision Into a Business Test
The business opportunity is not selling a generic "AI that decides anything." It is removing one expensive review bottleneck while keeping exceptions and final actions under human control. Copy this into any AI assistant to find a narrow pilot.
Design a decision-triage pilot
Find one buyer, one bounded choice, and a test you can measure.
Prompts, Code, Reports, and Official Sources
- Creator account: Pat Simmons' full video and AI for Mortals article with every published build prompt.
- Run the demos: GitHub repository, Wikipedia race, auto-blur report, and bad-takes report. The interactive race asks for provider keys; inspect a third-party page's security posture before entering one, or run the local repo.
- Official API facts: OpenAI Decisions guide and pricing, API changelog, key security, and data controls.
- Comparison: TypeSafe's Jev introduction and current provider listing.