AI App Development

OpenAI Decisions API: Pat Simmons Tests 7 Real Uses

Pat Simmons put OpenAI's new Decisions API into real apps instead of stopping at a playground example. His video and written breakdown cover seven everyday use cases, a game experiment, full build prompts, measured costs, and a direct comparison with Jev. This guide puts the demonstrations, source code, and important limits in one place.

The short answer

The Decisions API is a public-beta API for small, fast, bounded decisions on text or images. It can judge a yes/no proposition, choose from supplied options, or score an ordered scale. Pat's strongest examples use it to choose the next step while ordinary code handles timers, playback, editing, or UI. His demos are promising prototypes, not proof of production reliability or universal superiority to Jev.

Watch Pat's Full Test

Credit: The experiments, prompts, measurements, and demo repository are Pat Simmons' work at AI for Mortals. Numbers below are creator-reported results for his setup, not independent benchmarks. Safety and deployment cautions are editorial additions.

What the Decisions API Actually Does

According to OpenAI's API guide, a request to POST /v1/decisions uses gpt-6-luna in this beta. The three output shapes are a probability that a statement is true, a distribution over choices you supply, and a score across ordered levels. The service accepts text and images; it is not a general chat response or a tool that automatically clicks, blurs, or edits anything. Your application must perform those actions.

OpenAI lists $0.10 per million input tokens and no output-token charge for this endpoint. That is only the decision call: screenshots, transcription, model preparation, storage, infrastructure, retries, and a human reviewer still count in the cost of a finished job. OpenAI describes it as faster than calling its Responses API for these narrow tasks; actual end-to-end latency depends on your entire pipeline.

Seven Use Cases, With Their Results and Limits

The video opens with a separate AssaultCube experiment. An initial screenshot-only approach failed. Pat then exposed game state, used plain code for movement and aiming, and let the Decisions API choose bounded actions. His reported 32 kills to 5 deaths against three easiest offline bots is a hybrid software result, not proof that a vision model learned to play unaided. The game demo code is public.

DemoWhat the decision model chosePat's reported resultImportant limit and source
1. Focus orbWhether a screen capture matches the user's focus task.Eight of eight test screens classified as expected; about $0.11 per hour at his check interval.Screen capture can disclose private material. Code.
2. Mac voice controlWhich permitted desktop or browser command follows a transcription.17 of 17 synthetic test commands; 362 ms median end-to-end in his test.Not a live-microphone reliability test; OS permissions matter. Code.
3. X feed cleanerWhether a post looks like an ad or engagement bait.16 of 16 synthetic feed posts categorized as expected; about $0.0063 per 100 posts.Live feeds and changing page layouts need further testing. It folds posts, not publishes on your behalf. Code.
4. Auto-blurWhether a sampled video frame contains private information, and roughly where.84 of 144 frames flagged in a 72-second synthetic clip; about $0.06 for the scan.The original frames are sent for analysis before the resulting blur. Do not assume this protects confidential footage. Review the report and code.
5. Wikipedia raceWhich allowed outgoing link is most promising toward a target article.All six routes reached the target; OpenAI was faster in four, Jev cheaper in six.A six-route comparison, not a general benchmark. See the interactive demo or local code.
6. Calorie trackerFood category, portion estimate, then a matching food entry.A working iPhone prototype combining image decisions with USDA food data and plain-code arithmetic.Portion and nutrition accuracy are unvalidated. The written prompt is available; this app's code was not in the linked repo when checked. Never bundle a production API key in a mobile app.
7. Bad-take finderWhether a transcript segment is a flub, restart, or keeper.Seven takes cut across 4.75 minutes of raw video; about $0.0091 in Decisions API calls.Transcription dominated runtime and adds cost. It does not repair a stumble inside a kept line. See the report and code.

Pat's original article contains the long, copyable build prompts for each example. The MIT-licensed repo includes READMEs and setup notes for the public code. Some demos assume macOS or a particular browser. Inspect permissions and dependencies before running a repo locally.

What the Jev Comparison Shows

TypeSafe's Jev is another small decision-oriented model. Pat's six Wikipedia races are useful for one question: which model helps this link-selection loop finish faster at lower cost? OpenAI finished faster in four races; Jev cost less in all six. His app also performs retrieval, link collection, and retries, so the outcome is an application-level result, not a clean inference benchmark.

At the published endpoint rates, Jev's provider lists about $0.042 per million input tokens through OpenRouter, while OpenAI lists $0.10 per million for Decisions. Prices, routing, modality, and model revisions can change. The practical answer is to run your own labeled cases and compare task accuracy, abstentions, latency, total cost, and privacy requirements, not just token price.

How to Start Without Shipping an Unsafe Prototype

  1. Pick a bounded decision. Start with three to five unambiguous choices, such as needs_review, safe_to_queue, and irrelevant. If you need paragraphs or arbitrary JSON, use a generation endpoint instead.
  2. Try the Decisions playground. Compare predicate, choice, and score outputs against the official examples. Access and billing may depend on your API account.
  3. Build a small labeled test set. Include common cases, ambiguous cases, and expensive mistakes. Evaluate the decision separately from the software action it triggers.
  4. Keep the key server-side. OpenAI's API security guidance says not to expose secret keys in browser or mobile clients. A prototype prompt that injects a key into an iPhone build is not a production pattern.
  5. Add an abstain path and approval. Low-confidence or high-impact outputs should wait for a person. For recordings and screenshots, decide whether uploading the unredacted original is permissible before sending it to any API.
  6. Measure a complete job. Track correct decisions, false positives, false negatives, wall-clock latency, non-decision model costs, review time, and actual user outcome. Recheck data controls for your region and workload.

Jump to the Part You Need

TimeChapter
00:00Game opener and introduction
00:44Start building the Pomodoro timer
01:47How the API works
03:43Decisions API vs Jev
04:59Getting started
05:30Focus orb
06:54Mac voice control
10:47X feed cleaner
12:37Auto-blur private information
14:00Wikipedia race
17:13iPhone calorie tracker
19:35Editing out bad takes
24:43Final verdict

Turn a Small Decision Into a Business Test

The business opportunity is not selling a generic "AI that decides anything." It is removing one expensive review bottleneck while keeping exceptions and final actions under human control. Copy this into any AI assistant to find a narrow pilot.

Business idea prompt

Design a decision-triage pilot

Find one buyer, one bounded choice, and a test you can measure.

Ready to copy

Prompts, Code, Reports, and Official Sources

Common questions

Is the Decisions API a general-purpose chatbot?
No. OpenAI describes a public-beta API for fast, bounded choices, predicate probabilities, and ordered scores. Use the Responses API when you need long free-form or structured generation.
Are all seven demos available as open-source code?
Pat links a repository with folders for six named applications plus the AssaultCube opener. The iPhone calorie tracker is described with a prompt in his article, but its source was not present in the linked repo when checked.
Does Pat prove OpenAI is better than Jev?
No. His six-route Wikipedia race showed OpenAI faster in four routes and Jev cheaper in all six. That small, task-specific comparison is not a general model benchmark.
Share
X LinkedIn Reddit
Build Yours

Want a system
like this one?

Book a free 30-minute call. We map your situation, identify the highest-impact automation, and figure out if we are a fit.

Book Free 30-min Call