Direct Answer
An AI forward deployed engineer is the person who closes the gap between a powerful model and a workflow that reliably produces business value. The FDE learns how work is really done, identifies the right intervention point, builds the system, creates evidence that it works, deploys it into the existing stack, and stays close enough to users to improve it.
That is why this role is becoming more valuable. Access to capable models is broadening, but deployment remains specific. Two companies can buy the same model and get radically different results because their data, process design, integrations, permissions, evals, adoption, and operational discipline are different.
The interview's most useful framework is not the million-dollar headline. It is the delivery loop: audit the workflow, build evals, deploy carefully, measure the result, and repeat. Do that well and you are already practicing the core of the job, whether your title says FDE, applied AI engineer, solutions architect, implementation engineer, or AI consultant.
Credits: the interview is hosted by Greg Isenberg with Vas from Varick Agents. Follow Greg on X and Vas on X. The original episode description also links to The FDE Blueprint and Varick's FDE in 30 Days resource.
Source Note
The supplied transcript is the source for Vas's career framework, audit-evals-deployment loop, compensation discussion, client examples, and 30-day plan. Those are interview claims and recommendations, not promises that following the plan will produce a job or a specific income.
I checked the role and compensation details on 21 July 2026 against current public listings from OpenAI and Palantir. OpenAI describes FDEs as owning discovery, technical scoping, system design, build, production rollout, adoption, and eval-driven feedback. Its San Francisco listing publishes a base range of $162,000 to $280,000 plus equity. Palantir's New York FDSE listing publishes $135,000 to $200,000, with possible stock, sign-on, and other incentives. Location, seniority, equity value, company stage, performance, and negotiation can materially change total compensation.
The episode repeats the widely quoted claim that 95% of generative AI pilots fail. The underlying 2025 MIT NANDA report is more specific: it says 95% of organizations in its sample saw no measurable return and only 5% of task-specific tools reached production. The report also calls its findings preliminary and notes selection bias, inconsistent success definitions, interview-based estimates, and a six-month observation period. Treat the number as a warning about workflow fit and measurement, not a universal law for every AI project.
Link Map
| Resource | Status | Why it matters |
|---|---|---|
| FDE: The $1M/Year AI Job Explained | Primary interview | Vas's definition of the role, Palantir framing, audit-evals-deployment loop, compensation discussion, and 30-day plan. |
| Greg Isenberg and Vas | Creator and guest credits | The host's channel and the guest's public profile. |
| Varick Agents | Guest company | Varick's stated process: opportunity audit, process architecture, production deployment, and ongoing system optimization. |
| The FDE Blueprint and FDE in 30 Days | Episode resources | The companion learning routes supplied in the video description. |
| Palantir FDSE role | Official job listing | Customer embedding, architecture, application building, stakeholder work, travel expectations, qualifications, and a published salary range. |
| Palantir Ontology overview | Official documentation | How Palantir connects business objects, relationships, actions, functions, and governance into an operational layer. |
| OpenAI FDE role | Official job listing | A current definition of end-to-end frontier-model deployment, adoption, eval-driven feedback, technical scope, travel, and compensation. |
| OpenAI evals guide and business eval framework | Official technical guidance | Specify what good means, measure real examples, analyze errors, and improve the system. |
| Anthropic model selection | Official technical guidance | Build task-specific benchmark tests, use actual prompts and data, compare edge cases, and weigh performance against cost. |
| MIT NANDA: The GenAI Divide | Preliminary research report | The source behind the 95% claim, including methodology, findings, success definition, and limitations. |
| NIST AI RMF and AI Resource Center | Public risk guidance | Govern, map, measure, and manage AI risk; document go/no-go decisions, oversight, testing, and operational outcomes. |
| OWASP AI Agent Security Cheat Sheet | Security guidance | Least privilege, tool allowlists, human approval, isolation, audit logs, monitoring, memory controls, and incident response. |
Episode Guide
| Time | Topic | Practical takeaway |
|---|---|---|
| 00:00 | The AI deployment opportunity | Access to intelligence is spreading; applied judgment and implementation become the differentiators. |
| 02:03 | What is an FDE? | An engineer who sits close to the real problem, users, systems, and business outcome. |
| 04:09 | Palantir's role | Embed with customers and shape the software around the operational reality. |
| 06:16 | Where intelligence belongs | Use LLMs for judgment-heavy ambiguity, normal code for rules, and humans for high-consequence decisions. |
| 11:26 | What FDEs earn | The rare combination of engineering and consulting can command high pay, but the ceiling is not the norm. |
| 14:59 | Two kinds of judgment | Technical architecture and stakeholder communication reinforce each other. |
| 17:40 | How work gets done | Observe the workflow, find leverage, prototype, prove, and operate. |
| 20:40 | Audit, evals, deployment | The core delivery loop and the structure used throughout the episode. |
| 22:56 | Which LLM? | Learn one stack first, but make production decisions from task-specific evidence. |
| 27:36 | Audit | Map the real process, including exceptions, handoffs, systems, and decision points. |
| 31:47 | Evals | Use a representative dataset and failure taxonomy to make quality visible. |
| 32:57 | Deployment | Integrate with the stack people already use, then expand autonomy gradually. |
| 38:59 | 30-day plan | Build, harden, measure, and defend one agentic workflow. |
What an AI FDE Actually Does
A good FDE is not dropped into a company to install a chatbot. The job starts before the model choice and continues after the first deployment. OpenAI's current listing is unusually clear: the FDE owns discovery, technical scoping, system design, build, and production rollout, then measures success through adoption, workflow impact, and eval-driven feedback.
| Stage | FDE responsibility | Evidence produced |
|---|---|---|
| Discovery | Interview operators, observe work, map systems, exceptions, controls, and current costs. | Workflow map, pain-point inventory, baseline metrics, risk register. |
| Scoping | Choose one valuable boundary and define what the system will and will not do. | Project brief, definition of done, acceptance criteria, owner. |
| Architecture | Decide what belongs in code, an LLM, retrieval, human review, or an existing product. | System diagram, data contracts, permission model, threat model. |
| Build | Create integrations, application logic, model calls, interfaces, queues, and review gates. | Working system, tests, telemetry, runbook. |
| Evaluate | Test representative tasks and edge cases; classify failures and improve the system. | Golden dataset, rubrics, pass rates, failure taxonomy, cost and latency report. |
| Deploy | Run in shadow mode, train users, migrate carefully, monitor, and increase autonomy by risk tier. | Launch plan, approvals, rollback path, dashboards, incident ownership. |
| Adopt and improve | Watch actual usage, support operators, capture exceptions, and feed lessons back into product and research. | Adoption, quality, override, outcome, and ROI trends. |
This is why the role is demanding. It combines software engineering, product management, solutions architecture, operations, change management, security, and consulting. The best FDE is comfortable writing code in the morning, interviewing a finance operator after lunch, and explaining risk and ROI to an executive later that day.
The Palantir Blueprint
Palantir describes its forward deployed software engineer role as the original blueprint: engineers work side by side with customers, rapidly understand hard problems, make architecture decisions, build custom applications, and carry projects from idea to deployment. That embedded posture is the defining feature.
The deeper lesson is the Ontology. Palantir's documentation defines it as an operational layer that connects data and models to real-world objects, relationships, actions, functions, and security. In plain English, the software represents how the company actually works, not only the tables it stores.
You do not need Palantir to use this mental model. For any FDE project, map four things:
- Data: what facts are required and who is allowed to access them?
- Logic: what rules, calculations, models, and judgment determine the decision?
- Action: what changes in the real system after the decision?
- Security: who can initiate, approve, inspect, reverse, and audit that action?
That is a better design starting point than asking, "Where can we add an agent?" The agent should serve a decision or workflow, not become an expensive object searching for a use case.
Deciding Where Intelligence Belongs
One of Vas's strongest points is that not every step deserves an LLM. Reliable systems use different forms of computation for different jobs.
| Use | Best fit | Examples |
|---|---|---|
| Deterministic code | Stable rules, arithmetic, validation, authorization, state transitions, and compliance constraints. | Tax calculations, required fields, spending limits, duplicate checks, API schemas. |
| LLM or multimodal model | Ambiguous language, unstructured documents, flexible classification, synthesis, drafting, or tool planning. | Extracting clauses, triaging support cases, drafting an exception memo, summarizing a meeting. |
| Human review | High-consequence, novel, subjective, regulated, reputational, or irreversible decisions. | Wire transfers, contract approval, medical action, firing, public claims, destructive system changes. |
| Hybrid | Most serious production workflows. | Model proposes, code validates, policy checks, human approves, system executes, logs capture the result. |
A practical rule: use the model where uncertainty creates value, and use normal software where certainty is available. Then place human approval where the cost of a wrong action exceeds the speed benefit of automation.
The $1M Salary Reality
The episode's title is deliberately provocative. Vas says some roles can reach seven figures when salary, equity, scarcity, business impact, and seniority align. That may happen, especially at high-growth companies or for senior people who create major commercial value. It should not be treated as the standard market rate.
| Evidence checked 21 Jul 2026 | Published cash range | Important context |
|---|---|---|
| OpenAI FDE, San Francisco | $162K-$280K base | Equity offered; 5+ years of relevant experience; hybrid; up to 50% travel. |
| Palantir FDSE, New York | $135K-$200K salary | Possible restricted stock, sign-on, and incentives; travel up to 25%; 1+ year post-college experience. |
| Interview ceiling | Up to $1M total annual compensation | Guest-reported exceptional outcome, not a published typical salary or guarantee. |
The honest career message is still attractive. Public base ranges are already strong, and the role offers unusual leverage because the engineer can influence revenue, cost, risk, adoption, and product direction. But compensation follows evidence. Build the ability to ship systems and explain their value before optimizing for the most dramatic number.
Two Kinds of Judgment
Vas divides the role into communication judgment and engineering judgment. The distinction is useful because weakness on either side can sink a deployment.
| Engineering judgment | Communication judgment |
|---|---|
| Choose the right system boundary. | Ask operators how work really happens, not how the policy says it happens. |
| Separate deterministic logic from model reasoning. | Translate technical tradeoffs into risk, cost, time, and business outcomes. |
| Design schemas, retries, fallbacks, permissions, queues, and observability. | Set expectations about uncertainty, review, rollout, and failure. |
| Build evals and analyze failure clusters. | Negotiate success criteria with the people accountable for the workflow. |
| Operate the system under real load and exceptions. | Earn enough trust for users to report failures instead of quietly abandoning the tool. |
This is the art-plus-science combination discussed in the interview. It is rare because many engineers avoid messy stakeholder work while many consultants cannot inspect or repair the system they recommend.
The Core Loop: Audit -> Evals -> Deployment
The episode reduces the job to three connected phases. The order matters.
- Audit: find the workflow worth rebuilding and define the baseline.
- Evals: convert "this feels good" into repeatable evidence about quality, cost, latency, and risk.
- Deployment: put the system into real operations with permissions, monitoring, user ownership, and a rollback path.
Then the loop repeats. Production reveals edge cases that the audit missed. Those cases become evals. The improved system returns to production. The FDE's job is not to make uncertainty disappear; it is to make uncertainty measurable, bounded, and operationally manageable.
Phase 1: Audit the Real Workflow
Start by observing a workflow, not interviewing the executive who only sees its output. Talk to the people doing the work. Ask them to walk through the last real example, including the awkward spreadsheet, copied email, exception, approval delay, and workaround.
Questions an FDE should answer
- What starts the workflow, and what counts as finished?
- Who owns the outcome, and who touches the work?
- Which systems, files, inboxes, and informal channels are involved?
- Where does work wait, repeat, fail, or require re-entry?
- Which decisions are rules, and which require judgment?
- What are the common exceptions and the dangerous rare cases?
- What data is sensitive, regulated, or contractually restricted?
- How is quality measured today?
- What is the current cost in time, money, errors, delay, and opportunity?
- Who can approve a pilot and who can stop production?
Score opportunities on business impact, repeatability, data readiness, integration effort, failure cost, executive ownership, and measurability. The best first project is not always the flashiest. It is the one with a clear boundary, enough volume to learn, a qualified reviewer, and an outcome the business already cares about.
Phase 2: Turn Uncertainty Into Evidence
OpenAI defines evals as tests that check model outputs against criteria you specify. Its business framework is simple: specify what good means, measure the system under real conditions, then improve based on errors. Anthropic gives similar advice: build benchmark tests for the actual use case, use real prompts and data, test edge cases, and compare performance and cost.
The unit of evaluation should usually be the complete workflow outcome, not a sentence from the model. A support agent can produce a polite answer and still fail if it chose the wrong account, skipped a policy check, or triggered the wrong action.
Build a representative eval set
- Normal examples that represent most real traffic.
- Hard examples that require context or multi-step tool use.
- Known past failures and costly exceptions.
- Adversarial or untrusted inputs, including prompt injection attempts.
- Incomplete, contradictory, duplicated, and stale data.
- Cases that must escalate to a human.
- Examples across languages, teams, customer types, and document formats when relevant.
Measure more than accuracy
| Dimension | Example metric | Failure question |
|---|---|---|
| Task success | Correct outcome rate | Did the workflow finish correctly? |
| Safety | Unauthorized-action rate | Did the agent exceed its scope or ignore an approval? |
| Escalation | Correct handoff rate | Did it know when not to act? |
| Reliability | Completion and recovery rate | What happens when an API, model, or dependency fails? |
| Latency | P50 and P95 completion time | Is it fast enough for the real workflow? |
| Cost | Cost per successful task | Did cheaper execution create more retries or review work? |
| Adoption | Weekly active users and override rate | Do operators trust and use it? |
| Business impact | Hours, cycle time, errors, revenue, or risk avoided | Did the system change the outcome the sponsor funds? |
The transcript uses a useful 50-run example: 41 pass and nine fail. The next step is not to celebrate 82%. Classify the nine failures. Were they caused by missing context, a bad tool result, weak instructions, a schema error, model judgment, or an impossible case? Fix the system component responsible, rerun the set, and add the failures to regression tests.
Phase 3: Deploy Into Existing Systems
Production is where an impressive demo meets permissions, rate limits, inconsistent data, impatient users, procurement, security, and the one old system nobody documented. Vas argues for building on top of the tools the company already uses, such as Salesforce, SAP, NetSuite, Workday, Concur, Expensify, or Gong, instead of demanding a full migration first.
A sensible autonomy ladder looks like this:
- Offline replay: run the system on historical cases without affecting live work.
- Shadow mode: process live inputs but compare recommendations with human decisions.
- Assist: draft or recommend while a human explicitly approves every action.
- Bounded execution: automate low-risk cases inside strict rules and escalate exceptions.
- Expanded autonomy: widen scope only after evidence, monitoring, and incident handling are stable.
Every stage needs an owner, threshold, and rollback condition. "The model is better now" is not a release criterion. "The system passed 196 of 200 representative cases, escalated all high-risk exceptions, stayed under $0.18 per successful task, and ran two weeks in shadow mode without unauthorized actions" is much closer.
The 30-Day FDE Portfolio Plan
Vas compresses the learning path into four weeks. This will not turn a beginner into a senior enterprise engineer. It can produce a credible artifact that proves you understand the job beyond prompting.
| Week | Build | Required proof |
|---|---|---|
| 1: Complete the loop | Choose one real workflow. Connect the model to the minimum tools and data. Add a human checkpoint and an audit trail. | A working demo, workflow map, definition of done, sample inputs and outputs. |
| 2: Harden it | Add structured schemas, validation, retries, timeouts, idempotency, failure handling, permissions, and explicit exception paths. | Failure-mode table, test results, runbook, recovery demo. |
| 3: Make it measurable | Create a golden dataset, task-specific evals, model comparison, cost tracking, latency tracking, and business-impact estimate. | Evaluation report, failure taxonomy, cost per successful task, baseline comparison. |
| 4: Defend it | Prepare the architecture, tradeoffs, security boundaries, rollout, adoption plan, economics, and executive narrative. | Five-minute demo, technical design, one-page business case, risk register, 30/60/90-day rollout. |
The order is smart. Week one proves the happy path. Week two proves that you understand software. Week three proves that you understand probabilistic systems. Week four proves that you can explain why the company should trust and fund the work.
A Copy-Ready Portfolio Project Brief
Use this original brief to design your first FDE project. Keep the workflow bounded and replace the bracketed fields with a real process you can observe.
PROJECT: [workflow name]
BUSINESS OWNER
[name or role accountable for the outcome]
CURRENT WORKFLOW
- Trigger:
- Inputs:
- Systems involved:
- Human roles:
- Decision points:
- Common exceptions:
- Definition of finished:
BASELINE
- Monthly volume:
- Median cycle time:
- Error or rework rate:
- Human time per case:
- Current direct cost:
- Risk or opportunity cost:
SYSTEM BOUNDARY
- The agent may:
- The agent may not:
- Deterministic rules:
- Model judgment:
- Human approval required when:
EVALUATION
- Representative cases:
- Known failures:
- Minimum task success rate:
- Required escalation rate:
- Maximum unauthorized-action rate: 0
- Maximum cost per successful task:
- Maximum P95 latency:
DEPLOYMENT
- Offline replay:
- Shadow-mode period:
- Pilot users:
- Monitoring owner:
- Incident owner:
- Rollback trigger:
- Production approval:
BUSINESS CASE
- Hours saved:
- Cycle-time reduction:
- Errors avoided:
- Revenue or risk impact:
- Estimated monthly system cost:
- Expected payback period:
For a portfolio, include a two-minute failure demo. Deliberately break an API, remove a required field, inject an untrusted instruction, or submit a high-risk case. Show that the system stops, escalates, logs the reason, and recovers. That tells an employer more than another perfect-path agent video.
Security and Operations Are Part of the Job
An FDE gives models access to real data and tools. That makes security an architectural requirement, not a legal review at the end. NIST recommends explicit governance, measurement, oversight, documentation, and go/no-go decisions. OWASP recommends least privilege, tool allowlists, human approval for high-impact actions, isolated execution, logging, monitoring, memory controls, and incident response.
- Give the agent a dedicated identity, never a shared employee credential.
- Grant the smallest data and action scopes needed for the current workflow.
- Separate read, draft, approve, and execute permissions.
- Require human approval for destructive, financial, external, legal, or irreversible actions.
- Treat email, documents, web pages, tickets, and tool outputs as untrusted inputs.
- Validate structured outputs and tool parameters outside the model.
- Log prompts, retrieved context, model version, tool calls, decisions, approvals, outputs, and errors with appropriate access controls.
- Redact secrets and sensitive personal data from prompts, telemetry, and test datasets.
- Set budgets, rate limits, timeouts, loop limits, and emergency stop controls.
- Document rollback, incident triage, user notification, evidence preservation, and post-incident review.
The goal is not maximum autonomy. It is justified autonomy. Every additional permission should be earned by evidence and matched with a way to observe, interrupt, and reverse the action.
Which Model Should an FDE Use?
Vas recommends becoming deeply productive with one stack before trying to be model agnostic. That is sensible for learning. Production selection should come later, after the workflow and eval set exist.
Compare candidate models on:
- Successful workflow completion, not benchmark rank.
- Tool-use reliability and structured-output adherence.
- Performance on the domain's difficult and dangerous cases.
- Latency at normal and peak demand.
- Cost per successful task, including retries and human review.
- Context requirements, caching, and data residency.
- Provider retention, security, compliance, and availability terms.
- Fallback behavior when the primary model or provider fails.
A strong FDE may route simple classification to a smaller model, complex analysis to a stronger model, arithmetic to deterministic code, retrieval to a controlled search layer, and high-risk decisions to a person. Model routing is an output of the evals, not an aesthetic preference.
Who Should Pursue This Career?
This path fits people who enjoy both systems and people. You do not need to arrive with every skill, but you should be willing to develop across both sides.
| Starting background | Likely advantage | Gap to close |
|---|---|---|
| Software engineer | Architecture, coding, debugging, deployment. | Discovery, executive communication, process mapping, ROI. |
| Solutions engineer or consultant | Stakeholders, scoping, demos, adoption, commercial judgment. | Production code, testing, security, observability, operations. |
| Data or ML engineer | Data pipelines, modeling, measurement, experimentation. | Application UX, workflow ownership, tool integration, change management. |
| Domain operator | Deep workflow and exception knowledge. | Programming, system design, APIs, deployment, eval tooling. |
| Automation builder | Fast integration and workflow prototyping. | Software reliability, threat modeling, formal evals, enterprise governance. |
Look beyond the exact title when searching. Relevant openings may be called forward deployed engineer, forward deployed software engineer, applied AI engineer, AI solutions architect, technical deployment lead, implementation engineer, field engineer, or AI transformation engineer.
Bottom Line
The real opportunity is not that a new title suddenly guarantees a million dollars. It is that businesses now need people who can convert increasingly capable models into reliable operations. That work is difficult, specific, and valuable.
Start by doing the job before asking for the title. Choose one repeated workflow. Audit it. Build the smallest useful system. Add an eval set, failure handling, permissions, logs, human review, cost tracking, and a business metric. Run it in shadow mode. Document what failed and what improved. Then explain the system twice: once like an engineer and once like the executive paying for it.
A polished demo shows that you can build. A measured deployment shows that you can be an FDE.
Sources
- Greg Isenberg and Vas: FDE interview
- Varick Agents
- Palantir: Forward Deployed Software Engineer
- Palantir: Ontology overview
- OpenAI: Forward Deployed Engineer
- OpenAI: Working with evals
- OpenAI: How evals drive AI for business
- Anthropic: Choosing a model
- MIT NANDA: The GenAI Divide, State of AI in Business 2025
- NIST AI Risk Management Framework
- OWASP AI Agent Security Cheat Sheet