Agent Native
FDE Interview
Agent Native·An interview book for AI engineers

Forward Deployed Engineering Interview

Forward deployed engineering interviews ask the same things a customer will: how to serve a model to people who are waiting, how to keep retrieval fresh and lawful, how to build an agent that stops when it should, how to prove the system works, and how to tell an executive the plan slipped. This book takes the questions that keep coming up in FDE, applied AI and AI architect loops, gives each one an answer an interviewer can grade, and then teaches the engineering underneath it, with the numbers, code and interactive figures you need to answer the follow-ups, current to October 2026.

14 chapters + appendix·Questions, answers and runnable code·Current to October 2026·Agent Native, 2026

Chapter 1 · The role and the loop

What the Job Is and How the Loop Tests It

A customer, the forward deployed engineer and the rounds of the interview loop

What forward deployed engineers own in 2026?

Between May 4 and July 2, 2026, forward deployed engineering became a line of business. OpenAI launched a deployment company with “more than $4 billion of initial investment”; AWS put $1 billion behind an FDE organization; Microsoft committed $2.5 billion and 6,000 people; and Anthropic's new services company was valued at $1.5 billion, according to the WSJ as reported by TechCrunch. Forrester puts the total at roughly $9 billion, as Business Insider reported.

The job began as Palantir's internal name for an engineer who serves “one customer, many capabilities”, and its interview changed as fast as the money. Meta now expects candidates to use an AI assistant built into the interview, Sierra replaced its coding rounds with a two-hour build using any AI tools, and Anthropic forbids AI in live interviews unless told otherwise.

Who this book is for and how to use it

This book is for engineers preparing for forward deployed engineer (FDE), applied AI, GenAI and AI architect interviews, and for engineers in those jobs who want sharper perspectives. It assumes you write Python and have shipped something that calls a model API. The chapters that need inference servers, security reviews or executive briefings build them up from the mechanism.

Its spine is a list of 33 questions. Each one gets a card that says what the interviewer is testing, a model answer you could say in two to three minutes, the engineering underneath (sourced numbers, runnable code, figures you can drag or step through), follow-ups to try before you open the answers, and the red flags that sink an answer.

Sources carry labels. Company pages, postings, filings, papers and first-hand accounts we read are stated as fact, with a link; news reports are attributed by name; a candidate's post is one person's account; claims found only in search summaries are left out or named as unconfirmed. Postings, pay and counts are snapshots from October 7, 2026, and the market section includes a script that refreshes them.

With a loop two weeks away, read this chapter, find your rounds on the competency map near its end, read those chapters, and say the model answers aloud against a timer. With six weeks, follow the plan in the appendix. If you already do the job, go straight to the depth sections and the follow-ups.

Dev and Delta: where the role came from

On April 8, 2019, Palantir's blog published “Dev versus Delta: Demystifying engineering roles at Palantir”. Software Engineers and Forward Deployed Software Engineers (FDSEs) were its “biggest engineering roles by headcount”, called Devs and Deltas internally; Delta was “a throwback to our early days, when each team in Business Development was named after a letter in the NATO alphabet.” A Dev's focus is “one capability, many customers”, a Delta's is “one customer, many capabilities”, and Palantir's careers page for students uses the same phrases today.

Deltas “measure success in terms of impact on the customer's goal”, and the core skill the post names is “Technical decomp”, breaking a high-level problem down to “the lines of code you need to solve it”. Deltas also contribute code to the core product, validate larger feature requests against the product roadmap, and move between the two roles.

Switch between DEV and DELTA, then pick FEEDBACK to see the path that joins the two loops. Hover over or tab to any box to read what Palantir and its former engineers say happens there.

Palantir's 2020 S-1 put that path into a filing (“Our FDEs are our first line in identifying research and development opportunities for our platforms”) and showed the economics. Its sales model “often requires us to spend months and invest significant resources working with customers on pilot deployments at no or low cost to them”, and “Custom code and consulting teams” are “unsustainable alternatives”. In that model FDE time is an investment in adoption, and the platform keeps it from becoming a consulting business. Most 2026 criticism of the role asks whether a given employer still has the platform half.

Palantir still files roles under the old names: on October 7, 2026 its public Lever board listed 85 postings on the Dev team, 67 on Delta and 33 on Echo, its deployment strategists. The New York FDSE posting calls itself “The Original Forward Deployed Software Engineer”, asks for one year of post-college experience and “travel up to 25%”, pays $135,000–$200,000, and promises “increasing your pain threshold to deliver real value”. When the title began is disputed (a16z says 2011, Wikipedia cites a 2009 Science article), so the timeline starts from the 2019 post.

From Delta to a line of business

  1. AIP bootcamps
    2023
    Palantir's FY2024 10-K: bootcamps “deliver real workflows on actual customer data in days.”
  2. OpenAI starts an FDE team
    January 2025
    Two FDEs at first, per Pragmatic Engineer; 39 engineers by November, per Business Insider.
  3. a16z: Trading Margin for Moat
    June 2025
    22 of the 311 open roles on OpenAI's careers page are forward deployed or solutions engineering.
  4. Ramp's FDE team
    August 2025
    From 2 to 16 FDEs in 18 months, under the mantra “always be scoping”.
  5. Capital arrives
    May to July 2026
    About $9 billion pledged to FDE vehicles; the next section steps through it.
  6. Frontier Deployed Engineers
    October 2026
    Anthropic's Frontier Academy commits $100M to train 10,000 partner engineers by the end of 2027.
  7. First Frontier Deployed Engineer badges
    Early 2027
    A 12-week residency leading a real Claude use case, then a final assessment.

Nabeel Qureshi, who joined Palantir in 2015 when it had about 1,500 employees, left a detailed first-hand account. FDEs “were typically expected to ‘go onsite’ to the customer's offices and work from there 3-4 days per week”, and he spent a year in Toulouse on Airbus. On pilots he wrote: “you'd have a company buying an 8-12 week pilot, and we'd spend all 8-12 weeks just getting data access, and the final week scrambling to have something to demo.” Palantir's fix was product: role-based access controls, row-level policies, security markings and audit trails built into the data layer.

When access takes the whole pilot, the demo gets one week in eight to twelve, so treat data access as a day-zero dependency with an owner and a date (Chapter 13 builds it into a 90-day plan). Palantir solved that problem once, in the platform, which is the judgment behind the next question.

Related question · The role

A customer asks for a feature your platform does not have. Do you build it for them, or build it into the platform?

What it tests: Whether you weigh speed for one customer against reuse across many, know who makes the call, and plan for who maintains the result.

A strong answer

I first scope the request down to the outcome the customer is measured on, because sometimes a workaround with what the platform already does gets them there; Ramp's FDEs describe closing a gap of about three engineer-days on a customer call that way. If the need is specific to this customer's systems, I build it as an extension on their side, documented and owned by their engineers. If I have seen the same request at two or three customers, I take it to the product team with the deployments and evals as evidence, and they decide. I avoid a private fork of the core product, which Palantir's S-1 files under “unsustainable alternatives”.

Follow-ups

Two weeks in, the customer's data owner still has not granted access. What do you do?
Treat access as the critical path. Find out what the owner needs to say yes (often a security review, a list of who sees which fields, and an audit trail), bring it in writing for the one workflow in the pilot, and build on synthetic or masked data meanwhile, as OpenAI's FDEs do in early scoping according to Pragmatic Engineer. Put the access date on the sponsor's plan so a slip surfaces in week two.
Who maintains the custom code after you leave?
The customer, and the plan says so from week one. AWS describes its engagements as moving customer engineers “from observers to co-builders to autonomous operators”: they pair on the code, the runbook lives in their repository, and the last milestone is a change they ship without you.

The 2026 market in numbers you can defend

Market numbers come up in “why FDE, why now” answers, and the ones in circulation mix index values with counts and stocks with flows. Cite each dataset with its definition:

  • Lightcast, US postings appearing over a period (a flow): “from approximately 200 in 2024 to around 1,200 in 2025 - a sixfold increase”, and “More than 5,200 postings appeared in the first seven months of 2026 alone.”
  • Bloomberry, title matches on about 1,000 postings in Revealera data: up 1,165% from January–October 2024 to January–October 2025, or 12.65 times. It differs from Lightcast's sixfold on title rules, geography and deduplication.
  • Indeed, via Business Insider, an index with January 2025 set to 100: 643 in April 2025 and 5,330 in April 2026. Business Insider corrected its story (“The figures cited are indexed values relative to a January 2025 baseline”), but summaries still quote 5,330 as a count of jobs.
  • Plank's census, postings live on one day (a stock): 1,333 FDE postings at 565 companies on September 12, 2026, 388 of them (29.1%) with pay. Plank sells FDE services; its raw data is public under CC BY 4.0.

More than 5,200 postings in seven months is at least 743 a month, 7.4 times the 2025 rate, and at that pace 2026 would close near 8,900; that is our extrapolation, not a Lightcast forecast. For a growth rate with a fair denominator, Alexey Grigorev's builtin.com scrapes show FDE-titled listings rising 4.2 times between February 4 and July 22, 2026, against 2.3 times for all AI-engineering listings.

Name the dataset
Postings did not grow to 5,330: that is an index level, 53.3 times January 2025. “Sixfold” is Lightcast's US flow from 2024 to 2025, and “twelvefold” is Bloomberry's title match for January to October. Say which one you mean, with its date.

The capital is easier to pin down, because the companies announced it.

The figure opens on the last announcement. Press NEXT to start from December 2025 and watch the bar add the four FDE vehicles; dashed entries are press-reported figures, or amounts that may not be new money.

OpenAI's Deployment Company is “majority-owned and controlled by OpenAI” with “more than $4 billion of initial investment”, and its Tomoro acquisition brings “approximately 150 experienced Forward Deployed Engineers and Deployment Specialists”. AWS is “backed by a $1 billion investment”, and Microsoft's Frontier Company is “a $2.5B investment ... embedding 6,000 industry and engineering experts at customers”. Anthropic's services company, named Ode with Anthropic in July, gives no figure: the $1.5 billion is the WSJ's, via TechCrunch, and may be a valuation rather than capital. And GeekWire reports that Microsoft would not say whether its $2.5 billion is new money. The four sum to $9.0 billion, which matches Forrester's estimate.

On pay, our recomputation of Plank's data gives a median posted base midpoint of $192,500 across the 388 postings with pay: $242,500 at frontier labs (16 postings) and $150,000 at seed-stage or stealth companies (30). Bloomberry's 2025 median was $173,816, and crowd-sourced Levels.fyi data, stock included, puts Palantir's US FDE median at $255,000 across 227 submissions. Supply is short: Patrick Kellenberger of Betts Recruiting told the WSJ, as Pragmatic Engineer reported, that “only maybe 10% of the market” wants the role.

To refresh these numbers before an interview, query the boards. Ashby, Greenhouse and Lever serve public job-board APIs, and fde_postings.py counts FDE-titled and wider FDE-family titles with the title rules behind this chapter's counts.

fde_postings.py
python
"""Count FDE-titled and FDE-family postings on public job boards.

Python 3.11+ and requests. The boards and title rules are the ones behind
the counts in Chapter 1 (read 2026-10-07); boards change daily.
"""
import re

import requests

FDE_TITLED = re.compile(r"forward[- ]deployed|deployed engineer", re.I)
FDE_FAMILY = re.compile(
    r"deployment (lead|strategist|engineer|specialist)"
    r"|agent (deployment|product manager|pm|development)"
    r"|software engineer, agent"
    r"|applied ai (engineer|architect|strategist)"
    r"|technical deployment",
    re.I,
)

BOARDS = [  # (company, ATS, board slug)
    ("OpenAI", "ashby", "openai"),
    ("Sierra", "ashby", "sierra"),
    ("Cohere", "ashby", "cohere"),
    ("Anthropic", "greenhouse", "anthropic"),
    ("Databricks", "greenhouse", "databricks"),
    ("Scale AI", "greenhouse", "scaleai"),
    ("Palantir", "lever", "palantir"),
]


def titles(ats: str, slug: str) -> list[str]:
    """Every live posting title on one public board."""
    if ats == "ashby":  # {"apiVersion": ..., "jobs": [{"title": ...}, ...]}
        url = f"https://api.ashbyhq.com/posting-api/job-board/{slug}"
        return [job["title"] for job in requests.get(url, timeout=30).json()["jobs"]]
    if ats == "greenhouse":  # {"jobs": [{"title": ...}, ...], "meta": {...}}
        url = f"https://boards-api.greenhouse.io/v1/boards/{slug}/jobs"
        return [job["title"] for job in requests.get(url, timeout=30).json()["jobs"]]
    if ats == "lever":  # a JSON array of postings; the title is "text"
        url = f"https://api.lever.co/v0/postings/{slug}?mode=json"
        return [post["text"] for post in requests.get(url, timeout=30).json()]
    raise ValueError(f"unknown ATS: {ats}")


def kind(title: str) -> str:
    if FDE_TITLED.search(title):
        return "titled"
    return "family" if FDE_FAMILY.search(title) else "other"


if __name__ == "__main__":
    print(f"{'company':<12}{'total':>7}{'titled':>8}{'family':>8}{'share':>8}")
    for company, ats, slug in BOARDS:
        kinds = [kind(t) for t in titles(ats, slug)]
        titled = kinds.count("titled")
        family = titled + kinds.count("family")  # the family includes titled roles
        share = titled / max(len(kinds), 1)
        print(f"{company:<12}{len(kinds):>7}{titled:>8}{family:>8}{share:>8.1%}")

On October 7, 2026 those rules gave Databricks 99 FDE-titled postings of 888, Palantir 80 of 314, OpenAI 26 of 823 (102 in the family), Anthropic 6 of 639 (57 in the family) and Sierra none of 197 by title but 69 in the family, since Sierra calls the role Software Engineer, Agent. The pay medians come from Plank's JSON:

plank_median.py
python
"""Median posted base pay in Plank's FDE census (CC BY 4.0).

Python 3.11+ and requests.
Usage: python plank_median.py [field] [region], for example
       python plank_median.py segment "United States"
"""
import sys
from statistics import median

import requests

FIELD = sys.argv[1] if len(sys.argv) > 1 else "segment"
REGION = sys.argv[2] if len(sys.argv) > 2 else None

data = requests.get("https://joinplank.com/fde-jobs.json", timeout=30).json()
jobs = data["jobs"] if isinstance(data, dict) else data
if REGION:
    jobs = [job for job in jobs if job.get("region") == REGION]


def midpoint(job: dict) -> float | None:
    """salary is [min, max] when the posting discloses pay, else null."""
    pay = job.get("salary")
    if not pay or None in pay:
        return None
    return (pay[0] + pay[1]) / 2


paid = [(job, mid) for job in jobs if (mid := midpoint(job)) is not None]
print(f"{len(jobs)} postings, {len(paid)} with pay ({len(paid) / max(len(jobs), 1):.1%})")
print(f"median base midpoint: ${median(m for _, m in paid):,.0f}")

groups: dict[str, list[float]] = {}
for job, mid in paid:
    groups.setdefault(job.get(FIELD) or "unknown", []).append(mid)
for name, mids in sorted(groups.items(), key=lambda kv: -median(kv[1])):
    print(f"{name:<32} n={len(mids):<4} median ${median(mids):,.0f}")

What the postings say you will own

Postings are the most specific public description of the job. OpenAI's FDE posting for San Francisco, published September 11, 2026, says FDEs “own discovery, technical scoping, system design, build, and production rollout” and measures success through “production adoption, measurable workflow impact, and eval-driven feedback that changes product and model roadmaps”. It asks for five or more years including customer-facing work and travel up to 50%, pays $185K–$300K plus equity, and describes no interview steps. Its April 2025 predecessor paid $220K–$280K for four or more years; the band nearly doubled in width while its midpoint fell $7,500, which we read as one title stretched over more levels.

The same posting spells out the behavioral bar: “Scope work, sequence delivery, and remove blockers early”; “Make trade-offs between scope, speed, and quality; adjust plans to protect delivery”; “Spot risks early and adjust without slowing down”; “Model calm and judgment when the stakes are high”. Each phrase maps to a round. Scoping is the customer or decomposition case (Chapter 13), build is the design and coding rounds, adoption and eval-driven feedback are the evaluation questions of Chapters 9 and 10, and calm judgment is the behavioral round (Chapter 14).

Anthropic's London FDE posting asks FDEs to “Deliver technical artifacts for customers like MCP servers, sub-agents, and agent skills that will be used in production workflows” and to “Travel frequently (25-50%)”, for £225,000–£255,000. Its text asks for four or more years while its application form asks about eight. On October 7 Anthropic's FDE individual-contributor roles were open only in London, Paris and Munich, and US forward-deployed delivery was posted as Technical Deployment Lead. The Harness Engineering book covers the agent loops, tools and context behind those deliverables.

Most employers pair the engineer with a role that does not write production code. Anthropic's Technical Deployment Lead owns “value measurement and ROI” and states the split: “You won't write production code, but you will own the technical direction of engagements alongside FDEs.” OpenAI's Deployment Lead is measured on “delivery reliability (milestones hit, low reopen/churn), operating leverage (patterns reused across deployments), judgment under pressure, and product impact”, and Cursor staffs “a pod of one Strategist and 1-2 FDEs”. Cursor's posting also tells FDEs to “Measure the outcome that matters (revenue, cost, hours saved, error rate, cycle time, escaped defects), not seat adoption”. Know which half of the pair you are interviewing for.

Employer and titlePosted base payTravelExperience asked
Palantir, Forward Deployed Software Engineer$135,000–$200,000up to 25%1+ years
Palantir, Deployment Strategist$110,000–$170,00025–75%
OpenAI, Forward Deployed Engineer$185K–$300K plus equityup to 50%5+ years
OpenAI, Deployment Lead (FDE)$230K–$294Koften 25–50%7+ years
Anthropic, Forward Deployed Engineer (London)£225,000–£255,00025–50%4+ years (form asks 8+)
Anthropic, Technical Deployment Lead (US)$275,000–$380,00025–50%
Databricks, FDE and Sr. FDE$152,900–$210,155 and $182,000–$250,208AI FDE team: every 4–8 weeks6+ years (Sr.)
ServiceNow, Applied AI FDE$201,300–$352,300up to 30%10+ years
Ramp, Software Engineer, Forward Deployed$189K–$330K3+ years
Sierra, Software Engineer, Agent$180K–$390K

Ranges and travel as posted on each employer's public job board on October 7, 2026; blank cells were not in our read of the posting. Assuming a 46-week working year, 25–50% travel means 11.5 to 23 weeks on site, and a Palantir deployment strategist at 75% spends about 34.5.

Related question · Motivation

The role means 25 to 50 percent travel and long stretches inside a customer's building. Why does that work for you?

What it tests: Whether you know the travel load in weeks, have a reason the field suits you, and will still be there after the first engagement.

A strong answer

Because the job is done where the data and the users are, and that is where I do my best work. At 25 to 50 percent travel I expect 11 to 23 weeks a year on site, and I have planned for that. On my last project I spent two days a week at the customer's operations center, and the fixes that mattered came from watching people use the system. Pragmatic Engineer describes OpenAI's FDEs spending a couple of days on site for scoping and a few days a week during delivery, which is a rhythm that suits me. I would ask how you plan travel when two engagements overlap.

Follow-ups

What if the engagement needs you on site four days a week for a year?
Qureshi did exactly that at Airbus. Give your real limit and what would make it work, such as a rotation or an end date; “anything is fine” tells the interviewer you have not thought about it.
Would you take a version of the role with less travel?
Say what you would give up. In Bloomberry's data the sales-engineer-like FDE roles travel under 20 percent and code 30 to 40 percent of the time; the builder type travels 30 to 50 percent and codes 70 to 90 percent.

OpenAI grades its FDEs partly on feedback that changes roadmaps, so expect to be asked how you produce it.

Related question · The role

How do you get what you learn at customers back to the product and research teams?

What it tests: Whether you treat field signal as part of the job, with a mechanism, a cadence and evidence attached.

A strong answer

I write it down in a form the product team can act on. Each deployment produces three kinds of signal: failure cases with the eval that shows them, requests that repeat across customers, and workarounds I had to build. I keep a field log per engagement and send a short monthly readout with the top issues, how many customers hit each, and the eval results. OpenAI runs this as a system, according to Pragmatic Engineer, with roughly bi-weekly research sessions, fortnightly readouts with the Head of Product, an “FDE Field notes” Slack channel and quarterly bootcamps. I send failures as evals, because that is the format a research team can use.

Follow-ups

Give an example of field feedback that changed a roadmap.
Use your own, with numbers: how many customers hit it, what the eval showed, what shipped. Pragmatic Engineer reports that OpenAI's FDE and solutions teams contribute to the Agents SDK, which replaced Swarm; that is the kind of outcome to bring.
What if the product team disagrees with your read?
Bring the eval set and ask what evidence would change their mind. In Palantir's model, Deltas validate larger requests against the roadmap: the FDE makes the case, and the product team makes the call.

Three jobs under one title, and the critique

The next question is not on the list of 33, but this section's evidence is its answer, and it is the worked example that the section on grading takes apart.

Related question · The role

What is a forward deployed engineer, and how is the role different from a solutions architect, a sales engineer or a consultant?

What it tests: Whether you define the role by what the engineer owns after the sale, back the definition with evidence, and admit where the lines blur.

A strong answer, in about two minutes

An FDE owns a customer deployment from discovery to production, writes the code on the customer's systems, and is judged by whether people use what ships. OpenAI's posting puts it in one line: FDEs “own discovery, technical scoping, system design, build, and production rollout”, and success is “production adoption, measurable workflow impact, and eval-driven feedback”.

The neighboring roles differ on ownership and code. A solutions architect designs and prototypes around the sale; Pragmatic Engineer reports that OpenAI's solutions architects “rarely write code on customers' infrastructure”, while its FDEs do. A sales engineer supports a deal: in Bloomberry's analysis of about 1,000 FDE postings, even the sales-engineer-like FDE roles code only 30 to 40 percent of the time and pay commission, while the builder type codes 70 to 90 percent and no posting mentions a quota. A consultant, in a Palantir engineer's words, creates “a one-time analysis, recommendation, or solution to a specific problem”; a Delta measures success by impact on the customer's goal and sends what it learns back into one product.

The trade-off is that the title is used loosely. Bloomberry finds three different jobs under it, and Gergely Orosz expects FDE work to become indistinguishable from solutions architecture or consulting. So when I look at an FDE role I ask two things: is my time billed by the hour, and does field feedback have a path into the roadmap? I would ask you both about this team.

The evidence starts with Bloomberry's analysis, which finds three jobs under the title: a Builder FDE (60% of postings, 70–90% coding, 30–50% travel, $140–250K), a “Sales Engineer+” (30%, 30–40% coding, under 20% travel, $120–200K plus commission) and an internal tools or go-to-market engineer (10%, under 5% travel, $100–180K). Across all three, 70% of postings mention equity, 8% on-target earnings and none a quota.

Estimates of time spent coding disagree. Gergely Orosz's opinion is “~25% coding-related, 50% integration/plumbing, 25% meetings and customer hand-holding”; a survey by Perspective AI, a vendor that sells an AI interviewer, claims 31% code by an unaudited method; Bloomberry's builder type codes 70–90% by posting text. Quote the one that matches the role in front of you.

The critique is that the title is a rebrand. Orosz wrote in May 2026 that the role “is about to become indistinguishable from a solutions architect or consultant”, and summed it up as “You are a contractor who codes at a customer's office.” a16z's Tom Hollands writes that in 2011 Palantir took “their solutions engineers and integration engineers ... and gave them a new title”, which he calls title arbitrage. Chris Degnan, Snowflake's former chief revenue officer, called the FDE a “glorified professional services person” who will leave “a lot of technical debt”, and the top comment in a Hacker News thread called the role “just a title change for Solutions Architects”. A reply drew the line at building: SAs “provide guidance. FDEs actually build things alongside the customer engineers.”

The employers' answer rests on where the code runs and how the work is paid. Per Pragmatic Engineer, OpenAI's solutions architects “rarely write code on customers' infrastructure” while its FDEs “write code directly on customer infrastructure”, and Colin Jarvis, who heads the FDE team, said the team “avoids ‘services revenue’”, as Business Insider reported. AWS says its deployments are “structured around shared goals and business results, not billable hours”. Databricks is the counterexample: its postings say “FDEs are billable”. Ask which model you would be joining.

The coding-share argument also misses where engagements spend their time. In Qureshi's pilots it was data access. At Morgan Stanley, per the same Business Insider report, “technical scaffolding took six to eight weeks”, and pilots, evaluations and iteration with wealth advisors took “another four months”, so the build was a quarter to a third of the elapsed time. The rest, evaluation and trust, is the part a solutions architect hands off and an FDE owns.

Follow-ups

Is an FDE just a solutions architect with a better title?
Sometimes, and it is fine to say so. The test is ownership: who writes and runs the production code, and who answers for adoption. If the customer's team builds from your slides, it is a solutions role whatever the title says.
Is the FDE team a cost center?
At Palantir it began as an investment, with pilots “at no or low cost” to customers. a16z's Joe Schmidt argues for incentives “that allow services to be sold at cost”, while Databricks bills FDE time. Know which model the company runs, and say how you would make your work reusable under either.
If models keep improving, why do labs need FDEs at all?
Because most of a deployment's effort sits in integration, data access and evaluation. Eddie Siegel, CTO of Ode, told TechCrunch that “model selection matters, but it's not where the majority of calories are spent”.

Red flags

  • Defining the role by its title, or reciting “one customer, many capabilities” without saying what the engineer owns.
  • Claiming FDEs code all day: the estimates run from about 25% to 90% depending on the job type and who is counting.
  • Talking down solutions architects or consultants; the interviewer may have been one.
  • Ignoring the rebrand critique instead of saying how you would tell an FDE role from a relabeled services job.

The loop, from published pages and candidate reports

Few companies publish their FDE loop, and what they publish fits in one figure. Everything else comes from candidates and prep sites, and the two often get mixed.

OpenAI's interview guide covers all roles: résumé review (typically about a week), introductory calls, a skills-based assessment (“pair coding interviews, take-home projects, technical tests, etc.”), final interviews and a decision “within one week of your final interviews”. Finals are “4–6 hours of final interviews with 4–6 people over 1–2 days”, about 60 minutes per interviewer and virtual by default, and the published cadence alone implies at least four weeks of waiting. The engineering criteria are “well-designed solutions to the challenge, high-quality code, optimal performance, and good test coverage”, plus communication and collaboration, and the guide says “We are not credential-driven”.

Palantir's Getting Hired pages say interviews are “personalized”, that “Everyone starts with one or two phone interviews” of 20 to 45 minutes, and that the process rests on six competency guides: The Phone Interview, Writing Good Code, Analyzing the Efficiency of Code, Navigating Open-Ended Questions, Solving Technical Problems, and Working Inside Existing Systems. The open-ended guide describes problems “with multiple possible solutions that will incur different trade-offs” and advises “Deliver a functioning idea first, then expand it afterwards.” None of the six guides uses the word “decomposition”. The “decomp” round candidates describe maps onto Navigating Open-Ended Questions, and the “re-engineering” round onto Working Inside Existing Systems.

Intuit publishes an engineer timetable that totals 255 to 285 minutes of the candidate's time, with feedback “within 24 hours”. Databricks prints its Solutions Architect loop in the posting, ending in a “Build, Demo, Pitch!” presentation, and its careers page describes four to six onsite interviews over a two-to-three-month process. Snap runs a recruiter screen, an initial interview with a values question, a virtual loop on HackerRank that includes AI-assisted coding, and an offer. Sierra's April 2026 loop is in the next section.

Pick a company to see its published stages, the minutes it gives for each and its rule on AI in that round. PREP SITES shows the composite that prep sites describe, drawn dashed because it is secondary.

That composite comes from Aced (formerly Exponent), which describes a typical FDE loop of 5–8 stages over 3–6 weeks, with a decomposition case it calls “the hardest round” and a 4–8 hour take-home at some companies. Its Anthropic page calls the customer-conversation simulation “reportedly the decisive round”. Pass rates circulate too, such as about 40% for Palantir's decomposition round; they appear only on prep sites that cite each other, and we could not confirm any of them.

Candidate reports add detail, one person at a time. An OpenAI FDE candidate on Aced's experiences page described a one-week take-home (“a semantic search setup over Amazon products for ChatGPT”) and an “AI-enabled coding screen” that was “basically an easy LeetCode problem”, and wrote “The weirdest part was that the take home was basically the job.” A Palantir new-grad candidate on Reddit described two 45-minute decomposition rounds (assigning hospital rooms; costing a warehouse delivery network as a min-cost flow) in which the requirements changed mid-conversation. Snorkel AI's FDE screen, as described to one candidate, handed over a client brief, a ground-truth CSV and a set of model outputs, and asked what to measure.

Job descriptions rarely help: of 1,765 that Grigorev analyzed, about 80 (4.5%) describe the interview process. Ask the recruiter for the stages, their length and the AI rule for each round.

AI in the interview room: which rounds allow it

The largest change to technical interviews in 2025 and 2026 is the rule on AI tools, and it is not settled. Policies split three ways (banned unless stated, decided per format, and built in), and one company can apply different rules at different stages.

CompanyRule on AI toolsSinceSource
AnthropicNone in live interviews or take-homes unless stated; use Claude to prepareJuly 10, 2025Candidate AI guidance
OpenAIPer format: some interviews “intentionally allow” AI, others do not; prep materials say whichLive October 2026Interview guide
MetaBuilt-in assistant that candidates “are expected to use”; no outside toolsAI-enabled coding from October 2025 (secondary)Careers FAQ
Canva“AI-Assisted Coding” replaced “Computer Science Fundamentals”June 2025Engineering blog
SierraCoding rounds replaced by a 2-hour build with any AI toolsApril 22, 2026Company blog
SnapNone unless told; some rounds are AI-assisted codingLive October 2026Careers pages
AmazonNo GenAI unless permitted; violations “may result in disqualification”February 2025Business Insider (secondary)
CursorNone “other than autocomplete” in first technical screensJune 2025Business Insider (secondary)
Accenture, GenAI ArchitectYour own environment, “including any AI-assisted tools”Posting indexed June 2026Posting
GoogleAt least one in-person round; Gemini allowed in a code-comprehension pilot from the second half of 20262025 and 2026HC Magazine; Business Insider via an aggregator (secondary)

Surveys show the same split. CoderPad's State of Tech Hiring 2026 (about 650 respondents) found that 34% ban AI in technical interviews, 46% allow it and 20% decide case by case, and where it is allowed, 66% prioritize “catch & fix AI mistakes”. The UnchartedCareer tracker of 20 published policies, updated August 6, 2026, counts 10 bans, one that discourages AI, four that allow it, four that require it and one with no published rule. Meta's AI-enabled round is scored on problem solving, code quality, verification and communication, according to Hello Interview.

The case for bans rests on cheating data. In a January 2024 experiment, interviewing.io had candidates secretly use ChatGPT in 32 analyzed interviews: they passed 73% of verbatim LeetCode questions (11 interviews), 67% of modified ones (9) and 25% of custom questions (12), against a 53% platform average, and no interviewer reported suspecting cheating. Its 2025 survey of 67 interviewers found that 81% suspected AI cheating, and Gartner found that 6% of 3,000 job candidates surveyed in the second quarter of 2025 admitted interview fraud.

The figure draws each pass rate with its 95% Wilson interval; drag the slider to see how the intervals would narrow with more interviews at the same rates. As run, the verbatim and modified intervals (43.4–90.3% and 35.4–87.9%) cannot be told apart, and the custom arm (8.9–53.2%) overlaps the verbatim one by about ten points. At twice the interviews verbatim and custom separate; at ten times, verbatim and modified still overlap. Non-overlap is a stricter test than comparing two arms directly, which is how the post's own significance note can single out the custom arm. The practical finding survives the small samples. AI help stopped working on custom questions, and candidates report multi-part coding problems that add requirements as they go at OpenAI, Palantir and Anthropic.

The intervals come from wilson.py, which Chapter 9 reuses for eval pass rates; the Evaluating Agentic AI book covers small-sample statistics in depth.

wilson.py
python
"""Wilson score intervals for small-sample pass rates.

Python 3.11+, standard library only. Chapter 9 reuses wilson() for evals.
"""
from math import sqrt
from statistics import NormalDist


def wilson(passed: int, n: int, confidence: float = 0.95) -> tuple[float, float]:
    """Interval for passed/n that stays inside [0, 1] and behaves for small n."""
    if n == 0:
        return (0.0, 1.0)
    z = NormalDist().inv_cdf(1 - (1 - confidence) / 2)
    p = passed / n
    denom = 1 + z * z / n
    center = (p + z * z / (2 * n)) / denom
    half = (z / denom) * sqrt(p * (1 - p) / n + z * z / (4 * n * n))
    return (max(0.0, center - half), min(1.0, center + half))


def gap(a: tuple[float, float], b: tuple[float, float]) -> float:
    """Distance between two intervals; negative when they overlap."""
    return max(a[0], b[0]) - min(a[1], b[1])


# interviewing.io, January 2024: passes rebuilt from the published rates and n
ARMS = {"verbatim": (8, 11), "modified": (6, 9), "custom": (3, 12)}

if __name__ == "__main__":
    for name, (k, n) in ARMS.items():
        lo, hi = wilson(k, n)
        print(f"{name:<9} {k:>2}/{n:<2} = {k / n:6.1%}  95% CI [{lo:.1%}, {hi:.1%}]")
    for scale in (1, 2, 5, 10):  # the figure's slider: same rates, more interviews
        v, m, c = (wilson(k * scale, n * scale) for k, n in ARMS.values())
        print(f"x{scale:<2} verbatim vs custom {gap(v, c) * 100:+.1f} pts, "
              f"verbatim vs modified {gap(v, m) * 100:+.1f} pts")

Anthropic shows the tension from both sides. It bans AI in live interviews by default, yet its performance take-home, completed by more than 1,000 candidates since early 2024, allowed AI tools. In January 2026 Tristan Hume described how the models caught up until Claude Opus 4.5 matched the best human result within two hours, and when Anthropic open-sourced the old take-home, none of the first day's submissions below 1,300 cycles was valid: “In each case, a language model modified the tests to make the problem easier.”

Sierra went furthest. Its April 22, 2026 post says “We removed our coding and algorithms interviews and replaced them with an AI-native onsite.” The candidate plans a product with the interviewer, builds it alone “over 2 hours, using the AI tooling and frameworks of their choice”, then reviews the demo, the code, the path to production and how AI was used.

Related question · AI-native round

You have two hours and any AI tools. Plan a small product with the interviewer, build it alone, then demo it and defend the code and its path to production.

What it tests: Scoping under a clock, staying in control of AI-written code, and judgment about what production would need.

A strong answer

In Plan I narrow to one user and one workflow and write down what done means: two or three behaviors I will demo and the checks that prove them. In Build the agent writes scaffolding and boilerplate, and I write or review line by line everything on the critical path. I run the tests after every agent change and read every diff to a test file, because a model that edits tests to pass them is the failure Anthropic saw in its own take-home. With half an hour left I stop adding features and harden what exists. In Review I demo the path that works, show where I overrode the AI, and list what production needs: authentication, data access, observability, evals in CI and a rollback.

Follow-ups

Why did you accept that block of AI-written code?
Say what you checked: the test you ran, the edge case you tried, the diff you read. Meta's AI-enabled round scores verification explicitly, according to Hello Interview.
At two hours the build is half done. How do you run the review?
Demo what works, say what does not, and spend the time on the path to production. Sierra says it is “hiring for strengths, not just an absence of weakness”.

Sierra's review asks how the candidate used AI, and Snap says AI fluency “may be evaluated”, so expect the next question wherever AI is allowed. Answer with your own project; the shape transfers.

Related question · AI fluency

How did you use AI tools on your last project, and how did you check what they produced?

What it tests: Whether you use AI deliberately, know where it fails, and verify its output with evidence you can describe.

A strong answer

On my last project, an integration between a ticketing system and a model API, a coding agent wrote the scaffolding, tests and SDK boilerplate, and I wrote the code that touched customer data. Every agent change ran the unit tests and a small eval set in CI, I read each diff before merging, and any change to a test file needed a stated reason. It failed me twice: it invented a field missing from the customer's API version, and it fixed a flaky test by loosening the assertion. The checks caught both, and since then I give the agent the exact API version and schema up front.

Follow-ups

Where would you not use an AI tool?
Wherever you cannot verify the output cheaply: security-sensitive code without review, migrations against production data, and anything the customer's policy forbids.
How do you know the tool made you faster?
Measure your own work before and after, such as time to a passing build. CodeSignal's March 2026 vendor survey reports that 91% of 450 US engineers use agentic coding tools, which measures adoption; speed needs your own numbers.
Ask before every loop
Ask the recruiter, in writing, which rounds allow AI, which tools (the platform's built-in assistant or your own), and whether any take-home allows it. When AI is allowed, say what you ask it and check what it returns, out loud. When it is not, keep it closed: Amazon's guidelines warn that violations “may result in disqualification”, and Meta authorizes only the assistant inside its interview environment.

Where the 33 questions came from

The candidate who posted the 33 questions grouped them into seven themes: inference, RAG, agents and orchestration, guardrails and responsible AI, evaluation and observability, coding, and system design with FDE cases. No company published these questions, so we traced each one before writing about it.

Most of the list matches prep-site text. In our comparison with two pages from Aced (formerly Exponent), its definitive 2026 FDE guide and its 45+ AI engineer interview questions, 13 questions match near-verbatim, 8 match the same scenario in different wording, and 12 have no match: 21 of 33, or 64%. The FDE guide was online by May 28, 2026, when the University of Kentucky's career center reposted it (Harvard's FAS career center followed on August 1), and it contains 14 of the 33, among them the per-user and global rate limiter and the slipped deployment with a CTO on the line.

The overlap shows where the wording came from and says nothing about whether interviewers ask the questions. Aced says every question on its list “is pulled from real interviews candidates reported to us”, which we cannot check because its question pages put company statistics behind a paywall. The list counts as one source family with company names attached, and a prep site's “asked at OpenAI” tag is a claim to verify. The copying runs through GitHub too: one 2026 company-wise repository matches IGotAnOffer's Anthropic and OpenAI lists almost word for word.

Corroboration varies widely. The best-corroborated question is batching LLM requests on one GPU while users wait, reported by at least six source families. An Anthropic infrastructure engineer's loop report from February 2026 describes a system design round on an inference API with variable-length requests, GPU memory, a priority queue, streaming, dynamic batching, the KV cache and autoscaling, and says “if you only prepare for one thing at Anthropic make it this.” Chapter 2 answers it, and our post on why LLM inference is slow covers the background. At the other end, the AI recruiter (Question 15) and body-camera reports (Question 30) have no interview source, and micro versus macro F1 (Question 21) appears in none of the 25 question collections we downloaded and in no candidate report. We answer those three as likely design scenarios and label them that way.

Topic counts across the 25 collections, which are regex matches over README text and so lower bounds, show where the public supply concentrates: observability or tracing in 17, chunking and hallucination in 15, the KV cache, MCP, reranking and quantization in 13, LLM-as-judge in 12, prompt injection in 11, Kafka and voice agents in 3, HIPAA in 2, and micro versus macro F1 in none. Weight your preparation by corroboration and by your target company's postings, and expect follow-ups that no list contains: OpenAI says its finals are designed to “stretch you beyond your comfort zone”.

The competency map and the question index

The seven themes map onto rounds of the loop and onto chapters of this book. The usual round for each theme is our reading of candidate reports and prep pages, since no company publishes which question goes in which round.

Each square is a question, with its chapter underneath. White squares match an Aced page near-verbatim, hatched ones match its scenario, solid black ones come from other sources, and dashed ones have no interview source. Filter with the buttons, then hover over or tab to a square for its chapter, usual round and the number of source families that report it.

The near-verbatim matches cluster in inference, coding and the design cases. The questions with no Aced match are mostly deeper technical ones, such as the KV cache, autoscaling, document versioning, LangGraph, MCP and Kafka. Evaluation runs through every round, and Aced calls “How do you know it's working?” “the differentiator question”.

The question index, by chapter:

  • Chapter 2 (Batching and the KV Cache): batching requests on one GPU while users wait (1) and why the KV cache limits concurrency (2).
  • Chapter 3 (Scaling, Compressing and Diagnosing a Model Server): autoscaling inference on Kubernetes (3), distillation versus quantization (4) and diagnosing high latency (5).
  • Chapter 4 (Retrieval from First Principles to 50M Patient Records): ingesting large tables (7), prompting versus RAG versus fine-tuning (8) and RAG over 50M patient records under HIPAA (28).
  • Chapter 5 (Structuring Agent Work): a claims agent on a token budget (9), a LangGraph summarizer (11) and multi-agent orchestration (12).
  • Chapter 6 (Customer-Facing Agents): a support agent with tools, memory and human handoff (10) and a voice calling agent (14).
  • Chapter 7 (Securing the Action Surface): MCP with authentication, RBAC and schema-validated calls (13) and guardrails (16).
  • Chapter 8 (Regulated Decisions): an AI recruiter for sales hiring (15) and responsible AI for regulated decisions (17).
  • Chapter 9 (Evaluation: Proving It Works): how you know a system works (18), LLM judges (19) and micro versus macro F1 (21).
  • Chapter 10 (Operating in Production): hallucination and repetition (20), regressions and rollback (22) and agent observability (23).
  • Chapter 11 (Data Pipelines): document versioning (6) and Kafka into an AI pipeline (27).
  • Chapter 12 (The Coding Round): a rate limiter (24), async retries (25) and a RAG pipeline over a folder (26).
  • Chapter 13 (The Case and Design Round): search from 1.5 s to under 100 ms (29), body-camera audio to a reviewed report (30) and fraud detection across three acquisitions (31).
  • Chapter 14 (The Behavioral Round): a system you owned (32) and a three-week slip told to the CTO (33).

How an answer gets graded

Companies that publish how they interview describe structured interviews scored against rubrics, and the selection research backs them. In Sackett, Zhang, Berry and Lievens's 2022 re-analysis, structured interviews have the highest mean operational validity of the common predictors, .42, against .19 for unstructured interviews, so they explain about 4.9 times as much of the variance in job performance (the squares are .176 and .036). Their 80% credibility interval runs from .18 to .66, which means validity depends on how the interview is run.

The published processes have the same parts. Google's re:Work guide, updated in March 2026, defines structured interviewing by vetted questions, recorded feedback, standardized rubrics (outstanding, solid, borderline, poor) and trained, calibrated interviewers. Stripe's Atlas guide publishes a 1-to-4 communication rubric in which a 1 “talks exclusively about things other than their personal work and contributions” and a 4 conveys “their personal trajectory and impact”. Amazon says answers “should include metrics or data where applicable”.

Every model answer in this book follows one shape, built for rubrics like these:

  1. Claim. Answer the question in the first sentence, in your own words.
  2. Number. Give the figure that makes the claim checkable, with its source or its arithmetic: a memory budget, a pass rate with its interval, a timeline. It is also what the interviewer will probe.
  3. Trade-off. Name what your choice costs and when you would choose differently. Palantir defines its open-ended competency by solutions that “will incur different trade-offs”, and OpenAI asks FDEs to trade off “scope, speed, and quality”.
  4. Next step. End on what you would do or ask next, the instinct behind Palantir's “Deliver a functioning idea first, then expand it afterwards”.

The FDE answer earlier in this chapter follows the shape. Its first sentence is the claim, that an FDE owns a deployment from discovery to production and is judged by adoption. The numbers are Bloomberry's 70–90% coding for builders against 30–40% for the sales-engineer type. The trade-off is the loose title, with the critique named. The next step is the pair of questions turned back to the interviewer. Spoken, it takes about two minutes, inside the two to three that CMU's Tepper guide recommends for a STAR story.

Motivation answers are graded the same way, and they can decide an outcome: a Palantir FDSE candidate reported being asked “Why Palantir / Why Government / Why FDSE” in round after round and being rejected on mission alignment.

Related question · Motivation

Why forward deployed engineering instead of a product engineering role?

What it tests: Whether your motivation fits the job as it is (customers, travel, ambiguity) and rests on something you have done.

A strong answer

I want to be accountable for whether a system gets used, and in this role that accountability sits with the FDE. On my last project the changes that moved adoption came from sitting with the operations team twice a week, and that was the work I did best. FDE work is that loop full time: discovery, a build on the customer's systems, and a deployment judged by production adoption, which is how your posting measures it. I know the cost, which is depth in one codebase and a predictable week. I also know the critique that some FDE roles are consulting with a better title, so I looked for a team whose field feedback reaches the product.

The same answer at three levels

The model answer above is pitched here: a deployment you owned, a strength on the customer side, the cost you accept, and the critique with a test for it. Expect “tell me about a deployment that was not adopted” next, and have one ready.

Follow-ups

What would you miss about product engineering?
Name it: depth in one codebase and long ownership of one system. Then say how you would keep some of it, for example by contributing fixes from the field to the core product, as Palantir's Deltas describe doing.
Why our FDE team rather than another company's?
Quote what this team's posting measures and connect it to something you have done. OpenAI grades adoption and eval-driven feedback; Cursor asks for outcomes such as hours saved instead of seat adoption.

Chapter 1 in one page

Key takeaways
8 items
  • 1Palantir's 2019 definitions still frame the job: a Dev ships one capability to many customers, a Delta serves one customer with many capabilities, and the core Delta skill is technical decomposition.
  • 2OpenAI's 2026 posting defines the FDE by ownership from discovery to production rollout and grades it on production adoption, workflow impact and eval-driven feedback; most employers pair the FDE with a lead who writes no production code.
  • 3About $9.0 billion was pledged to FDE vehicles between May and July 2026 (OpenAI more than $4B, Anthropic's services company $1.5B as reported, AWS $1B, Microsoft $2.5B), and two of the four figures carry caveats.
  • 4Quote market numbers with their definitions: Lightcast's 5,200-plus US postings in seven months is a flow, Plank's 1,333 live postings is a stock, and Indeed's 5,330 is an index level.
  • 5One title covers three jobs (builder, Sales Engineer+, internal tools), and estimates of coding time run from about 25% to 90%; ask which one you are interviewing for.
  • 6Companies publish little about their loops beyond OpenAI's stages and 4–6 hour finals, Palantir's six competencies, Intuit's timetable, Databricks' loop string and Sierra's Plan, Build, Review; round names such as decomp, and every pass rate, come from prep sites.
  • 7AI rules split three ways (banned unless stated, per format, built in), so ask the recruiter about each round, and narrate and verify when AI is allowed.
  • 821 of the 33 questions match two prep pages and batching on one GPU is the best corroborated; every answer in this book follows one shape of claim, number, trade-off and next step.

What to do on Monday: pick the company you are interviewing with and run fde_postings.py against its board. Read its FDE posting line by line and map each responsibility to a round, the way the postings section does for OpenAI. Email the recruiter three questions: the stages and their length, which rounds allow AI and with which tools, and whether there is a take-home. Then record yourself answering “What is an FDE, and how is it different from a solutions architect?” in under three minutes, and check the recording against the four parts: claim, number, trade-off, next step.

Chapter 2 · Inference I

Batching and the KV Cache

Requests joining a continuous batch on one GPU, with the KV cache below

Why one decode step costs nearly the same for one user or forty until their cache fills memory, why memory rather than compute caps how many fit, and the arithmetic that turns a config.json into a concurrency limit an interviewer can check.

Batching requests on one GPU is the best-corroborated question in this book's bank. Aced lists it first among its AI engineer questions and attributes it to Anthropic, IGotAnOffer lists two variants, and an Anthropic infrastructure candidate described a design round that covered dynamic batching, the KV cache and a priority queue. Its companion, what the KV cache is and why it limits concurrency, appears in 13 of the 25 question collections the researchers downloaded. These are prep-site compilations and candidates' accounts, and no 2025 or 2026 report the researchers found names a concrete KV-sizing or vLLM-tuning question at OpenAI, Anthropic, NVIDIA, Databricks or Cohere; Chapter 1 explains how we weigh such sources.

The two questions share one answer. Decode is limited by memory bandwidth, so batching is nearly free; the batch you can run is limited by the cache you can hold, so memory sets concurrency; and the scheduler keeps that cache full of useful work without making anyone wait long for the next token. Forward deployed engineers meet this on customer hardware, and NVIDIA's posting for a Senior Solutions Architect, AI Inference asks for "Deep knowledge of modern inference best practices including disaggregated serving, multi-tier KV cache management, speculative decoding, quantization and custom inference kernels" and names forward deployed engineering as relevant background. The reference engine in this chapter is vLLM v0.31.0, released 2026-10-05; Chapter 3 continues with autoscaling, quantization and latency diagnosis.

Two phases, two bottlenecks

Every generation request runs in two phases. Prefill pushes the whole prompt through the model in one forward pass: each weight loaded from memory is multiplied against thousands of token vectors, so the tensor cores set the pace. A useful rule of thumb puts prefill at about 2 × parameters × prompt tokens FLOPs, plus attention terms that grow with the square of the prompt length.

For a 70B model and an 8,192-token prompt that is 1.15 × 1015 FLOPs. One H100 at its dense FP8 rate of 1,979 TFLOPS and 50% utilization needs about 1.16 s before the first token; eight H100s need about 0.15 s and one B200 about 0.51 s (the book's arithmetic on NVIDIA's spec sheets). MLPerf's interactive Llama 2 70B scenario allows 450 ms of time to first token at p99, so a long prompt on a single GPU misses the target before it has queued behind anyone.

Decode then produces one token per sequence per step. Each step reads every weight, plus the cache of every sequence, from high-bandwidth memory, and does only a couple of FLOPs per byte it reads. Bandwidth sets the pace, and dividing bytes by bandwidth gives a floor. Llama 3.3 70B in FP8 is about 70 GB; an H200 streams 4.8 TB/s; so no step finishes in under about 14.6 ms, about 69 tokens per second for a single user. A B200 at 8 TB/s brings the floor to about 8.8 ms, and Llama 3.1 8B in BF16 on an H100 sits at about 4.8 ms. Databricks gives the measured counterpart: a 7B FP16 model at 14 ms per token moves about 1 TB/s, 50% model bandwidth utilization on 2 TB/s hardware.

That floor is why batching works. The weights cost the same 14.6 ms whether a step serves one sequence or forty, so each added sequence brings only its own cache reads and a sliver of compute. Anyscale's engineers put it in one line in 2023: "LLM inference throughput is largely determined by how large a batch you can fit into high-bandwidth GPU memory." The roofline behind the split is the subject of The Real Reason LLM Inference Is Slow, and the memory hierarchy underneath it of GPU Architecture and Programming.

The figure computes one decode step of Llama 3.3 70B with FP8 weights on one H200 from those numbers, 320 KiB of BF16 cache per token (160 KiB in FP8), compute at two FLOPs per weight per sequence against 1,979 dense FP8 TFLOPS at 50% utilization, and the 54.7 GB cache budget of Section 7. Attention FLOPs and kernel overheads are left out, so every time is a floor.

Drag BATCH from 1 to 20 at 8K: the step floor rises from 15.1 ms to 25.8 ms, per-user speed falls from 66 to 39 tokens per second, and total throughput climbs from 66 to 776. At 21 the cache no longer fits. Switch to FP8, set CONTEXT to 1K and push BATCH past about 140, and the compute bar overtakes memory: that is where decode stops being bandwidth-bound. The dashed line is MLPerf's 40 ms interactive limit on time per output token. That limit, and the 15 ms MLPerf sets for gpt-oss-120b and DeepSeek-R1 in their interactive scenarios, sit only one to three times above these floors (the book's arithmetic), which is why the room for batching shrinks as interactivity targets tighten.

The two phases set the two latency numbers users feel: time to first token (TTFT), which covers queueing, prefill and the network, and inter-token latency (ITL), the gap between streamed tokens that decode sets. Benchmarking tools define them slightly differently, and Chapter 3 pins the definitions down.

Related question · Inference

Why is prefill compute-bound while decode is bound by memory bandwidth, and how do TTFT and inter-token latency trade against each other?

What it tests: Whether you can explain the two phases by how much work each byte read buys, and tie them to the two latency numbers users feel.

A strong answer, in one minute

Prefill runs the whole prompt through the model in one pass, so each weight read from memory is used for thousands of tokens: matrix-matrix work that saturates the tensor cores. Decode produces one token per sequence per step, so each weight read serves only the sequences in the batch, and the step time is set by how fast HBM streams the weights and the cache. TTFT is mostly queueing plus prefill; inter-token latency is the decode step. They trade through the scheduler. Giving a long prefill a whole step improves its TTFT and stalls every running user's next token; splitting it into chunks does the reverse.

Follow-ups

How long is prefill for an 8K prompt on a 70B model?
About 1.15e15 FLOPs at two FLOPs per parameter per token. One H100 at 1,979 dense FP8 TFLOPS and 50% utilization needs about 1.16 s before any queueing, already past a 450 ms TTFT target; eight H100s bring it to about 0.15 s.
What does a faster memory system buy you?
A lower decode floor. With about 70 GB of FP8 weights the floor is about 14.6 ms per token on an H200 (4.8 TB/s), about 8.8 ms on a B200 (8 TB/s) and about 3.6 ms on Rubin (19.2 TB/s), by the book's arithmetic. Prefill time moves with FLOPs instead.

Question 1: batching requests on one GPU while users wait

Question 1 · Inference

Design a system that batches LLM requests on a single GPU while every user waits on an open request for the answer.

What it tests: Whether you know the three kinds of batching, size the batch from memory instead of a constant, protect latency with admission control, and can name the settings that trade time to first token against inter-token latency.

A strong answer, in three minutes

I would start with the workload, because it decides the rest: prompt and output length distributions, peak concurrency, and a latency target for time to first token and time per output token. MLPerf's interactive Llama 2 70B limits are a good default to quote, 450 ms and 40 ms at p99.

Batching pays because decode is bandwidth-bound. Every step reads all the weights whether it serves one sequence or forty; for Llama 3.3 70B in FP8 on one H200 that floor is about 14.6 ms, and extra sequences ride along almost free until their KV cache or the compute catches up.

The design has two layers. In front, a bounded admission queue with a deadline per request and a fast rejection when it is full, so overload becomes a 429 instead of a timeout. Inside, an engine that batches continuously: it re-forms the batch every iteration, so a finished sequence frees its slot at once and a waiting request joins on the next step. A fixed batch with a max-wait timer suits embedders and rerankers; for generation it wastes slots on length variance, and Anyscale measured up to 23x over naive static batching on OPT-13B.

Memory is the binding limit, so I size concurrency from the KV cache: paged blocks, chunked prefill under a per-step token budget so one long prompt cannot stall everyone, preemption when blocks run out, and priority scheduling for interactive traffic. Then I would prove it with a load test at realistic burstiness, reporting p50 and p99 TTFT, inter-token latency and goodput.

Where this question comes from
Aced lists a single-GPU batching design, up to 100 inputs per batch with users waiting synchronously, first among its AI engineer questions and attributes it to Anthropic. IGotAnOffer's Anthropic page lists a batched inference system in which 100 requests take as long as one, and its system design guide a review of a junior developer's batching design. One candidate's account of an Anthropic infrastructure loop describes an inference API round covering variable-length requests, GPU memory, a priority queue, streaming, dynamic batching, the KV cache and autoscaling. All are secondary reports, not published company questions.

Frame the workload before drawing boxes. If the interviewer gives no latency target, borrow MLPerf's: MLCommons set the 40 ms limit after finding a median of 20 to 50 tokens per second critical for a seamless chat experience. Ask for the length distributions too, since 500-token and 30,000-token prompts need different schedulers.

The answer prep sites give is a request-level batcher with two triggers. Aced's model answer flushes when the batch reaches its maximum size (100 in the question) or when a wait window of 5 to 20 ms expires, buckets requests by length, pushes back when the queue grows and mentions continuous batching for token generation. That design is right for a model that runs one fixed batch per call: an embedder, a reranker, a classifier, or a generate() that pads a batch and runs it to completion. Here is a version with the parts interviewers probe: two flush triggers, a bounded queue that refuses work instead of growing, a deadline per request, and a batch that fails as a unit without taking the worker down.

micro_batcher.py
python
"""micro_batcher.py: request-level batching for a model that runs one fixed
batch per call (an embedder, a reranker, a classifier, a static generate()).
Python 3.11+, standard library only. In front of vLLM, SGLang or TensorRT-LLM,
keep the bounded queue and the deadlines and drop the wait window: the engine
re-forms its batch every decode iteration.
"""
import asyncio
import time
from collections.abc import Callable, Sequence


class Overloaded(Exception):
    """Raised at admission. Map it to HTTP 429 or 503 with a Retry-After header."""


class MicroBatcher:
    def __init__(self, run_batch: Callable[[Sequence[str]], list[str]], *,
                 max_batch: int = 32, max_wait_ms: float = 10.0, max_queue: int = 256):
        self.run_batch = run_batch          # blocking call that owns the GPU
        self.max_batch = max_batch          # trigger 1: the batch is full
        self.max_wait = max_wait_ms / 1000  # trigger 2: the oldest request has waited enough
        self.queue: asyncio.Queue[tuple[str, asyncio.Future[str]]] = asyncio.Queue(max_queue)
        self.worker: asyncio.Task[None] | None = None

    async def submit(self, item: str, timeout_s: float = 30.0) -> str:
        if self.worker is None:
            self.worker = asyncio.create_task(self._loop())
        fut: asyncio.Future[str] = asyncio.get_running_loop().create_future()
        try:
            self.queue.put_nowait((item, fut))  # backpressure: refuse now, never block
        except asyncio.QueueFull:
            raise Overloaded(f"{self.queue.maxsize} requests already waiting") from None
        return await asyncio.wait_for(fut, timeout_s)  # the caller's own deadline

    async def _loop(self) -> None:
        while True:
            batch = [await self.queue.get()]  # sleep until there is work
            deadline = time.monotonic() + self.max_wait
            while len(batch) < self.max_batch:
                try:
                    async with asyncio.timeout(deadline - time.monotonic()):
                        batch.append(await self.queue.get())
                except TimeoutError:
                    break
            live = [(x, f) for x, f in batch if not f.done()]  # skip callers that gave up
            if not live:
                continue
            try:
                outs = await asyncio.to_thread(self.run_batch, [x for x, _ in live])
            except Exception as exc:  # one bad batch fails its callers, not the server
                for _, f in live:
                    if not f.done():
                        f.set_exception(exc)
                continue
            for (_, f), out in zip(live, outs, strict=True):
                if not f.done():
                    f.set_result(out)


if __name__ == "__main__":
    import random

    def fake_model(batch: Sequence[str]) -> list[str]:  # stand-in for one GPU call
        time.sleep(0.015 + 0.0005 * len(batch))         # illustrative timings only
        return [x.upper() for x in batch]

    async def main() -> None:
        mb, lat, shed = MicroBatcher(fake_model), [], 0

        async def one(i: int, at: float) -> None:
            nonlocal shed
            await asyncio.sleep(at)
            t0 = time.perf_counter()
            try:
                await mb.submit(f"req {i}")
                lat.append((time.perf_counter() - t0) * 1000)
            except Overloaded:
                shed += 1

        arrivals = [random.expovariate(1500) for _ in range(2000)]  # 1,500 requests/s
        await asyncio.gather(*(one(i, sum(arrivals[: i + 1])) for i in range(2000)))
        lat.sort()
        print(f"served {len(lat)}, shed {shed}, p50 {lat[len(lat) // 2]:.0f} ms, "
              f"p99 {lat[int(len(lat) * 0.99)]:.0f} ms")

    asyncio.run(main())

Three decisions in that file are worth defending out loud. The wait window spends TTFT on purpose: 10 ms of waiting against a 450 ms budget buys fuller batches. The queue bound turns overload into an immediate 429 or 503 instead of a timeout thirty seconds later; the demo offers 1,500 requests per second to a stand-in model that serves about 1,000 at full batches, sheds the excess, and keeps latency for admitted requests bounded by the queue length. And run_batch runs in a worker thread, so a blocking GPU call never stalls the event loop that accepts requests.

In front of a generation engine the file changes shape. vLLM, SGLang and TensorRT-LLM schedule per iteration, so holding requests back to fill a batch mostly adds latency. Forward each request at once and let the engine's waiting queue do the batching; keep the bound, the deadline and the rejection path, and treat vllm:num_requests_waiting as the depth of that queue. Chapter 12 builds the rate limiter and the retry loop that sit on either side of it. The rest of this question happens inside the engine.

Static, dynamic and continuous batching

Three policies decide when a request enters the GPU's batch and when it leaves. Static batching collects a fixed group and runs it until the longest member finishes. Dynamic, or request-level, batching forms each group by size and time, as the file above does, and still runs the group to completion. Continuous batching, which the Orca paper (OSDI 2022) called iteration-level scheduling, "schedules execution at the granularity of iteration (instead of request)": after every decode step, finished sequences leave and waiting ones take their places. Orca paired it with selective batching, which batches only selected operations, and on GPT-3 175B reported 36.9x the throughput of NVIDIA FasterTransformer at the same latency.

The figure runs eight requests of different lengths through four batch slots.

Step through STATIC first. R4 finishes after two iterations, and its slot stays hatched for six more until R2 ends, while R5 to R8 wait in the queue. Switch to CONTINUOUS and the same eight requests finish five iterations sooner, because R5 takes R4's slot on the next step. No step got faster; the GPU stopped holding empty seats while requests waited.

Every speedup in this area is measured against a baseline, and the baselines differ by an order of magnitude. Quote the pair or neither.

Reported gainAgainst what, on whatSource
36.9x throughput at equal latencyNVIDIA FasterTransformer, GPT-3 175BOrca, OSDI 2022
Up to 23x throughput (vLLM); 8x from continuous batching alone (Ray Serve, TGI); 4x from FasterTransformer static batchingHugging Face Pipelines static batching; OPT-13B on one A100 40GB; 1,000 requests of 512 input tokens with exponentially distributed output lengthsAnyscale, June 2023
2-4x throughput at the same latencyFasterTransformer and OrcavLLM paper, SOSP 2023
Up to 24x and up to 3.5x throughputHugging Face Transformers and TGI; LLaMA-7B on A10G, LLaMA-13B on A100 40GB, ShareGPT lengthsvLLM launch post, June 2023
2.6x serving capacity under tail-latency limitsvLLM; Mistral-7B on one A100Sarathi-Serve, OSDI 2024

The baselines explain the spread. Anyscale's naive one fell to 81 tokens per second at its highest variance of output length, while the vLLM paper compared against systems that already batched well. All of these gains end where the KV cache of the batch no longer fits in memory.

Related question · Inference

Why did continuous batching replace static batching for LLM serving?

What it tests: Whether you can name the waste in static batching and the change that removed it, and quote the gain with its baseline.

A strong answer, in one minute

Static batching runs a fixed group until its longest request finishes, so short requests hold their slots idle and new arrivals wait for the whole batch. When output lengths vary, as they do in chat, much of the batch is waiting; Anyscale's naive static baseline fell to 81 tokens per second at the highest variance it tested. Orca's iteration-level scheduling re-forms the batch after every decode step: a finished sequence leaves, a waiting one joins, and the GPU stays full. Orca reported 36.9x FasterTransformer's throughput at the same latency on GPT-3 175B. Paged KV memory then let the larger batches fit, which is where vLLM's further 2-4x over Orca and FasterTransformer came from.

Follow-ups

What does continuous batching not fix?
Memory and interference. Admitting more sequences per step helps only while their KV cache fits, which is why vLLM paired it with paging; and a new request's prefill still competes with everyone's decode, which chunked prefill addresses (Section 4).
Someone quotes a 23x speedup for continuous batching. Your response?
Ask for the baseline. Anyscale's 23x compares vLLM with Hugging Face Pipelines static batching on OPT-13B on one A100 40GB, and continuous batching alone gave 8x there. Against Orca and FasterTransformer at equal latency, the vLLM paper reports 2-4x.

The scheduler's knobs in vLLM V1

In vLLM this is the V1 scheduler's job. V1 became the default engine in v0.8.0 (March 2025) and the only one in v0.11.0 (October 2025), and async scheduling became the default in v0.14.0 (January 2026), per the release notes. Each step it hands the GPU a set of tokens, decode tokens for running sequences and prompt tokens for new ones, and a few settings shape that set.

max_num_batched_tokens is the token budget of one step. The scheduler spends it on decodes first, one token per running sequence, then fills what is left with prefill, chunking any prompt that does not fit; vLLM's optimization guide says chunked prefill is enabled by default whenever possible. Sarathi-Serve (OSDI 2024) introduced chunked prefills with stall-free schedules. The guide states the trade: smaller budgets such as 2048 give better inter-token latency, larger ones better TTFT, and for throughput it recommends more than 8192, especially for smaller models on large GPUs.

The figure shows why the budget trades one latency against the other. A long prompt arrives while other users are decoding.

With NO CHUNKING the prefill fills one long step and every running user waits through it for the next token. SMALL BUDGET splits it into four chunks that share steps with the decodes: the longest gap shrinks, and the new prompt's first token lands after the dashed line, later than it would unchunked. LARGE BUDGET sits between the two. The units are illustrative; the direction is the one the vLLM guide describes. Chunking does not remove interference under bursts: in their DistServe retrospective, Hao AI Lab's authors write that "a single large prefill can inflate TPOT by 2~30x, especially under bursty workloads", one argument for running prefill and decode on separate GPUs.

The other settings decide how large the batch may grow. Their defaults depend on the GPU and moved during 2026; the table pins them to vLLM's main branch in October 2026 (arg_utils.py and cache.py).

SettingDefaultWhat it trades
max_num_batched_tokens16384 on GPUs with at least 160 GiB (B200, B300); on H100 and H200, 16384 offline and 8192 for the OpenAI server; on smaller GPUs, A100 included, 8192 offline and 2048 for the serverSmall favors inter-token latency; large favors TTFT and throughput
max_num_seqs1024 on H100, H200, B200 and B300; 256 on smaller GPUsMore sequences per step against more preemption as contexts grow
gpu_memory_utilization0.92 since v0.20.0 (0.9 through v0.19); kv_cache_memory_bytes overrides itCache capacity against headroom for anything else on the GPU
enable_prefix_cachingOn, with sha256 block hashesReuse of shared prefixes; salt it per tenant (Section 8)
kv_cache_dtypeauto; fp8, fp8_e4m3 and fp8_e5m2 availableHalf the cache bytes against calibration work (Section 9)
--scheduling-policyfcfs; priority availableArrival order against per-request priority

When a running sequence needs a block and none is free, V1 preempts, by recompute rather than swap by default: the victim's blocks are freed, it returns to the waiting queue, and its cache is recomputed when it resumes. Under first-come-first-served the victim is the most recently scheduled running request; under the priority policy it is the one with the largest priority value, the later arrival losing a tie (the V1 scheduler's source). The optimization guide's remedies for frequent preemption: raise gpu_memory_utilization, lower max_num_seqs or max_num_batched_tokens, or add tensor or pipeline parallelism.

Priority scheduling lets one engine serve a chat product and a batch job at once. Start the server with --scheduling-policy priority and send an integer priority in the request body: lower values run earlier, the default is 0, any other value is rejected unless the server runs the priority policy, and an X-Vllm-Priority header overrides the body (vLLM's server docs). Put together, one engine step runs like this.

One engine step in vLLM V1

Step 1 / 5
Order the queues
The scheduler keeps a running list and a waiting queue. Under the priority policy the waiting queue is ordered by priority value, lowest first, then by arrival; under fcfs, by arrival alone.

A configuration for the running example, Llama 3.3 70B with FP8 weights on one H200, each flag annotated:

vllm_serve.sh
bash
#!/usr/bin/env bash
# vllm_serve.sh: Llama 3.3 70B on one H200 (141 GB) for interactive chat.
# Flags as of vLLM v0.31.0 (October 2026). Several defaults moved during
# 2025 and 2026, so pin the version and write the values out.
set -euo pipefail

# --quantization fp8_per_tensor   FP8 weights quantized at load from the BF16
#                                  checkpoint: about 70 GB instead of about 140 GB
# --gpu-memory-utilization 0.92   the default since v0.20.0 (0.9 before); weights,
#                                  activations and the KV cache share this slice
# --kv-cache-dtype fp8            160 KiB per token instead of 320 KiB; without
#                                  calibration every scale is 1.0 (Section 9)
# --max-model-len 32768           the longest request: 5 GiB of FP8 cache, about a
#                                  tenth of the budget
# --max-num-seqs 64               the H200 default is 1024; 64 sequences average
#                                  about 5K tokens each in the ~334K-token budget
# --max-num-batched-tokens 2048   token budget per step: decodes first, prefill
#                                  chunks fill the rest; small favors ITL (the
#                                  OpenAI-server default on H100 and H200 is 8192)
# --enable-prefix-caching         on by default; written out for the reviewer
# --scheduling-policy priority    the request field "priority" orders the queue:
#                                  lower values first, ties by arrival
vllm serve meta-llama/Llama-3.3-70B-Instruct \
  --quantization fp8_per_tensor \
  --gpu-memory-utilization 0.92 \
  --kv-cache-dtype fp8 \
  --max-model-len 32768 \
  --max-num-seqs 64 \
  --max-num-batched-tokens 2048 \
  --enable-prefix-caching \
  --scheduling-policy priority \
  --port 8000

# An interactive request that jumps the queue and keeps its cached prefix
# to one tenant (send priority 0 or higher for batch work):
#   curl -s localhost:8000/v1/chat/completions -H 'Content-Type: application/json' \
#     -d '{"model": "meta-llama/Llama-3.3-70B-Instruct", "priority": -10,
#          "cache_salt": "<random secret per tenant>",
#          "messages": [{"role": "user", "content": "Summarize this ticket."}]}'

The values follow from the arithmetic in Sections 6 and 7 and are starting points for a load test: vllm bench serve with --burstiness (a Gamma shape, where 1.0 is Poisson) and --goodput (latency targets in milliseconds) shows whether they hold. Four metrics show how the scheduler copes: vllm:num_requests_waiting (the queue), vllm:num_requests_running (the batch), vllm:kv_cache_usage_perc (cache pressure, where 1 means 100 percent) and vllm:num_preemptions. Chapter 3 autoscales on the queue and the cache pressure.

Pin the version before you quote a default
vLLM's defaults moved during 2025 and 2026. gpu_memory_utilization went from 0.9 to 0.92 in v0.20.0; vllm:gpu_cache_usage_perc became vllm:kv_cache_usage_perc; v0.15.0 removed vllm:time_per_output_token_seconds in favor of vllm:inter_token_latency_seconds; and the batching defaults depend on GPU memory and on whether you run the server or offline. Tutorials from 2025 still use the old names. Say which version a number belongs to.

Related question · Inference

A junior engineer's design document for a batched inference service lands on your desk. What do you look for?

What it tests: Whether you can critique a design against the mechanics of batching and memory, rank the problems, and suggest fixes without rewriting the document.

A strong answer, in two minutes

I read it for six things, roughly in the order they would hurt in production. First, the batching policy: a fixed batch size with a timeout in front of a generation model is static batching, and I would ask them to hand batching to an engine that schedules per iteration. Second, a KV budget: the document should say how many tokens of cache the GPU holds and derive the maximum sequences and context from it. Third, the queue: bounded, with a deadline, a rejection path and a metric for its depth. Fourth, targets for TTFT and inter-token latency at p99, measured separately. Fifth, how a long prompt and a priority request are handled. Sixth, a load test at realistic burstiness. I would write each comment as a question with a number attached.

Follow-ups

The document sets the maximum batch to 256 because bigger batches are faster. Your comment?
Ask what 256 sequences of their median context cost in cache. For Llama 3.3 70B in BF16 at 8K that is about 687 GB against about 54.7 GB available on one H200 with FP8 weights, so the engine would preempt long before 256, and every added sequence also lengthens each step for everyone.
They plan to autoscale on GPU utilization. Your comment?
Google's GKE guidance notes that DCGM_FI_DEV_GPU_UTIL measures the duty cycle and does not measure how much work is done, and recommends scaling on the server's queue size or batch size. Chapter 3 builds the autoscaler.

Follow-ups on Question 1

What happens when the KV cache fills up?
vLLM V1 preempts a running request by recompute: frees its blocks, queues it again and recomputes its cache when it resumes. Watch vllm:num_preemptions; the documented remedies are more gpu_memory_utilization, a lower max_num_seqs or max_num_batched_tokens, or more tensor or pipeline parallelism.
A user sends a 32K-token prompt. What happens to everyone else?
Without chunking, its prefill runs as one long step and every running user waits through it for their next token. With chunked prefill, decodes go first and the prompt is spread across steps under max_num_batched_tokens. Even so, Hao AI Lab found that one large prefill can inflate TPOT by 2-30x under bursty load.
When would you split prefill and decode onto different GPUs?
When TTFT and inter-token latency targets conflict and you need to tune them separately or protect tail latency. vLLM's own docs say disaggregated prefill does not improve throughput, and the cache must move: an 8K prompt on a 70B model is about 2.68 GB of BF16 cache, about 54 ms over 400 Gb/s RDMA (the book's arithmetic). Chapter 3 returns to it as a latency fix.
How do you keep interactive users ahead of batch jobs on the same server?
Run the priority policy and send a lower priority value on interactive requests; under memory pressure the scheduler also preempts the largest value first. If batch traffic is heavy, a separate pool is simpler to reason about. NVIDIA reports that Dynamo's agent hints (latency sensitivity, expected output length, cache control) gave up to 4x lower TTFT on Hopper, a vendor number.

Red flags

  • Presenting a fixed batch size with a timeout as continuous batching.
  • Claiming larger batches always raise throughput, with no word on KV memory or per-user latency.
  • Quoting 23x without its baseline: naive static batching of OPT-13B on one A100.
  • An unbounded queue, so overload shows up as timeouts instead of fast rejections.
  • Treating TTFT and inter-token latency as one number, or not knowing that decode is bandwidth-bound.

Question 2: what the KV cache is and why it limits concurrency

Question 2 · Inference

Explain what the KV cache is, and why it, more than compute, limits how many users one GPU can serve.

What it tests: Whether you can derive the cache size from a model's shape, turn free memory into a concurrency limit, and name the architectural and system levers that move it.

A strong answer, in three minutes

During generation every attention layer needs the keys and values of every earlier token. The KV cache keeps them in GPU memory, so a decode step computes K and V for one new token instead of the whole prefix. It grows linearly with context length and with the number of sequences in the batch.

Its size comes straight from config.json: two (K and V) times layers times KV heads times head dimension times bytes per element. Llama 3.3 70B has 80 layers, 8 KV heads and a head dimension of 128, so in BF16 that is 327,680 bytes, 320 KiB per token: 2.5 GiB for an 8K conversation and 40 GiB for one at 128K.

Concurrency is free memory divided by bytes per sequence. Put Llama 3.3 70B with FP8 weights on one 141 GB H200: at vLLM's default 0.92 utilization, about 54.7 GB is left for the cache after 70 GB of weights and roughly 5 GB of activations. That is about 167K tokens, so about 20 sequences of 8K in BF16, or about 41 with an FP8 cache. Compute is nowhere near its limit at that batch; memory ran out first. Decode also reads the whole cache every step, so long contexts cost bandwidth as well as capacity.

The levers come in two kinds. Architecture: grouped-query attention with 8 KV heads instead of 64, DeepSeek's latent attention at 68.6 KiB per token, sliding windows and hybrid linear attention. System: paging to remove fragmentation, prefix sharing, FP8 or FP4 storage, and offload to CPU memory or storage.

This question is absent from Aced's lists. Outcome School's company-wise collection tags it to OpenAI, xAI, Mistral, Amazon, Apple, NVIDIA, Together and Character.AI and asks candidates to derive the formula, though the researchers found that the collection mirrors other prep lists. Expect the derivation and the arithmetic, done aloud.

The arithmetic was already visible in 2023. Anyscale's back-of-envelope for a 13B model came to about 1 MB of state per token: an A100 40GB holding 26 GB of weights leaves 14 GB, about 14K tokens, so at most about 28 sequences of 512 tokens or 7 of 2,048. The vLLM paper did it exactly for OPT-13B: 2 (K and V) × 5,120 hidden × 40 layers × 2 bytes = 800 KB per token, up to 1.6 GB for one 2,048-token request, with about 65% of an A100 40GB holding weights and close to 30% holding the cache.

Compute stays far from its limit at those batch sizes. In the decode figure of Section 1, 20 sequences of 8K need about 2.8 ms of compute per step against about 25.8 ms of memory traffic. Memory fills long before the tensor cores are busy, and everything it holds is read again on every step. The rest of this question is arithmetic: how many bytes per token a model needs (Section 6), how many tokens fit on a GPU (Section 7), how the engine manages them (Section 8) and how to shrink them (Section 9).

From config.json to bytes per token

NVIDIA's inference optimization guide writes the cache per token as 2 × layers × (heads × head dimension) × bytes per element, where for grouped-query and multi-query models the heads are the KV heads. Every input is a field in the model's config.json.

KV cache per token: MHA, GQA, MQA
\text{bytes per token} = 2 \times L \times H_{kv} \times d_{\text{head}} \times b
L = num_hidden_layers, H_kv = num_key_value_heads, d_head = head_dim (or hidden_size / num_attention_heads), b = 2 for BF16 and 1 for FP8.

Llama 3.3 70B's config.json gives 80 layers, 64 attention heads, 8 KV heads and a head dimension of 128: 2 × 80 × 8 × 128 × 2 = 327,680 bytes, 320 KiB per token in BF16. Plain multi-head attention, with 64 KV heads, would need eight times as much; that eightfold saving is what grouped-query attention bought.

DeepSeek's multi-head latent attention (MLA) needs a different formula. Each layer caches one compressed latent vector of width kv_lora_rank (512) and one RoPE key of width qk_rope_head_dim (64) shared by all heads, so the head count drops out: (512 + 64) × 61 layers × 2 bytes = 70,272 bytes, about 68.6 KiB per token for DeepSeek-V3. An ordinary multi-head cache for the same 128 heads, with keys 192 wide and values 128 wide, would be 61 × 128 × (192 + 128) × 2 = 4,997,120 bytes, about 4.77 MiB, 71 times more. A GQA-style reading of this config gives a wrong answer, which is why the code below checks for kv_lora_rank first.

KV cache per token: MLA
\text{bytes per token} = (r_{kv} + d_{\text{rope}}) \times L \times b
r_kv = kv_lora_rank, d_rope = qk_rope_head_dim; the number of attention heads does not appear.

Hybrid models add one step: count only the layers whose cache grows with context. The gpt-oss model card describes 120b as alternating dense layers with banded-window layers that see 128 tokens, so its 18 full layers cost 2 × 18 × 8 × 64 × 2 = 36,864 bytes per token while the 18 window layers hold a fixed 4.5 MiB per sequence. Qwen3.8-27B lays out 64 layers as 16 groups of three Gated DeltaNet layers and one full-attention layer; only the 16 full layers (4 KV heads, head dimension 256) grow, at 64 KiB per token, while the linear-attention layers keep a fixed-size state. Llama 4 Scout and Maverick attend within 8,192-token chunks in 36 of their 48 layers (the transformers config), so past 8K only 12 layers grow. The script below reads these shapes from the config keys that carry them: layer_types, full_attention_interval, linear_attn_config.full_attn_layers and sliding_window.

kv_calc.py
python
"""kv_calc.py: KV cache bytes per token from a Hugging Face config.json, and
how many sequences fit in a KV budget. Python 3.11+.

    pip install huggingface_hub
    python kv_calc.py
"""
import json
from dataclasses import dataclass

from huggingface_hub import hf_hub_download

BYTES = {"bf16": 2, "fp16": 2, "fp8": 1}  # per stored element of K or V


@dataclass
class KV:
    per_token: int  # bytes each token adds, summed over the layers that grow
    per_seq: int    # bytes a sequence holds whatever its length (sliding windows)
    shape: str


def load(repo_id: str) -> dict:
    # Gated repos (meta-llama) need HF_TOKEN; mirrors such as unsloth's do not.
    with open(hf_hub_download(repo_id=repo_id, filename="config.json")) as f:
        cfg = json.load(f)
    return cfg.get("text_config", cfg)  # multimodal configs nest the language model


def kv_bytes(cfg: dict, dtype: str = "bf16") -> KV:
    b, layers = BYTES[dtype], cfg["num_hidden_layers"]
    if "kv_lora_rank" in cfg:  # MLA: one latent plus one shared RoPE key per layer
        r, rope = cfg["kv_lora_rank"], cfg["qk_rope_head_dim"]
        mla = (cfg.get("linear_attn_config") or {}).get("full_attn_layers")
        n = len(mla) if mla else layers  # Kimi Linear and K3: only MLA layers grow
        return KV(n * (r + rope) * b, 0, f"({r} + {rope}) x {n} layers x {b} B")
    kv_heads = cfg.get("num_key_value_heads", cfg["num_attention_heads"])
    head_dim = cfg.get("head_dim") or cfg["hidden_size"] // cfg["num_attention_heads"]
    types = cfg.get("layer_types")
    if types is None and (k := cfg.get("full_attention_interval")):  # Qwen3-Next style
        types = ["linear_attention" if (i + 1) % k else "full_attention" for i in range(layers)]
    types = types or ["full_attention"] * layers
    full, sliding = types.count("full_attention"), types.count("sliding_attention")
    per_layer = 2 * kv_heads * head_dim * b  # K and V
    window = cfg.get("sliding_window") or 0
    shape = f"2 x {full} layers x {kv_heads} kv heads x {head_dim} x {b} B"
    if sliding:
        shape += f", plus {sliding} layers capped at {window} tokens"
    return KV(full * per_layer, sliding * per_layer * window, shape)


def sequences_that_fit(kv: KV, context: int, hbm_gb: float, weights_gb: float,
                       util: float = 0.92, activations_gb: float = 5.0) -> float:
    """vLLM reserves util x HBM; weights and activations come out first (GB = 1e9 B)."""
    budget = (hbm_gb * util - weights_gb - activations_gb) * 1e9
    return budget / (kv.per_token * context + kv.per_seq)


if __name__ == "__main__":
    for repo in ["unsloth/Llama-3.3-70B-Instruct", "deepseek-ai/DeepSeek-V3",
                 "openai/gpt-oss-120b", "Qwen/Qwen3.8-27B"]:
        kv = kv_bytes(load(repo))
        at_128k = (kv.per_token * 131_072 + kv.per_seq) / 2**30
        print(f"{repo:31} {kv.per_token:>7,} B/token  {at_128k:5.2f} GiB at 128K  {kv.shape}")
    llama = load("unsloth/Llama-3.3-70B-Instruct")
    for dtype in ("bf16", "fp8"):  # one H200, FP8 weights (about 70 GB)
        for ctx in (8_192, 32_768):
            n = sequences_that_fit(kv_bytes(llama, dtype), ctx, hbm_gb=141, weights_gb=70)
            print(f"Llama 3.3 70B, one H200, {dtype} KV, {ctx // 1024}K: {n:.1f} sequences")

# unsloth/Llama-3.3-70B-Instruct  327,680 B/token  40.00 GiB at 128K  2 x 80 layers x 8 kv heads x 128 x 2 B
# deepseek-ai/DeepSeek-V3          70,272 B/token   8.58 GiB at 128K  (512 + 64) x 61 layers x 2 B
# openai/gpt-oss-120b              36,864 B/token   4.50 GiB at 128K  2 x 18 layers x 8 kv heads x 64 x 2 B, plus ...
# Qwen/Qwen3.8-27B                 65,536 B/token   8.00 GiB at 128K  2 x 16 layers x 4 kv heads x 256 x 2 B
# Llama 3.3 70B, one H200, bf16 KV, 8K: 20.4 sequences   (32K: 5.1)
# Llama 3.3 70B, one H200, fp8 KV, 8K: 40.8 sequences    (32K: 10.2)

The rows below come from each model's config.json, read on 2026-10-07, and an FP8 cache halves each of them; the last row is DeepSeek's own figure for a model that already stores its cache in FP4.

ModelAttentionArithmetic (BF16)Per token
Llama 3.1 8BGQA, 8 KV heads2 × 32 × 8 × 128 × 2128 KiB
Llama 3.3 70BGQA, 8 KV heads2 × 80 × 8 × 128 × 2320 KiB
Llama 3.1 405BGQA, 8 KV heads2 × 126 × 8 × 128 × 2504 KiB
Qwen3-32BGQA, 8 KV heads2 × 64 × 8 × 128 × 2256 KiB
Qwen3.8-27BHybrid, 16 of 64 layers full2 × 16 × 4 × 256 × 264 KiB
gpt-oss-120b18 full + 18 window (128)2 × 18 × 8 × 64 × 236 KiB + 4.5 MiB per sequence
Llama 4 Scout, Maverick12 full + 36 chunked (8,192)2 × 12 × 8 × 128 × 248 KiB past 8K
DeepSeek-V3/R1, Kimi K2MLA, 512 + 64(512 + 64) × 61 × 268.6 KiB
DeepSeek-V3 if it were MHA128 heads, K 192, V 12861 × 128 × 320 × 24,880 KiB
DeepSeek V4.1-FlashCompressed, FP4 cacheDeepSeek's figure890 bytes

Related question · Inference

From their config files, compute the KV cache per token for DeepSeek-V3 and for Llama 3.3 70B. Which serves more 32K-token users per GPU, and why?

What it tests: Whether you read GQA and MLA configs correctly and keep cache density apart from what fits on a GPU once the weights are counted.

A strong answer, in two minutes

Llama 3.3 70B: 2 × 80 layers × 8 KV heads × 128 × 2 bytes, 320 KiB per token, so a 32K sequence holds 10 GiB. DeepSeek-V3 caches a 512-wide latent and a 64-wide RoPE key per layer: (512 + 64) × 61 × 2, about 68.6 KiB per token, 2.14 GiB at 32K. Per gigabyte of cache DeepSeek-V3 holds about 4.7 times as many users: the 54.7 GB an H200 has left after Llama's FP8 weights would hold about 24 of its 32K sequences against about 5 of Llama's. Per GPU the comparison flips, because DeepSeek-V3's weights need many GPUs before the first token; DeepSeek's own V3 deployment used a minimum of 32 GPUs for prefill and 320 for decode. So I would answer both ways: per gigabyte of cache, DeepSeek-V3; per GPU you already own, Llama, because DeepSeek-V3 does not run on one.

Follow-ups

Does the BF16 figure for DeepSeek hold in production?
Not exactly. DeepSeek's stacks store FP8 latents with scales: FlashMLA's format for V3.2 takes 656 bytes per token per layer (512 FP8 values, 16 bytes of scales and 128 bytes of BF16 RoPE values, kept unquantized for accuracy), and DeepSeek's own chart puts V3.2 at 48,068 bytes per token.
Where does head_dim come from when the config omits it?
hidden_size divided by num_attention_heads. Qwen2.5-72B has a hidden size of 8192 and 64 heads, so 128, and 2 × 80 × 8 × 128 × 2 gives the same 320 KiB per token as Llama 3.3 70B.

From bytes per token to users per GPU

Bytes per token become users in three subtractions. vLLM claims gpu_memory_utilization of the card, 0.92 by default; the weights come out first, then activations and CUDA graphs; the rest is cache. For Llama 3.3 70B with FP8 weights on one H200: 141 GB × 0.92 = 129.7 GB, less about 70 GB of weights and about 5 GB of activations, leaves about 54.7 GB. That holds about 167K tokens with a BF16 cache or 334K with FP8: about 20 sequences of 8K or 5 of 32K in BF16, about 41 or 10 in FP8. This is the book's arithmetic and the 5 GB is an estimate, so treat the results as planning numbers to confirm in a load test.

Paging changes what the number means. The engine allocates blocks as tokens arrive, so the budget is a pool of about 167K token slots shared by sequences of any length: twenty users at 8K cost the same as 160 at 1K, and a capacity answer states the length distribution it assumes. Parameter count predicts little. Mistral-7B-v0.3 on one 80 GB GPU leaves about 59.1 GB of cache, about 451K tokens. Qwen2.5-72B in BF16 needs about 145 GB of weights: two 80 GB GPUs leave about 1.8 GB of cache (about 5,500 tokens, not viable), and four leave about 149 GB, about 455K tokens, the 7B model's budget on four times the hardware. (Both examples skip the activation estimate, so they are upper bounds.)

The calculator holds the 54.7 GB budget fixed and lets you change the model, the cache precision and the context length.

Pick LLAMA 3.3 70B at 8K to see the 20.4 sequences from above, then switch to FP8 KV and watch the count double. DEEPSEEK-V3 fits about 95 sequences of 8K in the same budget and GPT-OSS-120B about 178. Drag CONTEXT to 128K and Llama holds one sequence with room for a quarter of another. Because the budget is fixed, this compares the models per gigabyte of cache; on real hardware DeepSeek-V3's weights need a multi-GPU deployment and gpt-oss-120b's MXFP4 checkpoint is 60.8 GiB, so redo the subtraction for the deployment you are sizing.

Holding the cache is one cost and reading it is another. Every decode step reads the cache of every running sequence: at 20 × 8K in BF16 that is about 53.7 GB per step, nearly as much as the weights, which is why the step floor in Section 1 rises from 15.1 ms to 25.8 ms as the batch fills.

Size the cache for three 2026 configs

Exercise

Compute bytes per token, GiB for one 128K sequence, and how many 32K sequences fit in the 54.7 GB budget, for a GQA model with 4 KV heads, a hybrid MLA model and a large GQA model. Give yourself ten minutes.

Steps
  1. Write kv_bytes_per_token(cfg) for GQA and MLA configs; for MLA, count only the layers listed in linear_attn_config.full_attn_layers when the key is present.
  2. Apply it to the three config fragments in the starter.
  3. For each model, compute GiB at 131,072 tokens (divide by 2^30) and the number of 32K sequences in the budget (decimal gigabytes, as in this section).
  4. Check one model by hand before trusting your code.
Starter
kv_exercise.py
python
"""Starter: the KV cache of three models from their config.json values."""
QWEN3_235B = {"num_hidden_layers": 94, "num_attention_heads": 64,
              "num_key_value_heads": 4, "head_dim": 128}
KIMI_K3 = {"num_hidden_layers": 93, "kv_lora_rank": 512, "qk_rope_head_dim": 64,
           # 24 entries: 4, 8, ..., 92, 93; the other 69 layers are KDA (linear)
           "linear_attn_config": {"full_attn_layers": [*range(4, 93, 4), 93]}}
GLM_4_5 = {"num_hidden_layers": 92, "num_attention_heads": 96,
           "num_key_value_heads": 8, "head_dim": 128}
MODELS = [("Qwen3-235B-A22B", QWEN3_235B), ("Kimi K3", KIMI_K3), ("GLM-4.5", GLM_4_5)]
BUDGET = (141 * 0.92 - 70 - 5) * 1e9  # bytes of cache on one H200, as in Section 7


def kv_bytes_per_token(cfg: dict, bytes_per_elem: int = 2) -> int:
    raise NotImplementedError  # GQA and MLA; count only the layers whose cache grows


for name, cfg in MODELS:
    per_token = kv_bytes_per_token(cfg)
    # print: bytes per token, GiB for one 131,072-token sequence, 32K sequences in BUDGET
Success criteria
  • - Bytes per token for all three models, with the arithmetic written out
  • - Kimi K3 counted on its 24 MLA layers, not all 93
  • - A one-line explanation of why the hybrid model fits more than ten times as many 32K sequences as the large GQA model
Hint
Kimi K3 caches a 512-wide latent plus a 64-wide RoPE key on 24 of its 93 layers; the other 69 are linear-attention (KDA) layers with a fixed-size state.
Reveal solution
kv_exercise.py
python
# Replace the starter's function and loop with these.
def kv_bytes_per_token(cfg: dict, bytes_per_elem: int = 2) -> int:
    if "kv_lora_rank" in cfg:  # MLA: one latent plus one shared RoPE key per MLA layer
        mla = (cfg.get("linear_attn_config") or {}).get("full_attn_layers")
        layers = len(mla) if mla else cfg["num_hidden_layers"]
        return (cfg["kv_lora_rank"] + cfg["qk_rope_head_dim"]) * layers * bytes_per_elem
    return 2 * cfg["num_hidden_layers"] * cfg["num_key_value_heads"] * cfg["head_dim"] * bytes_per_elem


for name, cfg in MODELS:
    per_token = kv_bytes_per_token(cfg)
    gib_128k = per_token * 131_072 / 2**30
    fit_32k = BUDGET / (per_token * 32_768)
    print(f"{name:16} {per_token:>7,} B/token  {gib_128k:5.2f} GiB at 128K  {fit_32k:4.1f} x 32K")

# Qwen3-235B-A22B  192,512 B/token  23.50 GiB at 128K   8.7 x 32K
# Kimi K3           27,648 B/token   3.38 GiB at 128K  60.4 x 32K
# GLM-4.5          376,832 B/token  46.00 GiB at 128K   4.4 x 32K

Related question · Inference

Estimate the KV memory of a 70B-class model at 128K context. Does the model fit on one 80 GB GPU, and what single-stream speed should you expect?

What it tests: Whether you can do the arithmetic aloud, notice that the weights already fill the card, and offer options instead of stopping at no.

A strong answer, in two minutes

One 128K sequence of Llama 3.3 70B holds 40 GiB of BF16 cache, 20 GiB in FP8. The weights are about 140 GB in BF16 and about 70 GB in FP8. On an 80 GB H100 at vLLM's default 0.92, 73.6 GB is usable, so even FP8 weights leave a few gigabytes before activations and nothing for a 128K cache. It does not fit. Single-stream speed follows from bandwidth: BF16 weights split across two H100s decode at about 21 ms per token, at most about 48 tokens per second, and FP8 weights on one H200 at about 14.6 ms, about 69. For 128K contexts I would offer FP8 weights on an H200 or across two GPUs, an FP8 cache, a lower context limit for most users, or a model with a smaller cache per token.

Follow-ups

Two 80 GB GPUs, then. Is that enough for a 72B model in BF16?
Barely, and not usefully. The book's arithmetic for Qwen2.5-72B: about 145 GB of weights against 147.2 GB usable at TP=2 leaves about 1.8 GB of cache, about 5,500 tokens. At TP=4 the cache grows to about 149 GB, about 455K tokens.
How would you serve 128K contexts to many users anyway?
Cap the context per tier, store the cache in FP8, share common prefixes, offload cold blocks to CPU memory, and prefer a model with a smaller cache per token. Section 9 goes through the levers.

Paging, prefix sharing and preemption

Related question · Inference

Explain PagedAttention, then implement the block manager for a paged KV cache with copy-on-write prefix sharing.

What it tests: Whether you know why contiguous allocation wastes memory and can turn block tables, reference counts and copy-on-write into working code.

A strong answer, in two minutes, before the code

Systems that gave each request one contiguous region of cache, sized for its maximum length, left most of that memory reserved or fragmented. PagedAttention stores the cache in fixed blocks of 16 tokens, like pages of virtual memory. Each sequence keeps a block table from logical to physical blocks, allocated on demand, so only its last block can be partly empty. Blocks carry reference counts, so samples of one prompt share its blocks, and a sequence about to write into a shared, partly filled block copies it first. The attention kernel reads keys and values through the block table. Then I would write the manager: admit, append, fork and release.

The evidence is in the vLLM paper (SOSP 2023). Before paging, "only 20.4% - 38.2% of the KV cache memory is used to store the actual token states in the existing systems"; the launch post put the waste at 60% to 80%, the same finding framed the other way, and paging brought it under 4%. The paper chose 16-token blocks as "large enough to efficiently utilize the GPU and small enough to avoid significant internal fragmentation in most workloads". With the memory back the batch grows: 2-4x the throughput of FasterTransformer and Orca at the same latency, and, per the launch post, LMSYS cut its serving GPUs by 50%. Sharing saves more: 6.1% to 9.8% of memory for parallel sampling and 37.6% to 55.2% for beam search on Alpaca, and 16.2% to 30.5% and 44.3% to 66.3% on ShareGPT.

Automatic prefix caching extends sharing across requests. vLLM's design doc hashes each block from its parent block's hash, its own token IDs and extra keys (LoRA IDs, multimodal input hashes, cache salts). A new request walks its prompt block by block and reuses every block whose hash is already cached. "We only cache full blocks," the doc says, so a prompt's partial tail is always computed. sha256 has been the default hash since v0.11, with sha256_cbor and xxhash as alternatives. Freed blocks join the tail of a least-recently-used free queue and stay reusable until they are evicted, which is how a shared system prompt survives between conversations.

Sharing opens a side channel: a cached prefix answers faster, so response times can reveal what another tenant sent. A per-request cache_salt goes into the hash of the first block, so only requests with the same salt reuse each other's blocks. vLLM's request schema asks for a salt that is random, protected from third parties and long enough to be unpredictable, such as 43 base64 characters (256 bits).

The bookkeeping fits in one class. This is a teaching model, not vLLM's code: block tables, reference counts and copy-on-write as the paper describes them, plus the hash chain, the salt and the LRU free queue from the prefix-caching design doc.

block_manager.py
python
"""block_manager.py: the bookkeeping of a paged KV cache. Block tables and
copy-on-write as in the vLLM paper, plus prefix caching on a hash chain of
full blocks as in vLLM's docs. A teaching model, not vLLM's code (Python 3.11+).
"""
import hashlib
from dataclasses import dataclass, field

BLOCK = 16  # tokens per block, vLLM's default


@dataclass(eq=False)
class Block:
    id: int
    refs: int = 0
    tokens: list[int] = field(default_factory=list)
    hash: str | None = None  # set once the block is full


def chain(parent: str | None, tokens: list[int], salt: str) -> str:
    # The salt enters the first block's hash; each later hash covers its parent's,
    # so the salt, or one changed early token, reaches every block after it.
    seed = parent if parent is not None else f"salt:{salt}"
    return hashlib.sha256(f"{seed}|{tokens}".encode()).hexdigest()


class BlockManager:
    def __init__(self, num_blocks: int):
        self.free = [Block(i) for i in range(num_blocks)]   # the front is evicted first
        self.cache: dict[str, Block] = {}                     # full-block hash -> block
        self.seqs: dict[str, tuple[list[Block], str]] = {}  # id -> (block table, salt)

    def _take(self, tokens: list[int]) -> Block:
        if not self.free:
            raise MemoryError("out of KV blocks: preempt a sequence")
        b = self.free.pop(0)
        if b.hash and self.cache.get(b.hash) is b:
            del self.cache[b.hash]  # reusing a freed block evicts the prefix it held
        b.refs, b.tokens, b.hash = 1, list(tokens), None
        return b

    def _seal(self, b: Block, parent: str | None, salt: str) -> None:
        if len(b.tokens) == BLOCK:  # only full blocks are cached
            b.hash = chain(parent, b.tokens, salt)
            self.cache.setdefault(b.hash, b)

    def admit(self, seq: str, prompt: list[int], salt: str = "") -> int:
        """Build the block table for a new prompt; return the tokens found in cache."""
        table, hits = [], 0
        for i in range(0, len(prompt), BLOCK):
            chunk, parent = prompt[i:i + BLOCK], (table[-1].hash if table else None)
            b = self.cache.get(chain(parent, chunk, salt))
            if b is not None and hits == i:  # an unbroken run of full blocks from the start
                if b.refs == 0:
                    self.free.remove(b)  # a freed block that still holds this prefix
                b.refs, hits = b.refs + 1, hits + BLOCK
            else:
                b = self._take(chunk)
                self._seal(b, parent, salt)
            table.append(b)
        self.seqs[seq] = (table, salt)
        return hits

    def append(self, seq: str, token: int) -> None:
        """One decode step: write a token, copying the last block first if it is shared."""
        table, salt = self.seqs[seq]
        if len(table[-1].tokens) == BLOCK:
            table.append(self._take([]))
        elif table[-1].refs > 1:  # copy-on-write: never write into a shared block
            table[-1].refs -= 1
            table[-1] = self._take(table[-1].tokens)
        table[-1].tokens.append(token)
        self._seal(table[-1], table[-2].hash if len(table) > 1 else None, salt)

    def fork(self, seq: str, child: str) -> None:
        """Parallel sampling or beam search: the child shares every block."""
        table, salt = self.seqs[seq]
        for b in table:
            b.refs += 1
        self.seqs[child] = (list(table), salt)

    def release(self, seq: str) -> None:
        for b in reversed(self.seqs.pop(seq)[0]):  # the last block is evicted first
            b.refs -= 1
            if b.refs == 0:
                self.free.append(b)  # LRU: still cached until it reaches the front
block_demo.py
python
from block_manager import BlockManager

m, system = BlockManager(num_blocks=64), list(range(1000, 1040))  # a 40-token prompt
print(m.admit("a", system + [1, 2, 3]))        # 0: cold cache
print(m.admit("b", system + [7, 8]))           # 32: two full blocks reused, 8 tokens new
m.fork("b", "b2")                              # a second sample of the same prompt
m.append("b2", 99)                             # copy-on-write of the shared last block
print(len(m.free))                             # 59: a took 3 blocks, b 1, b2's copy 1
print(m.admit("c", [999] + system[1:]))        # 0: one changed first token, no hits
print(m.admit("d", system, salt="tenant-42"))  # 0: same tokens, another tenant's salt
for seq in ("a", "b", "b2"):
    m.release(seq)
print(m.admit("e", system))                    # 32: freed blocks still hold the prefix

The demo walks through each behavior: a cold miss, two shared blocks, a copy-on-write, a changed first token that misses every block after it, a salted request that misses identical tokens, and a released prefix that still hits. When the free list is empty, _take raises, which is the point where Section 4's scheduler preempts a victim. The victim's full blocks stay in the cache until evicted, so a resumed request can find part of its own prefix waiting.

Follow-ups

Why 16 tokens a block?
The paper's trade-off: large enough to use the GPU efficiently, small enough to avoid much internal fragmentation in most workloads. Only the last block of each sequence can be partly empty.
nvidia-smi shows free GPU memory, yet the server says its cache is full. How?
vLLM reserves its share of GPU memory at startup (0.92 by default) and carves the cache blocks out of it, so the cache can be full while nvidia-smi shows the remaining slice as free. The fix is a budget decision: raise gpu_memory_utilization or set kv_cache_memory_bytes, shorten contexts, or shrink the bytes per token.

Across replicas a prefix helps only if the request lands where its blocks are. In llm-d's benchmark (8 vLLM pods on 16 H100s, Qwen-32B, 150 customers with 6,000-token contexts, cache demand at 73% of cluster capacity), precise prefix-cache-aware scheduling held p90 TTFT at 0.542 s against 31.083 s for approximate prefix routing and 94.865 s for load-aware routing, and KServe reported 3x the output tokens per second and 2x lower TTFT for Llama 3.1 70B on four MI300X after enabling it. Chapter 3 covers routing.

Related question · Inference

After a prompt change, the prefix-cache hit rate fell from 70% to 20%. What happened, and how do you find it?

What it tests: Whether you know how prefix caches match (full blocks, hash chains from the first token) and debug from metrics and token diffs instead of guesses.

A strong answer, in one minute

Something near the start of the prompt now differs between requests. Prefix caches match only identical leading full blocks, and each block's hash covers the one before it, so a timestamp, a user name, a reordered tool list or a reworded system prompt in the first block makes every later block miss. Truncation breaks it from the other side; the LMCache authors report that context truncation can cut the hit ratio by half. To find it, compute hits over queries from vllm:prefix_cache_hits and vllm:prefix_cache_queries per route, diff the token IDs of two consecutive requests, and find the first block that differs. Then move stable content first and per-request content last.

Follow-ups

What did the 2026-07-28 MCP revision change here?
The revised spec says servers SHOULD return tools/list in a deterministic order, to improve prompt-cache hit rates. A tool list that reorders between calls breaks the prefix at the tools block. Chapter 7 covers the revision.
Does the same apply to a hosted API's prompt cache?
Yes: provider caches also match from the first token, so the same edits break them. Chapter 4 of Harness Engineering lists the client-side edits that silently break a cached prefix.

Shrinking the cache: architecture, FP8 and FP4, offload

Two kinds of lever shrink the cache: change what the model stores, or change how the engine stores it. The first kind moved furthest in 2025 and 2026.

  • Grouped-query attention. Several query heads share each key and value head, so 8 KV heads against 64 query heads save eightfold. The GQA paper found quality close to multi-head attention at speed comparable to multi-query attention.
  • Latent attention. MLA caches a compressed latent instead of per-head keys and values. Against DeepSeek 67B, the DeepSeek-V2 paper reports a 93.3% smaller KV cache and 5.76 times the maximum generation throughput.
  • Windows and chunks. Layers that do not need the whole context get a cap: gpt-oss uses a 128-token band on half its layers, Llama 4 an 8,192-token chunk on three quarters of them.
  • Hybrid linear attention. Most layers keep a fixed-size state and only a minority attend fully: 16 of 64 in Qwen3.8-27B, 24 of 93 in Kimi K3. The Kimi Linear paper reports beating full MLA under the same training recipe while cutting KV cache use by up to 75% and decoding up to 6 times faster at a 1M context.
  • Compressed and sparse attention. DeepSeek-V4 (April 2026) combines compressed sparse attention with heavily compressed attention; its model card says "DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache compared with DeepSeek-V3.2" at 1M tokens. V4.1-Flash (September 2026) keeps its main cache in FP4, E2M1 values with one E4M3 scale per 16 channels, and reaches 890 bytes per token.

DeepSeek's chart on the V4.1-Flash model card puts the trend in four points.

From V1 in November 2023 (389,120 bytes per token, a GQA design with 95 layers and 8 KV heads, close to Llama 3.3 70B's 327,680, the dashed line) to V4.1-Flash in September 2026 (890 bytes), the cache per token fell 437 times, in steps of 8.1x, 13.7x and 3.9x. The chart mixes architecture with precision, since V3.2 is counted with FP8 latents and V4.1-Flash with an FP4 cache, so read it as the cache DeepSeek ships, precision included.

The engine-side lever is precision. vLLM's kv_cache_dtype takes auto, fp8, fp8_e4m3 (CUDA 11.8 or later, and ROCm) or fp8_e5m2, and an FP8 cache halves every GQA and MLA row of the table in Section 6. Scales are per tensor, or per attention head with the Flash Attention backend and llm-compressor calibration. NVIDIA's research blog reports that an NVFP4 cache cuts the footprint by up to about 50% against FP8 with under 1% accuracy loss, and beats an MXFP4 cache by about 5%; these are NVIDIA's own results. Quantizing the weights leaves the cache untouched, since the two are separate settings. Chapter 3 compares the 4-bit weight formats, and The Ultimate Guide to Local LLMs covers quantization on smaller hardware.

An FP8 cache needs calibration
Without calibration, vLLM sets every FP8 KV scale to 1.0. Its docs recommend calibrating on a dataset with llm-compressor; --kv-cache-dtype-skip-layers keeps sensitive layers, such as sliding-window ones, in their native dtype; and with Flash Attention 3 the queries are quantized to FP8 too. Run your evaluation set before and after the switch.

The last lever moves cold blocks out of HBM. LMCache, plugged into vLLM, reports up to 15x throughput on multi-round question answering and document analysis. Mooncake, Kimi's serving platform, uses the cluster's CPU, DRAM and SSD resources as one cache pool and handled 75% more requests under real Kimi workloads. NVIDIA's Dynamo 1.0 KV Block Manager reaches S3 and Azure blob storage and can pin blocks in host memory instead of deleting them. Moving a cache costs bandwidth: an 8K-token prompt on a 70B model is 2.68 GB of BF16 cache, about 54 ms over 400 Gb/s RDMA and about 3 ms over NVLink, while DeepSeek-V3's MLA cache for the same prompt is 0.58 GB (the book's arithmetic).

Providers price the difference. DeepSeek's API charges $0.15 per million input tokens on a cache miss and $0.003 on a cache hit for V4.1-Flash off-peak, as of October 2026, a 50x gap, and Anthropic bills cache reads at 0.1x the base input price (0.05x on Claude Opus 5.5). An agent that resends 100K tokens of context every turn pays DeepSeek $0.015 per turn on a miss and $0.0003 on a hit.

Related question · Inference

How would you serve 1M-token contexts for an agent product?

What it tests: Whether you turn KV arithmetic into a model choice, cache tiers and a price per turn instead of asking for bigger GPUs.

A strong answer, in two minutes

Start from the arithmetic. One 1M-token sequence of Llama 3.3 70B would need 320 GiB of BF16 cache, more than twice an H200's memory, so dense attention is ruled out before any tuning. Pick an architecture built for the length: DeepSeek reports that V4-Pro needs 10% of V3.2's KV cache at 1M tokens, and V4.1-Flash stores 890 bytes per token, about 0.9 GiB for a full 1M context. Then make reuse cheap. Agents resend long shared context every turn, so prefix caching, cache-aware routing and offload tiers keep that context out of prefill. Finally, price it per turn, where hits dominate: on DeepSeek's API a cached input token costs a fiftieth of an uncached one for V4.1-Flash off-peak.

Follow-ups

What does a cache hit save per turn at 100K tokens of context?
On DeepSeek V4.1-Flash off-peak, $0.015 per turn on a miss against $0.0003 on a hit, 50x, by the book's arithmetic on the October 2026 price sheet.
Where do offloaded blocks go?
CPU memory and SSD through connectors such as vLLM's OffloadingConnector and LMCache, or a cluster-wide pool such as Mooncake's; NVIDIA's Dynamo KV Block Manager also reaches S3 and Azure blob storage.

Follow-ups on Question 2

Compute the KV cache for an MLA model.
Use kv_lora_rank + qk_rope_head_dim per layer, independent of the head count. DeepSeek-V3: (512 + 64) × 61 × 2 bytes = 70,272 bytes, about 68.6 KiB per token in BF16. A plain multi-head cache with the same 128 heads would be about 4.77 MiB, 71 times more.
Why not cut every model to one KV head?
Quality. The GQA paper found that multi-head checkpoints can be uptrained to multi-query attention with 5% of the original pre-training compute, and that a few KV heads get close to multi-head quality at multi-query speed. Eight KV heads is the common compromise in the configs of Section 6.
What changes at 1M tokens?
Dense attention stops being an option: Llama 3.3 70B would need 320 GiB of BF16 cache for one sequence. The 2026 answers are architectural. DeepSeek reports that V4-Pro needs 10% of V3.2's KV cache at 1M tokens, and V4.1-Flash stores 890 bytes per token.
Can prefix caching leak data between tenants?
Response times can reveal whether a prefix was cached. vLLM's per-request cache_salt goes into the hash of the first block, so only requests with the same salt share cached blocks; give each tenant its own random salt.
Is an FP8 cache free?
It halves the bytes, but without calibration vLLM sets every scale to 1.0. The docs recommend calibrating with llm-compressor, and --kv-cache-dtype-skip-layers keeps sensitive layers in their native dtype. Evaluate on your own task before and after.

Red flags

  • Forgetting the factor of two for K and V, or multiplying by query heads instead of KV heads.
  • Treating MLA like GQA, or saying the cache scales with parameter count.
  • Claiming that quantizing the weights shrinks the KV cache.
  • Mixing up capacity (gigabytes of cache held) with bandwidth (bytes read on every step).
  • Equating the KV cache with prompt caching: every request has a KV cache, and reuse across requests is a separate layer on top.

Chapter 2 in one page

Key takeaways
8 items
  • 1Prefill is compute-bound and sets time to first token; decode is bandwidth-bound and sets inter-token latency. Llama 3.3 70B in FP8 on one H200 cannot decode faster than about 14.6 ms per token.
  • 2Batching works because one read of the weights serves every sequence in the step. It stops paying when the KV cache fills memory or compute catches up.
  • 3Continuous batching re-forms the batch every iteration. Quote its gains with their baselines: 36.9x over FasterTransformer (Orca), up to 23x over naive static batching (Anyscale), 2-4x over Orca and FasterTransformer (vLLM).
  • 4In vLLM V1, max_num_batched_tokens trades inter-token latency against TTFT, max_num_seqs caps the batch, preemption recomputes by default, and the priority policy runs lower values first.
  • 5KV bytes per token = 2 × layers × KV heads × head_dim × bytes; MLA uses (kv_lora_rank + qk_rope_head_dim) × layers × bytes; hybrid models count only the layers whose cache grows.
  • 6Llama 3.3 70B: 320 KiB per token, 2.5 GiB at 8K, 40 GiB at 128K. On one H200 with FP8 weights, about 20 sequences of 8K fit with a BF16 cache and about 41 with FP8.
  • 7Paging cut KV waste from 60-80% to under 4%. Prefix caching reuses full blocks only, one changed early token misses every block after it, and a per-tenant cache_salt closes the timing channel.
  • 8The cache shrinks fastest through architecture (GQA, MLA, hybrid and compressed attention; DeepSeek's cache per token fell 437x from V1 to V4.1-Flash), then through FP8 or FP4 storage and offload.

What to do on Monday: pick the model you serve, or the one you expect to be asked about, run kv_calc.py on its config.json, and write down three numbers: bytes per token, GiB per sequence at your p95 context, and the sequences that fit on your GPU after weights and activations. Then compare the last number with your server's --max-num-seqs and with last week's peak of vllm:num_requests_running, and lower the cap if it sits far above what fits.

Chapters 3 to 14: Serving, Retrieval, Agents, Security, Evaluation, Operations and the Rounds

Unlock autoscaling, compression and latency diagnosis, retrieval up to 50M patient records, agents under budget with LangGraph, support and voice agents, MCP access and guardrails, regulated decisions and bias testing, evaluation and judges, rollbacks and observability, data pipelines and Kafka, the coding round, the case and design round, the behavioral round, and the extended question bank with a six-week plan.

Join The Agent Foundry to unlock chapters 3 to 14 (autoscaling, compression and latency, retrieval up to 50M patient records, agents under budget, support and voice agents, MCP access and guardrails, regulated decisions, evaluation, operations, data pipelines, and the coding, case and behavioral rounds), the extended question bank with a six-week plan, and every future book on release.

Enter your email to continue. We'll send a one-click sign-in link and bring you back here.