Claude Certified Architect

Search the study guides

Contents

Claude Certified Developer – Foundations

Study guide for exam CCDV-F

Based on the official Developer – Foundations Exam Guide, version 1.0, effective July 2026. Anthropic marks the guide as subject to change, so confirm the current blueprint before scheduling.

The Developer – Foundations track tests whether you can build, integrate, and ship production-grade applications, agents, and workflows on Anthropic's Claude platform at a foundational level. It sits below the Architect – Professional exam in scope: it rewards hands-on mechanics — the API, the agent loop, tool schemas, streaming, context engineering, model and cost selection, debugging, security, and MCP — more than architecture or stakeholder negotiation.

This guide turns the official blueprint into a practical preparation plan. It does not reproduce or predict live exam content. The guide prose is Ravn-authored; the practice exam it points to is built from collected practice questions. Neither is real exam content, and no third party can reproduce the live item bank.

Exam at a glance

ParameterOfficial detail
CredentialClaude Certified Developer – Foundations
Exam codeCCDV-F
Items53
Item formatMultiple-choice and multiple-response; each item states how many responses to select
Time limit120 minutes
DeliveryProctored, online or at a Pearson VUE test center, per program policy
Passing score720 on a 100–1,000 scaled score
Exam fee$125 USD list; the amount at checkout reflects any partner-tier discount
Credential validity12 months from the date the credential is awarded
PrerequisitesNone. No course is required; the credential is awarded on exam performance alone
Score reportPass/fail, scaled score, and percent correct by domain

The schedule allows about 2 minutes 16 seconds per item. That is a pacing signal, not a target for every question: direct knowledge checks should leave time for multi-step scenarios and the multiple-response items that ask for a specific number of selections.

Registration starts on the Anthropic Partner Academy certification page. Scheduling and delivery then move to Pearson VUE.

Scoring

The exam is criterion-referenced: you are measured against a fixed performance standard, not against other candidates. The 720 cut score came from a formal standard-setting study in which subject-matter experts judged the performance expected of a minimally qualified candidate.

The score report shows percent correct per domain. Those section percentages do not determine pass or fail — only the total scaled score does. Use them to direct a retake, not to predict one.

Policies worth knowing before you register

PolicyDetail
IdentificationA valid, unexpired, government-issued photo ID whose name matches your registration exactly. Correct the name through certifications-support@anthropic.com before scheduling
AttemptsUp to 4 per rolling 12-month period, per exam; the fee applies to each attempt
Retake waits14 days after a first failure, 30 days after a second, 90 days after a third
DeliveryPearson VUE, online proctored or at a test center; closed book
LanguageEnglish only. Browser translation tools are prohibited during proctored testing
AccommodationsRequest through Pearson VUE and be approved before you schedule
RescheduleFree up to 24 hours before the appointment; changes inside 24 hours, or a no-show, forfeit the fee
ConfidentialityYou must accept a non-disclosure agreement before the exam begins; declining ends the session with no refund
RenewalReview what changed and complete a complimentary non-proctored assessment on time; a lapsed credential means retaking the full exam at full fee

Online proctoring needs Pearson's domains allowed on your network. If your device is locked down, a test center is the safer choice.

Who the exam is for

Anthropic targets technical professionals who build, integrate, and ship production-grade AI solutions on Claude. The intended audience is AI and ML engineers, technical leads, and senior software engineers at the intersection of business requirements and implementation. The minimally qualified candidate profile recommends:

These are recommendations, not prerequisites. The credential is not intended for non-technical or casual users, or for roles limited to prompt writing without broader application responsibility. Exam performance alone determines certification.

The eight-domain blueprint

DomainWeightWhat the exam expects
1. Agents and Workflows14.7%Build Claude agents and workflows with the Agent SDK, custom loops, and frameworks; decide workflow versus agent; use subagents and memory.
2. Applications and Integration33.1%Integrate Claude through the API and SDKs with streaming, batch, vision, and caching; apply software-engineering foundations; design applications and manage configuration.
3. Claude Code3.1%Operate Claude Code: Rules, Skills, Commands, Agents, memory, slash commands, modes, CLAUDE.md hierarchy, and settings.json.
4. Eval, Testing, and Debugging2.6%Identify error types, choose recovery strategies, and isolate failure origin between the integration layer and model output through trace analysis.
5. Model Selection and Optimization16.8%Reason about LLM fundamentals, model tiers and tradeoffs, technical substrate, and token/cost management including caching and batch.
6. Prompt and Context Engineering11.0%Engineer prompts and context: instruction clarity, few-shot, placement, output constraints, context drift and compaction, and structured output handling.
7. Security and Safety8.1%Apply secure-by-design principles, prompt-injection defense, guardrails and hooks, secrets and identity management.
8. Tools and MCPs10.6%Implement tools and function calling, build MCP servers, and choose among built-in tools, custom tools, Skills, and MCPs.

Take these eight domains and their weights from the official exam guide only. The exam is deliberately lopsided: Domain 2 alone is a third of the exam, and Domains 2 and 5 together are half. Domains 3 and 4 together are under 6%, so invest there lightly. The credential's purpose statement also lists designing and running evals as a capability, but on this blueprint evals sit inside the lightly weighted Domain 4 alongside debugging — do not over-invest in eval design the way the Architect track does.

Domain guidance

1. Agents and Workflows — 14.7%

Official objectives. The guide groups this domain into three skills:

Decide workflow versus agent before you write the first line, because the wrong choice is the most expensive early mistake and the easiest to defend in a demo.

Choose a workflow when…Choose an agent when…
You can enumerate the exact steps in codeYou can state the goal and the tools but not the path
Error cost is real and step-level guardrails matterThe path cannot be enumerated in advance
Standard observability is requiredNon-determinism is acceptable and actions are bounded by the toolset
Inputs are well-constrained to a known setUser inputs vary unpredictably in content and structure

The safe default is the cheapest pattern that survives your inputs: a single API call, then a workflow, then an agent. Promote one tier only when the current one cannot absorb the input variety you actually see. Mismatch the other way and you pay for it — a workflow chosen where an agent was needed collapses on the first input that leaves the coded path, and an agent chosen where a workflow suffices buys nondeterminism you never use.

The agent is the pattern; the wiring path is an implementation choice. Once you decide the task needs an agent, the pattern is constant: a loop that calls tools, manages context, and runs until a goal is met. Three wiring paths differ in how much runtime you own.

PathWho runs the loopWhat you ownChoose when
Raw loop against the Messages APIYour code, every iterationFull loop, tool execution, context, retries, exit conditionsYou need full control or are learning the loop
Agent SDKThe SDK, inside your processTool execution and the surrounding applicationYou want the loop, context handling, and tool scaffolding in your own Python or TypeScript environment
Claude Managed AgentsAnthropic runs the loop and sandboxThe application layer and a versioned agent definitionLong-running execution in minutes or hours, or you want a managed sandbox and no loop to build

Set settingSources explicitly in the Agent SDK rather than relying on a default: it decides whether filesystem-based configuration such as CLAUDE.md and Skills load into the agent. Managed Agents sessions are stateful and stored server-side, which currently rules them out for Zero Data Retention or HIPAA BAA workloads no matter how well they fit operationally — the governing constraint picks the path before convenience gets a say.

Wire the loop the same four ways on every path. Register tools, set a scoped system prompt, handle the tool-use loop by returning a tool_result block for every tool_use block the model issues, and define explicit exit conditions. A loop without exit conditions keeps requesting tool calls past what the task needs. Every tool_use block in an assistant turn must be answered by a matching-ID tool_result in the immediately following user turn; mismatched or missing IDs fail request validation, and no prompt fixes that.

Place the human-in-the-loop gate by worst-case cost. The question that decides where to insert a checkpoint is: what is the worst outcome if this step runs without a person checking it?

Insertion pointWhat it addresses
Before a destructive tool callIrreversible writes, deletes, or sends
After a planning stepA wrong plan executed correctly still produces the wrong outcome
On unexpected outputError flags, empty results, or out-of-bound values that retry alone will not fix

The characteristic failure is a loop that worked in a scratch directory and then edited a production file: validation passed on the file the agent edited, but no checkpoint sat between "validation passed" and "write committed to the live environment." If a tool can take an irreversible action in production, it needs a checkpoint before it runs, registered at design time.

Pick the memory scope at design time, not under production pressure. Four scopes trade cost against continuity.

ScopeWhat persistsWhen it fits
In-contextState in the active conversation, survives turnsShort sessions that fit the window and need no cross-session continuity
External storageState written to a database, read at session startState that must survive across sessions, users, or agent instances
Summarized memoryA condensed prior conversation injected at startLong-running dialogue where full history would outgrow the budget
StatelessNothing; each session independentSelf-contained jobs that finish and close out

Too much in-context state inflates every call because the model re-reads the full history each turn; too little persistent state strips memory across sessions. Measure expected state size — history plus system prompt plus tool schemas — against the window before choosing in-context as the default. Tool outputs in production are often three to five times larger than test fixtures, so a window that held twenty turns in development can fill at turn eight in production.

Be able to:

Practice artifact: take a multi-step task and write three designs for it — a workflow, a single agent, and an orchestrator-worker — then state which you would ship and the constraint that decides it.

2. Applications and Integration — 33.1%

Official objectives. The largest domain splits into six skills:

A third of the exam lives here, so weight your study accordingly. The unifying skill is turning a business requirement into running, configured, production-shaped code against the API.

Capture requirements before you wire. A business problem yields two kinds of requirement. Functional requirements name what the system must do, split across the model, existing systems, and people. Infrastructure requirements name where it runs: which cloud, which region, which identity model, which compliance posture. The expensive failure is sketching the architecture mid-discovery; a stakeholder who sees a confident design assumes the questions were already asked, and the questions stop. Write the constraint down — a residency rule, a retention rule, an approval gate — because an undocumented constraint becomes an infeasible system the moment production violates it.

Choose the request shape by who is waiting. The API offers three response shapes that solve different problems.

ShapeWhen it fits
Synchronous (one complete response)Short responses and backend jobs where no one is waiting
Streaming (server-sent events, pieces as the model generates)Long responses or a user watching, so output appears immediately instead of after a blank-screen wait
Asynchronous client (async/await in the SDK)You need concurrency without blocking your application thread; the request still returns in real time
Message Batches APIBulk offline workloads where no user is waiting; submits a large set, returns within a 24-hour window at lower per-token cost

Streaming and batch never compete for the same request: a request is either user-facing or it is not. Streaming changes how latency is perceived; batching changes the bill. On the newest models, sampling parameters such as temperature, top_p, and top_k return a 400 error and behavior is steered through prompting — confirm current parameter support in the API reference at build time.

Handle a stream without corrupting state. Acting on a half-built block is the bug to avoid. SSE chunks a tool_use block across several delta events; the block is whole only after the stream closes. Buffer the deltas by index, rebuild the tool call from the finished block, and only then execute it. Touch a partial tool_use and you feed malformed arguments downstream with no error signal. A dropped stream is not data to salvage — treat it as transient, retry the request in full, and discard anything you buffered.

Match multimodal input to its cost. Every image and PDF consumes context budget before the model reads a character of your prompt. A base64 payload inflates request size on every call, so a one-off image is fine but the same image sent repeatedly should be passed as a URL. Calculate image token cost before you commit a multimodal pipeline.

Apply software-engineering foundations as you would for any production service. REST and JSON are the substrate; the SDK is a thin convenience layer over the same REST API that handles authentication, request construction, retries, and parsing. Asynchronous programming keeps your application responsive while calls are in flight. Version control, code review, SDLC integration, and staged refactoring are scored as engineering discipline, not as Claude-specific trivia — the exam treats a Claude application as a system that must survive its own lifecycle.

Design the application around how Claude reads instructions. The same instruction is interpreted differently across interfaces — Claude Code, Desktop, claude.ai, the API, and SDKs — so a prompt that works in one surface is a draft for another, not a finished artifact. Hold content boundaries explicitly: separate trusted instructions from untrusted content the agent fetches. Keep session hygiene by trimming or summarizing history before each call; the window is a ceiling, not a target. Design schemas for the outputs you consume, and manage plugins as versioned dependencies rather than ad hoc copies.

Configuration is a production artifact, not a convenience. Five configuration surfaces recur, and each is a governance choice.

SurfaceWhat it controlsWhere it lives
CLAUDE.mdProject instructions loaded into every sessionRepo root
settings.jsonPermission mode, allow and deny rules, hooksUser, project, local, and enterprise levels
Model version pinWhich model snapshot the code callsIn code, by full model ID
Prompt versioningWhich prompt the code sendsAlongside the code
Plugin dependenciesWhich packaged capabilities are installedVersioned and tracked

Pin the model version, not the alias. A model alias such as Opus or Sonnet is convenient but resolves to a moving target; an upstream update then becomes a silent production change with nothing to roll back to. Pin the full model ID to a fixed snapshot, keep the prior version available, and gate promotion through your eval. The same discipline applies to prompts and plugins: version them alongside the code so a regression is a rollback rather than a hotfix.

Be able to:

Practice artifact: write the configuration files for one Claude application — a pinned model ID, a scoped settings.json with a deny rule, a CLAUDE.md held to the rules that change behavior, and a plugin declared as a versioned dependency — and state where each file lives and who can override it.

3. Claude Code — 3.1%

Official objectives. One skill:

At roughly one and a half scored items, this domain is worth an evening, not a week. Know the names and what each does; do not chase depth.

Claude Code runs the same agent loop in your terminal and adds a permission layer that gates every action. It works through a task in three phases — explore, plan, code — and plan mode holds it in the read-only explore phase, blocking edits and shell commands until you release a plan.

Permission modes trade speed against oversight. Default prompts before nearly every edit or command and is the baseline for any new or unfamiliar codebase. acceptEdits auto-approves reads, file edits, and common filesystem commands inside the working directory but still gates other shell commands and writes outside it. plan mode reads and proposes only. A bypass mode silences every confirmation prompt and also skips the protected-path guard the other modes keep, so it belongs only inside an isolated container. An allow-list mode pre-approves a named tool set and auto-denies everything else, built for locked-down CI.

Configuration layers by scope, and deny wins. User settings apply to every project on the machine; project settings apply to everyone who clones the repo; local settings are personal overrides that are git-ignored; enterprise managed settings cannot be overridden by users or projects. A deny rule always wins over an allow rule regardless of mode, and an enterprise-level deny is the most durable governance control because no individual developer can remove it.

Keep CLAUDE.md short. It loads into every session unconditionally, so every line you add reduces the weight of every other line. A file that grows to hundreds of lines dilutes the one rule that catches a real mistake. Hold it to the constraints that change behavior, and move path-specific guidance into rules files scoped by a paths glob in their frontmatter. Skills are the third mechanism: a SKILL.md file loads only when a request matches its description, so task-specific expertise inflates only the sessions that need it.

Be able to:

Practice artifact: assemble a settings.json for a trusted local refactor that auto-approves edits, never runs destructive shell commands, and denies reads to .env.production, then state where the human gate sits for a change to a deployment config.

4. Eval, Testing, and Debugging — 2.6%

Official objectives. One skill:

Despite the domain name, the single scored skill is debugging and error handling, not eval design. Treat evals as supporting material, not as the center of this domain.

Every failure starts with one question: is it retriable or terminal? If waiting and retrying the identical request could plausibly work, the failure is retriable; if not, it is terminal. On the Anthropic API, the status code tells you the bucket.

StatusClassHandling
429 (rate limit)RetriableExponential backoff with jitter, honor retry-after, cap attempts
529 (overloaded)RetriableBackoff; reflects Anthropic-side load, not a rate-limit signal
5xx server errors, 504 timeoutRetriableAnthropic-side faults that typically clear on retry
400 (bad request)TerminalNo retry; the identical request will fail again — fix or reject the input
401 (auth failure), 403 (forbidden)TerminalNo retry; a permissions or auth problem that time cannot fix
Tool result errorDepends on the causeReturn it to Claude with is_error: true so the model can react
Refusal (200, stop_reason refusal)TerminalNo retry; the model made a content decision. Raise it and log it

Retrying a terminal error wastes the retry budget and hides the real problem behind a wall of identical failures. When unsure, the safe default is to treat an error as terminal and raise it: a failure wrongly classed as terminal fails loudly and gets fixed, while one wrongly classed as retriable hammers a service.

Know what the SDK already retries before you write your own. The Anthropic client libraries retry transient failures with progressive delays up to a configurable cap. Two retry loops wrapped around the same call multiply attempts against a rate limit rather than capping them, so decide where retry lives: let the SDK handle transient cases and reserve your code for application-specific fallback, or turn the SDK down and own the whole path. Honor retry-after when present; treat your own backoff as the fallback when it is not.

Tool errors must come back to Claude explicitly, never silenced. Swallowing a tool failure is what manufactures a confident, wrong answer: the model reads an empty result as legitimate data and builds reasoning on it. Surface the failure instead — return it with is_error: true, which gives the model the signal to change tack, ask for clarification, or halt. An empty result is never neutral; silence is what turns a tool bug into a plausible-sounding response.

Use a trace to isolate the origin. Tests tell you a failure exists; a trace tells you which step produced it. Four test levels each catch a different break: unit tests isolate one function, functional tests check one Claude call's shape, integration tests exercise the handoff between two components, and end-to-end tests run the whole flow. Most silent production breaks live at the integration level, where each side passes its own tests while the handoff between them is broken. A trace turns "the case failed" into "step four raised a KeyError on a field the model did not return," which is the difference between a five-minute fix and a day spent tracing by hand.

Be able to:

Practice artifact: take a failing end-to-end run, read the trace, name the failing step, and classify the error as retriable or terminal with the handling that follows.

5. Model Selection and Optimization — 16.8%

Official objectives. Four skills:

This domain and Domain 2 together are half the exam. The fundamentals are the shared vocabulary the rest of the exam builds on.

Tokens are the unit of input, output, and cost. Claude reads tokens, not characters or words, and the characters-per-token average depends on the tokenizer and differs between model generations. Everything the model processes — prompt, history, tool definitions, tool results, and the response — is counted in tokens, and tokens are what the API bills and the context window measures. Think and budget in tokens, not words.

The context window is a fixed budget, not a free resource. Everything in a request shares it: system prompt, conversation history, documents, tool results, and the generated output. Two limits bite in different places. An input that already overflows fails validation before generation ever starts. An input that fits can still overflow while the model is generating — current models then halt and hand back what they produced so far, flagged with a model_context_window_exceeded stop reason, not an error. In neither case does the API quietly drop your oldest tokens; staying under the ceiling is your code's responsibility, which means trimming or summarizing history before every call.

Sampling makes generation non-deterministic. At each step the model produces a probability distribution over possible next tokens and samples from it; settings such as temperature shape that distribution. Because the choice is sampled, the same prompt run twice can return different wording even when both answers are correct. That changes how you test: a test that asserts the exact text of a response will be inconsistent, so assert on the property that must hold — a required field, a value in range, a structure that parses. On the newest Claude models, non-default sampling parameters are not accepted and behavior is steered through prompting; confirm current support at build time.

Separate model choice from reasoning mode. Two decisions, set independently: which model answers, and whether it thinks before it answers. On current models the thinking is adaptive — the model picks when and how deeply to reason, and you steer that depth with an effort setting, not a token allotment. The legacy budget_tokens field is deprecated and 400s on the newest generations. Spend the thinking budget where it pays:

ReasoningWorth itWasted
OnHard, multi-step problemsLookups, classification
OffA capable model, fast and directA small model that needs the extra reasoning to compete

Compose the levers: a strong model with thinking off is quick and straight; a smaller model with thinking on trades tokens for deliberation. The carry-back rule governs any turn that mixes reasoning with tools: every thinking block goes back to the API verbatim on the next turn. Its signature is the proof the reasoning was not altered; drop the block to save context and the following request fails.

Start at the balanced tier and move on evidence. The guide frames the model family around Opus, Sonnet, and Haiku. Sonnet is the balanced default for most production workloads; Haiku is built for speed and cost on tasks that fit its envelope; Opus handles demanding work above the Sonnet envelope. Move up only when an eval shows Sonnet missing your quality bar, and down only when an eval shows the quality drop is acceptable for the task — not merely to save cost. A model change is a release, gated by an eval against your cases, not a preference.

Guide-versus-course note. The exam guide's Domain 5 objective names three tiers — Opus, Sonnet, and Haiku. The current prep-course material describes a four-tier family that adds Fable as the most capable tier. The exam is written against the blueprint, so the three-tier framing is the exam-relevant one; the course's Fable tier reflects a lineup that has since evolved. Confirm the current family and identifiers at platform.claude.com before scheduling, and treat any tier count as version-sensitive.

Routing lets one system use more than one model. Before building a router, ask whether every request looks the same: if it does, pin one model and stop. The pattern earns its keep only when request types differ — send the bulk to a balanced default and divert the few categories that need a larger or smaller model, keyed off a cheap signal read from the request. The whole point is to pay for capability only on the calls that justify it.

Manage cost against the tail, not the average. A cost or latency problem almost always traces to a few measurable levers: model selection, prompt and context size, number of tool calls, and streaming versus batch. Instrument three metrics per call — token usage, latency, and error rate — from the start, so a cost spike becomes a row you can sort instead of a mystery on the invoice. Token distributions are usually skewed, so a cost model built on an average can understate spend by a wide margin.

Prompt caching reuses work already done on a stable prefix. The first request writes the prefix to a cache; follow-up requests that send identical content up to a marked breakpoint read from it at a fraction of the cost. The cache matches on an exact prefix, so a single changed character before the breakpoint invalidates it — which is why caching fits stable content such as a long system prompt or a large tool schema and works against anything that must reflect live state. The default lifetime is five minutes, refreshed on each read; a one-hour lifetime is available at additional cost. Caching only applies above a minimum token threshold, so short prompts see no benefit. Place the stable content first and the per-request content last: ordering is the mechanism, not configuration, and putting variable content at the top means the prefix changes every time and the cache never hits.

The Batch API trades latency for a lower bill. For non-urgent, high-volume work, submit a large set of requests and receive results within an asynchronous window at lower per-token cost. Batching and prompt caching compound when a scheduled job reuses the same context across many requests.

Be able to:

Practice artifact: build a cost model for one workload that names the model tier, the per-request token budget, the cacheable prefix, and whether each request is realtime or batch, then state which lever you would pull first if the bill tripled.

6. Prompt and Context Engineering — 11.0%

Official objectives. Three skills:

Context engineering is deciding in advance what enters the context window, what comes back out as a summary, and what never enters at all.

Four strategies keep a session in budget. Each loses a different kind of continuity.

StrategyWhat it doesWhat you lose
PruningJump back to an earlier message and continue from thereThe work done after the rewind point
CompactionSummarize history into a condensed version that preserves key informationDetails the summarizer did not capture
ClearingStart a new conversation with empty contextAll session context; persist what matters elsewhere
Subagent handoffDelegate a scoped task to an isolated context that returns a summaryVisibility into how the subagent reached its conclusion

What compaction preserves depends on how you write the summarizer. An under-specified summarizer drops task-critical state — which files were modified, which decision was made at a branch point, which error was resolved — and that loss is one of the most common sources of multi-session agent failure. Context isolation through subagents keeps per-turn cost low and makes long-horizon tasks tractable, but only worth the overhead where context cost is a real constraint; a short workflow does not need it.

Use the lightest prompting technique that clears the bar. Start with instruction alone. Add examples when showing the desired shape is easier than specifying it. Add explicit step-by-step reasoning only when the path genuinely determines the answer. Each step costs tokens and latency on every call.

ModeWhen it fits
Zero-shotThe task is simple and the output shape is obvious
One-shotA single example pins a structure a description keeps missing
Multi-shot (few-shot)The output has a specific structure, casing, or edge case that needs several examples

Model capability and prompting mode are the same optimization: a stronger model clears a task zero-shot that a smaller one only passes with a few examples. Run both down at once — start from the cheapest model and the fewest shots that satisfy your eval, then escalate capability or example count only where the eval demands it.

Diagnose before you re-prompt. Three failed revision passes is the cutoff for adding words: stop and ask which technique is absent instead. The trap is a prompt that swells rather than sharpens — six rounds, each longer than the last, none supplying the output constraint that was missing all along. Bloated prompts return bloated output and a latency hit for no accuracy gain. The fix is usually two lines: pin the exact format with an output constraint, and cover the ambiguous case with a single few-shot example.

Move output control from the prompt into the API when it must hold. A JSON-only instruction in the prompt passes the cases you tested and fails on the one you did not. Structured output closes that hole: pass the API a JSON schema and the grammar is enforced at generation, so a schema-violating response cannot be emitted. Two forms serve different surfaces — JSON outputs bind the final answer, and strict tool use validates the arguments bound for your tools before your code runs them. The cost is real and worth naming: the first request on a fresh schema is slower while the grammar compiles, an injected format prompt adds input tokens, and a guaranteed schema still does not guarantee success. A refusal or a max_tokens cut returns non-matching text, so your code reads stop_reason before it treats any response as parseable.

Treat confident output with skepticism. A model that produces fluent, confident content can still be wrong, with no error signal. Validate the response against the contract it must hold: parse defensively, check required fields, and do not trust confidence as a proxy for correctness. Validation confirms a value is the right shape; it cannot confirm it is the right value, which is why the eval — not the parser alone — guards the cases you did not think to test.

Be able to:

Practice artifact: take a classification prompt that returns the wrong shape, write the two-line fix (an output constraint and a covering few-shot example), then rewrite it with a JSON schema and name when the schema is worth its cost.

7. Security and Safety — 8.1%

Official objectives. Four skills:

Treat security as architecture, not as a prompt you add later. A rule that lives only in a prompt is a convention; a control that holds is one the system enforces and produces evidence for.

Prompt injection is the core threat for any agent that reads content it did not write. Trusting the human at the keyboard does not help, because the hostile instruction almost always hitches a ride inside content the agent retrieves — a fetched page, a document, a tool result — rather than arriving in the user's message. The model flattens its entire context into a single token stream with no native trusted/untrusted line, so buried instructions read as commands alongside your own prompt. The defense is structural: mark fetched and user-supplied content as data to inspect, never instructions to obey; fence it with delimiters; and gate any consequential action through constrain-and-log regardless of what that data says. No agent that ingests untrusted content is fully immune — the application has to hold the boundary as well.

Jailbreaks and prompt injections are different threats with the same defense. The two aim at different targets:

ThreatWhat it attacks
JailbreakThe model's own safety guardrails — get it to ignore them
Prompt injectionYour application's instructions — get it to follow the attacker's instead

One defense covers both: validate and constrain what reaches the model, and separately cap what the model is permitted to do as a consequence. Guarding the prompt while leaving the action open is the common gap — once the model is steered, the unconstrained action is where the damage lands.

Least privilege is the control that holds when every other defense fails. The same injection, run under two identities, lands as two different outcomes: with an identity that can write anywhere and read every secret, it is an incident; with an identity scoped to one output directory and its given input, it is a denied action and a line in the log. So scope every production agent to the narrowest permission set the task needs. The blast radius of a steered agent is bounded by what its identity can touch, which makes the auth configuration itself a target — anything that can rewrite it acts with that identity, so guard that configuration as tightly as the secret.

Secrets never travel with the configuration that references them. Once a key leaks, rotation is the only remedy — and a credential baked into committed code (a .mcp.json, a settings file, source) cannot be rotated, because overwriting the file never erases it from history. Keep the two apart: the config carries a variable reference, and the value is injected at execution time from an environment variable or a managed secret store. When several services share the value, a store centralizes it so one rotation reaches every consumer and every read is logged.

A hook is enforcement, not convention. A PreToolUse hook runs before a tool call executes and can exit with a non-zero code to block it, writing the reason to stderr the agent sees. A PostToolUse hook runs after the call and is the right place for automated side effects and audit logging. When multiple rules apply to the same action, the precedence is deny over ask over allow — a single deny rule blocks the action regardless of how many allow rules are present. The distinction that matters in a regulated environment: a rule in a prompt can be followed inconsistently, while a hook fires at every relevant tool call without exception. Use both layers — a CLAUDE.md instruction communicates intent, and a hook enforces it deterministically.

Screen at the right point in the path. Input screening decides whether a request reaches the model. Output screening decides whether a response reaches the user. Action authorization decides whether a side-effecting call may run. One filter at the far end of the path covers exactly one of these, and it sits downstream of the only step that cannot be undone. Choose the failure direction deliberately: a screening service that errors and passes traffic through gives the appearance of protection with none of its function, so where an unscreened request would cause harm, fail closed.

Be able to:

Practice artifact: write the minimal secure configuration for an agent that fetches untrusted web pages and writes to one protected path — a hook that blocks writes triggered by untrusted input, a deny rule on sensitive paths, a secret referenced by environment variable, and an audit-log line on every privileged action.

8. Tools and MCPs — 10.6%

Official objectives. Three skills:

Claude does not run your tools; it selects them and tells your code what to call. The tool-use loop has a boundary that is where most bugs live: your code defines a schema, the model reads it and decides whether and when to call the tool, your code executes the tool and returns a tool_result, and the model continues. If your application does not handle the return correctly, the model never gets the data it asked for and the loop breaks.

The schema is what drives selection. Tool selection is driven almost entirely by what you wrote in the schema — the name, the description, and the input schema. Write descriptions that name what the tool does and when to use it; a vague or overlapping description produces erratic routing. Too many tools with overlapping descriptions degrade selection quality as the surface grows, so start with the minimum set the task needs and add tools only when a specific gap is confirmed. Too few tools force the agent to hallucinate a path or return an incomplete result.

FailureFix
Wrong tool selectedThe description, not the model — name the task and when the tool applies
Malformed tool argumentsStrict tool use with an input schema, validated before your code runs
Tool error silencedReturn is_error: true so the model can react
Over-tooled agentRemove tools the task does not need; audit the set as you audit permissions

An MCP server separates tool definitions from any one application. Build the capability once as a process that exposes tools, resources, and prompts, and every MCP client that connects gets access without re-implementing the integration. Resources are read-only data the client fetches into context by address, useful when pulling it in directly is cheaper and more predictable than a tool call. Prompts are vetted instruction templates the client invokes by name, useful when specific wording produces materially better results than whatever a user would type.

Transport and scope are independent decisions that interact. Transport is how the client talks to the server; scope is who loads it.

TransportWhen it fits
stdioA local process on the same machine as the client — personal tools and dev servers
HTTPAny server that does not run locally — shared team servers and hosted integrations (recommended)
SSELegacy; superseded by HTTP, not recommended for new servers
ScopeWho loads itWhere it lives
LocalOnly you, one project~/.claude.json under the current project
UserYou, across all your projectsPersonal Claude settings
ProjectEveryone who clones the repo.mcp.json committed to the repo root
EnterpriseAll users, admin-controlledManaged settings that cannot be overridden

A stdio server cannot be project-scoped for sharing because it runs only on one machine, so match transport to where the server runs before choosing scope. Each connected server adds its tool definitions to the context pool, so connect only the servers a task needs.

Secrets in MCP configuration follow the same rule as everywhere else. An API key committed inline to .mcp.json enters repository history and cannot be removed by overwriting the file in a later commit; it must be treated as compromised and rotated. The configuration holds a variable reference, and the value lives in an environment variable. GitHub MCP authenticates with a personal access token passed as a header; Linear MCP uses an OAuth browser sign-in flow that issues and stores a token without anyone copying a secret. OAuth redirect URIs are registered per host, so a working staging connection does not mean production is configured — add the new host's redirect URI before moving an OAuth integration into a new environment.

Permission rules can target a single MCP tool, not the whole server. An MCP tool is identified as mcp__server__tool; an allow rule on one tool lets it run without a prompt while every other tool on the server still prompts, and a deny on one tool overrides an allow on the server. The API MCP connector adds an enabled flag per tool: the flag decides whether the model sees the tool at all, while a permission rule decides whether an exposed tool may run — a context-cost control and a governance control, often used together.

Choose among built-in tools, custom tools, Skills, and MCPs by reuse. Work down the list and stop at the first fit.

ApproachWhen it fits
Built-in toolThe platform already provides the capability you need
Custom tool, wired directlyOne application owns the integration and reuses nothing
Skill (SKILL.md)A reusable instruction set that loads on demand for a recurring task
MCP serverThe same tool surface must be reachable from more than one client and maintained independently

Hard-coding logic into a prompt is neither reusable nor maintainable; pasting live data into the context window gives no live access and wastes context; relying on a built-in tool to reach an arbitrary internal API does not work. The failure to recognize is a protocol carried forward out of habit: a shared MCP layer with exactly one client buys integration cost and no reuse, and the mirror-image failure is a developer-facing surface placed in front of users who are not developers.

Be able to:

Practice artifact: for one internal REST service, write two designs — a custom tool wired into one application and an MCP server shared across several — and state which you would ship and the reuse signal that decides it.

The official prep path

Anthropic publishes a Developer – Foundations prep path on the Partner Academy. It is not required — the exam guide states there is no single required course and that no resource guarantees a pass — but it is the only preparation material written against the developer material. The five modules and their run times:

#ModuleLengthPrimary domains
1MSO Foundations59 min5, 6
2Production-Grade Prompting, Agents & Tool Use209 min1, 2, 6, 8
3Claude Code, MCP & Integration142 min3, 8, 2, 7
4Production Engineering, Evals & Security211 min4, 5, 7, 1
5Accelerators & IP Contribution139 min2, 7

Total run time is about 12 hours 40 minutes.

Two things to notice. First, the course modules are organized by topic, not by exam domain, so no module maps one-to-one to a blueprint domain. Second, the module hours do not track the domain weights: Module 4 spends 211 minutes on evals, tracing, errors, cost, and security, yet Domain 4 (Eval, Testing, and Debugging) is only 2.6% and Domain 7 (Security and Safety) is 8.1%; meanwhile Domain 2 (Applications and Integration, 33.1%) is spread across Modules 2, 3, and 5 with no single home. Use the blueprint weights, not the module minutes, to budget your time.

Access: prep course path · all prep courses

What the official sample questions teach

Section 8 of the exam guide contains three sample items with full rationale. Read them in the source rather than a summary — the reasoning in the answer key is the most direct signal available about how items are constructed. What they demonstrate:

The shared pattern: the stem contains a discriminator — only cost matters and no user waits; the threat arrives through fetched content; the capability must be reusable across clients. The credited answer is the one that acts on it. Several distractors are defensible engineering practices that simply do not address what the stem describes.

Two habits follow. Read the stem for the variable that changed and the constraint that is binding before reading any option. Then check each attractive distractor against the discriminator: if it would be equally reasonable advice with the discriminator removed, it is almost certainly not the credited answer.

Ravn practice material

This study guide is Ravn-authored and is not real exam content. The practice items Ravn publishes for this track are collected practice questions; they rehearse the reasoning the blueprint rewards, and no third party can reproduce the live item bank. Treat a passing practice score as a readiness signal, never as a prediction.

The repository ships a browser practice exam for this blueprint at ccdf/dist/exam_en.html, drawing 53 questions per attempt across the eight CCDV-F domains. Do not substitute the Architect – Foundations exam (exam code CCAR-F, 60 items, five domains): its domain mix and item count do not represent Developer – Foundations coverage. For item-style calibration, the authoritative reference is the three sample questions in Section 8 of the official exam guide.

A four-phase preparation plan

Phase 1: Map your gaps

Copy the eight domains into a scorecard. Rate every domain from 0 to 3:

Calculate (3 - rating) × domain weight for each domain. Use the result to prioritize study time instead of reading every topic equally. Domain 2 dominates this scorecard; Domains 3 and 4 barely move it.

Phase 2: Build one reference application

Build or deeply review one end-to-end Claude application that includes:

One coherent application exposes cross-domain trade-offs better than eight disconnected demos.

Phase 3: Rehearse decisions

For each major design choice, practice this sequence:

  1. State the requirement and constraints.
  2. Name two or three viable options.
  3. Compare quality, latency, cost, security, and operability.
  4. Choose one and explain the evidence that would invalidate it.
  5. Describe the fallback and who owns it.

This mirrors the judgment the blueprint rewards.

Phase 4: Run timed items

Practice mixed-format items under the 120-minute limit. Include multiple-response items and force yourself to verify the requested number of selections. Review misses by failure type, not just topic: overlooked constraint, wrong request shape, misclassified error, misread discriminator, or a control placed at the wrong point in the path. Re-read the three official sample questions and their rationale before the run, because the answer key is the clearest available signal about how items are constructed.

Exam decision framework

When several options look plausible, prefer the answer that:

  1. Solves the stated requirement and respects every explicit constraint.
  2. Uses the least complex pattern that meets the need — a single call before a workflow before an agent.
  3. Matches the request shape to who is waiting — realtime or batch, streaming or synchronous.
  4. Removes unnecessary capability and privilege at the source.
  5. Places the control where the harm occurs — input, output, or action.
  6. Makes cost, latency, quality, and security trade-offs visible and measured.

Several tie-breakers recur often enough to be worth memorizing:

Avoid answers that rely on a larger model to solve an authorization or schema problem, add monitoring without reducing avoidable risk, optimize one metric while ignoring the stated service level, or introduce agentic complexity without a clear benefit. Be equally wary of answers that silence a tool error, commit a secret to a config file, retry a terminal error, or place a screening check after the irreversible action it was meant to prevent.

Readiness checklist

You are ready when you can do all of the following without notes:

Resources

Official — certification program

Everything in this guide's tables traces to one of these.

Official — product documentation

The exam guide directs candidates to the Claude API, models, prompt engineering, Claude Code, Skills, and MCP documentation.

Review the official exam guide again before registration. It contains the current policies for identification, accommodations, retakes, rescheduling, exam conduct, confidentiality, renewal, support, and appeals — and Anthropic marks it subject to change without notice.