Claude Certified ArchitectStudy Cheatsheet
Claude Certified Architect · Study Cheatsheet

Every question has one correct move and three traps.

The exam tests the same twelve principles over and over. Learn the principle, not the wording — the scenario changes, the underlying idea does not.

136
Questions
5
Domains
12
Principles
What the correct answer does What the traps do The domain it maps to

The 12 principles

Each card states the principle, what the correct answer does, the trap to avoid, and the domain(s) it maps to.

01

Surface conflicts; let the orchestrator or a human decide

D1Multi-agentSupport
Correct

When two credible sources disagree, do not silently choose one. Complete your portion, record both values with their provenance, flag the contradiction explicitly, and escalate the decision to the orchestrating agent or a human reviewer.

Trap

Quietly selecting one value, merging the figures without noting the conflict, halting the entire workflow, or inventing a resolution rule that was never defined.

7 exam questions that test this
  1. #1Multi-agent Research System

    A document analysis agent discovers that two credible sources contain directly contradictory statistics for a key metric: a government report states 40% growth, while an industry analysis states 12%. Both sources look credible, and the discrepancy could materially affect the research conclusions. How should the document analysis agent handle this situation most effectively?

    Which approach is most effective?

    CorrectD) Complete analysis with both numbers, explicitly annotate the conflict with source attribution, and let the coordinator decide how to reconcile the data before passing to synthesis.

    This approach preserves separation of responsibilities: the analysis agent completes its core work without blocking, preserves both conflicting values with clear attribution, and correctly passes reconciliation to the coordinator, which has broader context.

  2. #5Multi-agent Research System

    The web-search subagent returns results for only 3 of 5 requested source categories (competitor sites and industry reports succeed, but news archives and social feeds time out). The document analysis subagent successfully processes all provided documents. The synthesis subagent must produce a summary from mixed-quality upstream inputs. Which error-propagation strategy is most effective?

    Which error-propagation strategy is most effective?

    CorrectD) Structure the synthesis output with coverage annotations that indicate which conclusions are well-supported and where gaps exist due to unavailable sources.

    Coverage annotations implement graceful degradation with transparency, preserving value from completed work while propagating uncertainty to enable informed decisions about confidence.

  3. #20

    When researching "renewable energy adoption," the web search agent returns recent statistics (2024: 35% adoption) while the document analysis agent extracts data from internal reports (2022: 18% adoption). The synthesis agent incorrectly flags these as contradictory sources rather than recognizing the data shows growth over time.

    What change would best enable the synthesis agent to correctly interpret such temporal differences?

    CorrectA) Require subagents to include publication or data collection dates in their structured outputs.

    Correct. The synthesis agent misreads the data because it never sees the dates. Making each data point carry its own timestamp in the structured output lets synthesis reason about trends instead of contradictions.

  4. #159Customer Support Agent

    Your `get_customer` tool returns all matches when searching by name. Currently, when there are multiple results, Claude picks the customer with the most recent order, but production data shows this selects the wrong account 15% of the time for ambiguous matches. How should you address this?

    How should you address this?

    CorrectB) Instruct Claude to request an additional identifier (email, phone, or order number) when `get_customer` returns multiple matches before taking any customer-specific action.

    Asking the user for an additional identifier is the most reliable way to resolve ambiguity because the user has definitive knowledge of their identity. One extra conversational turn is a small price to pay to eliminate a 15% error rate caused by choosing the wrong account.

  5. #385

    Production reviews reveal inconsistent handling of uncertainty in final reports. Sometimes conflicting subagent findings are synthesized into a single confident statement (losing nuance), while other times reports over-hedge with excessive qualifications (becoming unhelpful). When the web search agent returns "industry analysts estimate $50B market size (methodology varies)" and the document analysis agent returns "peer-reviewed study estimates 35B(±7B, 95% CI)," the coordinator either picks one arbitrarily or produces vague statements like "the market may be 35B−50B depending on factors."

    What systematic approach best addresses this?

    CorrectC) Instruct the synthesis agent to structure reports with explicit sections distinguishing well-established findings from contested ones, preserving original source characterizations and methodological context.

    Correct. Report structure that keeps methodological context and separates settled vs. contested claims is how you get nuance without over-hedging.

  6. #390

    Your extraction pipeline processes contracts that frequently include amendments. When a contract contains both original terms and later amendments (e.g., original clause specifies "30-day payment terms" while Amendment 1 changes this to "45 days"), the model inconsistently extracts one value or the other with no indication of which applies.

    What's the most effective approach to improve extraction accuracy for documents with amendments?

    CorrectA) Redesign the schema so amended fields capture multiple values, each with source location and effective date.

    Correct. Amendments are structurally about versioned values. A schema with value + source location + effective date models the domain correctly and stops forcing the model to pick one.

  7. #498Conversational AI Architecture Patterns

    Over several turns discussing investment strategy, a user stated "I have a very low risk tolerance" and later "I want to maximize my returns." They now ask: "What should I invest in?"

    Which approach best ensures the recommendation aligns with the user's actual priority?

    CorrectA) Surface the contradiction and ask the user to clarify which matters more.

    When user preferences directly contradict each other, surfacing the conflict and asking for clarification is the only way to guarantee the recommendation aligns with the user's true intent. Any other approach involves making an assumption that may be wrong—maximizing returns and low risk tolerance are fundamentally incompatible goals that require a human decision.

02

Recover from transient failures yourself; escalate only when genuinely blocked

D5D1Multi-agent
Correct

For recoverable issues such as a timeout or a rate limit, retry with backoff on your own. Escalate to the orchestrator only when you cannot proceed, and include what failed and what you already attempted.

Trap

Escalating every minor hiccup to the orchestrator, or suppressing the failure and continuing as if nothing happened.

23 exam questions that test this
  1. #3Multi-agent Research System

    A document analysis subagent frequently fails when processing PDF files: some have corrupted sections that trigger parsing exceptions, others are password-protected, and sometimes the parsing library hangs on large files. Currently, any exception immediately terminates the subagent and returns an error to the coordinator, which must decide whether to retry, skip, or fail the whole task. This causes excessive coordinator involvement in routine error handling. What architectural improvement is most effective?

    Which improvement is most effective?

    CorrectD) Implement local recovery in the subagent for transient failures and escalate to the coordinator only errors it cannot resolve, including attempted steps and partial results.

    Handle errors at the lowest level capable of resolving them. Local recovery reduces coordinator workload while still escalating truly unrecoverable issues with full context and partial progress.

  2. #6Multi-agent Research System

    The document analysis subagent encounters a corrupted PDF file that it cannot parse. When designing the system’s error handling, what is the most effective way to handle this failure?

    Which approach is most effective?

    CorrectA) Return an error with context to the coordinator agent, allowing it to decide how to proceed.

    Returning an error with context to the coordinator is the most effective approach because it lets the coordinator make an informed decision—skip the file, try an alternative parsing method, or notify the user—while maintaining visibility into the failure.

  3. #8Multi-agent Research System

    The web-search subagent times out while researching a complex topic. You need to design how information about this failure is returned to the coordinator. Which error-propagation approach best enables intelligent recovery?

    Which error-propagation approach best enables intelligent recovery?

    CorrectA) Return structured error context to the coordinator including the failure type, the query executed, any partial results, and potential alternative approaches.

    Returning structured error context—including failure type, executed query, partial results, and alternative approaches—gives the coordinator everything needed to make intelligent recovery decisions (e.g., retry with a modified query or continue with partial results). It preserves maximum context for informed coordination-level decision-making.

  4. #10Multi-agent Research System

    During research, the web-search subagent queries three source categories with different outcomes: academic databases return 15 relevant papers, industry reports return “0 results,” and patent databases return “Connection timeout.” When designing error propagation to the coordinator, which approach enables the best recovery decisions?

    Which approach enables the best recovery decisions?

    CorrectD) Distinguish access failures (timeout) that require a retry decision from valid empty results (“0 results”) that represent successful queries.

    A timeout (access failure) and “0 results” (valid empty result) are semantically different outcomes requiring different responses. Distinguishing them allows the coordinator to retry the patent database while accepting the industry reports “0 results” as a valid, informative finding.

  5. #13Customer Support Agent

    After calling `get_customer` and `lookup_order`, the agent has all available system data but still faces uncertainty. Which situation is the most justified trigger for calling `escalate_to_human`?

    Which situation is most justified for escalation?

    CorrectC) A customer requests competitor price matching. Your policies allow price adjustments for price drops on your own site within 14 days, but say nothing about competitor prices. The agent should escalate for policy interpretation.

    This is a genuine policy gap: company rules cover price drops on your own site but do not address competitor price matching. The agent must not invent policy and should escalate for human judgment on how to interpret or extend existing rules.

  6. #15Customer Support Agent

    You are implementing the agent loop for your support agent. After each Claude API call, you must decide whether to continue the loop (run requested tools and call Claude again) or stop (present the final answer to the customer). What determines this decision?

    What determines this decision?

    CorrectA) Check the `stop_reason` field in Claude’s response—continue if it is `tool_use` and stop if it is `end_turn`.

    `stop_reason` is Claude’s explicit structured signal for loop control: `tool_use` indicates Claude wants to run a tool and receive results back, while `end_turn` indicates Claude has completed its response and the loop should end.

  7. #28

    Your codebase exploration tool stores session IDs to allow engineers to continue investigations across work sessions. An engineer spent an hour yesterday analyzing a legacy authentication module, building context about its architecture and dependencies. They want to continue today. The session ID is valid, but version control shows 3 of the 12 files the agent previously read were modified overnight by a teammate's merge.

    What approach best balances efficiency and accuracy?

    CorrectC) Resume the session and inform the agent which specific files changed for targeted re-analysis

    Correct. Keeps the expensive context you already built, while telling the agent exactly which 3 files to re-read — minimum waste, maximum accuracy.

  8. #32

    An engineer's exploration subagent spent 30 minutes analyzing a legacy payment system, reading 47 files and documenting data flows. The session was interrupted when the engineer's connection dropped. While away, a teammate merged a PR that renamed two utility functions. The engineer wants to continue the same exploration.

    What's the most effective approach?

    CorrectD) Resume the subagent from its previous transcript and inform it about the renamed functions.

    Correct. Keep the accumulated understanding, and give it a targeted delta about the renames so it can update its mental model — minimum waste, maximum accuracy.

  9. #36

    You're implementing the escalation logic for when the agent should call `escalate_to_human`. Your team proposes four different approaches for triggering escalation.

    Which approach will most reliably identify cases that genuinely require human intervention?

    CorrectA) Instruct the agent to escalate when the customer requests a human, when the issue requires policy exceptions, or when the agent cannot make meaningful progress.

    Correct. Escalation decisions are judgment calls about intent and progress — exactly what LLMs are good at. Clear criteria in natural language outperform rigid rules for the long tail.

  10. #37

    After investigating a billing dispute over 25+ turns, you've identified that duplicate charges occurred due to a payment gateway timeout triggering retry logic. The required refund ($847) exceeds your $500 authorization limit. You need to call `escalate_to_human`, and the human agent won't have access to your conversation transcript.

    What context should you pass to enable effective resolution?

    CorrectB) A structured summary: customer ID, root cause, refund amount, and recommended action.

    Correct. A structured handoff with identifiers, cause, amount, and recommended action is what a human agent needs to pick up the case instantly without re-investigating.

  11. #39

    During a billing dispute resolution, your agent successfully retrieves customer info via `get_customer` and order details via `lookup_order`, but when attempting to call `process_refund`, the tool returns a timeout error. The agent has enough information to explain the charges and verify refund eligibility, but cannot actually process the refund due to the backend failure.

    What approach best balances first-contact resolution with appropriate error handling?

    CorrectC) Explain the billing, confirm refund eligibility, acknowledge the system issue preventing immediate processing, and offer escalation or retry later

    Correct. Deliver the partial value you can (explanation + eligibility), be honest about the failure, and let the customer choose between human escalation or a retry. Classic graceful degradation.

  12. #40

    A customer writes: "I've been going back and forth on this return for days. I just want to speak to someone who can actually help me." The agent has confirmed via `lookup_order` that the return is straightforward—within policy and eligible for immediate processing.

    What should the agent do?

    CorrectA) Acknowledge frustration, inform them this is resolvable now, and offer to complete it or escalate

    Correct. Honor the feeling, give them the fast resolution path in writing, and preserve their choice. That's the customer-respect move that still leverages the agent's capability.

  13. #41

    When the agent calls `lookup_order` and receives order details showing the item was purchased 45 days ago, how does the agentic loop determine whether to call `process_refund` or `escalate_to_human` next?

    CorrectC) The order details are added to the conversation and the model reasons about which action to take.

    Correct. The agentic loop works by appending `tool_result` messages to the conversation and letting the model decide the next step on each turn. That's how 45 days → refund vs. escalate gets resolved.

  14. #42

    A customer sends: "This is frustrating. I've explained my issue twice and nothing is being resolved. I want to talk to a real person NOW." The agent has not yet called any tools to investigate their account.

    What should the agent do?

    CorrectA) Acknowledge the frustration and ask one targeted question to understand the specific issue before escalating.

    Correct. The customer has said 'twice' but you have no context yet. One acknowledging, focused question gives you a shot at first-contact resolution without dismissing the frustration or delaying a potential handoff.

  15. #43

    Your agent is handling a billing dispute. After calling `get_customer` and `lookup_order`, it identifies that the dispute involves a promotional pricing error requiring manager approval—beyond the agent's authorization level.

    How should the workflow handle this mid-process escalation?

    CorrectB) Compile a structured handoff with customer details, order info, and the identified issue before calling `escalate_to_human`.

    Correct. A structured brief (who, what order, what issue, why it exceeds auth) lets the human agent pick up instantly. That's the mid-process escalation pattern.

  16. #163Conversational AI Architecture Patterns

    Production monitoring shows your `search_catalog` tool fails 12% of the time: 8% are network timeouts that succeed when retried, and 4% are query syntax errors that never succeed regardless of retries. Currently both error types are returned identically, causing wasted retries.

    How should you modify the tool's error handling?

    CorrectC) Implement automatic retry with backoff for network timeouts inside the tool; return syntax errors immediately with parameter validation details.

    Handling retries at the tool level for transient errors is the correct abstraction boundary—the tool has definitive knowledge of the error type and can implement deterministic retry logic without relying on the agent to interpret a flag (D) or follow prompt-level instructions (A). Uniform backoff (B) wastes time on syntax errors that will never succeed.

  17. #169

    Production logs reveal inconsistent error handling: when `lookup_order` fails, the agent sometimes retries 5+ times (wasteful when the order ID doesn't exist), sometimes escalates immediately (premature for temporary network issues), and sometimes asks users for clarification (inappropriate when the issue is a backend permission error). Investigation shows your MCP tool returns uniform error responses: {"isError": true, "content": [{"type": "text", "text": "Operation failed"}]}. The agent cannot distinguish between error types.

    What's the most effective improvement?

    CorrectA) Enhance error responses with structured metadata: include errorCategory (transient/validation/permission), isRetryable boolean, and a description of what caused the failure.

    Correct. Give the agent the information it needs to make the right decision: category, retryability, and a human-readable cause. That replaces guessing with deterministic policy.

  18. #170

    When implementing your `lookup_order` MCP tool, the backend sometimes returns errors (e.g., "Order not found" or temporary database failures).

    What is the correct pattern for communicating these errors back to the agent?

    CorrectB) Return the error message in the tool result content with the isError flag set to true

    Correct. MCP's designed pattern: put the error text in the content field and mark isError=true. Claude sees both the failure flag and a readable message to reason about.

  19. #171

    Your `process_refund` tool returns two types of errors: technical errors ("503 Service Unavailable", "Connection timeout") that are transient (5% of calls), and business errors ("Order exceeds 30-day return window", "Item already refunded") that are permanent (12% of calls). Monitoring shows the agent wastes 3-4 turns retrying business errors that can never succeed. Currently, both error types return only a plain text message to Claude.

    What's the most effective way to reduce wasted retries while improving customer-facing response quality?

    CorrectA) Return structured error responses with retryable: false for business errors and a customer-friendly explanation for Claude to use.

    Correct. A retryable flag tells Claude deterministically 'don't retry,' and a ready-made customer-friendly message improves the outgoing reply. Fixes both problems at once.

  20. #391

    Your extraction system implements automatic retries when validation fails. On each retry, the specific validation error is appended to the prompt. This retry-with-error-feedback approach resolves most failures within 2-3 attempts.

    For which failure pattern would additional retries be LEAST effective?

    CorrectD) The model extracts "et al." for co-authors when the full list exists only in an external document not in the input

    Correct. No amount of retrying teaches the model information that isn't in the input. Retry-with-error-feedback only fixes mistakes the model could have gotten right from the source.

  21. #499Conversational AI Architecture Patterns

    Users refine playlist preferences over multiple conversation turns. Two messages after a user said "I love jazz," Claude asks "What genres do you enjoy?"

    What is the most likely cause?

    CorrectD) Your application isn't including prior messages in the `messages` array.

    Claude has no server-side memory—every API call is stateless. Without including the full conversation history in the `messages` array of each request, Claude has no knowledge of prior turns. Vector databases (A) and `session_id` (C) are not part of Claude's architecture; context window overflow (B) is impossible for two-message exchanges.

  22. #509

    The agent verifies customer identity through a multi-step process before resetting passwords. During testing, you notice that after the customer answers the third verification question, the agent asks them to provide their name again, as if the earlier exchange never happened.

    What's the most likely cause of this behavior?

    CorrectC) The conversation history isn't being passed in subsequent API requests.

    Correct. The API is stateless. Each request must include the full messages array. If you only send the latest turn, the model has no memory of earlier ones — exactly the 'ask for the name again' symptom.

  23. #512

    After your daily batch of 10,000 documents completes, 300 documents (3%) failed with "`context_length_exceeded`" errors. The results file identifies each failure by `custom_id`.

    What's the most cost-effective approach to process these failures?

    CorrectB) Resubmit only the 300 failed documents after chunking them into smaller pieces, then combine the partial extractions

    Correct. Targeted, and it addresses the actual cause (input too long). Chunk the oversized docs, extract per chunk, then merge — minimum tokens, fixes the specific failure mode.

03

Address the root cause, and scope each tool to least privilege

D2Multi-agent
Correct

If a tool repeatedly performs an unsafe action, constrain it to a narrower, safer capability rather than appending 'be careful' to the prompt. Grant each tool only the access it genuinely requires.

Trap

Leaving the over-powered tool unchanged and adding more warnings, or cleaning up the damage afterward instead of preventing it.

5 exam questions that test this
  1. #154Multi-agent Research System

    In your system design, you gave the document analysis agent access to a general-purpose tool `fetch_url` so it could download documents by URL. Production logs show this agent now frequently downloads search engine results pages to perform ad hoc web search—behavior that should be routed through the web-search agent—causing inconsistent results. Which fix is most effective?

    Which fix is most effective?

    CorrectA) Replace `fetch_url` with a `load_document` tool that validates that URLs point to document formats.

    Replacing a general-purpose tool with a document-specific tool that validates URLs against document formats fixes the root cause by constraining capability at the interface level. This follows the principle of least privilege, making undesired search behavior impossible rather than merely discouraged.

  2. #155Multi-agent Research System

    In testing, you observe that the synthesis agent often needs to verify specific claims while merging results. Currently, when verification is needed, the synthesis agent returns control to the coordinator, which calls the web-search agent and then re-invokes synthesis with the results. This adds 2–3 extra loops per task and increases latency by 40%. Your assessment shows 85% of these verifications are simple fact checks (dates, names, stats) and 15% require deeper research. Which approach most effectively reduces overhead while preserving system reliability?

    Which approach is most effective?

    CorrectD) Give the synthesis agent a limited-scope `verify_fact` tool for simple checks, while routing complex verifications through the coordinator to the web-search agent.

    A limited-scope fact-verification tool lets the synthesis agent handle 85% of simple checks directly, eliminating most loops, while preserving the coordinator delegation path for the 15% of complex verifications. This applies least privilege while significantly reducing latency.

  3. #161Customer Support Agent

    Production logs show the agent misinterprets outputs from your MCP tools: Unix timestamps from `get_customer`, ISO 8601 dates from `lookup_order`, and numeric status codes (1=pending, 2=shipped). Some tools are third-party MCP servers you cannot modify. Which approach to data format normalization is most maintainable?

    Which approach is most maintainable?

    CorrectA) Use a PostToolUse hook to intercept tool outputs and apply formatting transformations before the agent processes them.

    A PostToolUse hook provides a centralized, deterministic point to intercept and normalize all tool outputs—including third-party MCP server data—before the agent processes them. It’s more maintainable because transformations live in code and apply uniformly, rather than relying on LLM interpretation.

  4. #164

    The document analysis agent has a single `analyze_document` tool that takes a document and a free-text instruction parameter. During evaluation, requests like "extract the key financial metrics" often return narrative summaries, while "summarize the methodology" sometimes returns raw data tables. The synthesis agent reports that 35% of analysis results require re-requests with clarified instructions.

    What's the most effective way to improve reliability?

    CorrectA) Split the generic tool into purpose-specific tools—`extract_data_points`, `summarize_content`, `verify_claim_against_source`—each with defined input/output contracts.

    Correct. Free-text instructions put the semantics in prose, which the model interprets inconsistently. Purpose-specific tools give the model an explicit, well-typed contract to pick between.

  5. #168

    Your agent needs to insert a new helper function into the middle of a 150-line utility module, between two existing functions. The Edit tool fails because its `old_string` parameter cannot find unique text to match — the file has repetitive docstrings, variable names, and structural patterns.

    What's the most reliable way to complete this insertion?

    CorrectD) Use Read to load the file, add the function at the appropriate location, then Write the updated file

    Correct. When Edit's unique-match contract can't be satisfied in a repetitive file, fall back to Read → modify in memory at the intended line → Write the full file back.

04

When the wrong tool is selected, fix the names and descriptions first

D2Multi-agentSupport
Correct

An agent chooses a tool from its name and description. If it keeps selecting the wrong one, clarify those labels so each tool is distinct and unambiguous before changing anything else.

Trap

Blaming the model or re-architecting the system before simply disambiguating the tool definitions.

8 exam questions that test this
  1. #27

    The coordinator agent has `AgentDefinitions` configured for all four specialized subagents, each with appropriate descriptions, prompts, and tool restrictions. During testing, you notice the coordinator correctly reasons about when to delegate—it generates messages like "I'll ask the web search agent to find sources on this topic"—but no subagent execution ever occurs. The coordinator then proceeds as if the delegation happened and continues with incomplete information. Logs show no errors.

    What is the most likely cause?

    CorrectC) The coordinator's allowedTools configuration doesn't include "Task", so while it can reason about delegation, it cannot invoke the tool required to spawn subagents.

    Correct. Without the Task tool in allowedTools, the coordinator can talk about delegating but has no way to actually call a subagent — which matches the 'reasons about it, no execution, no errors' symptom.

  2. #153Multi-agent Research System

    Production logs show a persistent pattern: requests like “analyze the uploaded quarterly report” are routed to the web-search agent 45% of the time instead of the document analysis agent. Reviewing tool definitions, you find that the web-search agent has a tool `analyze_content` described as “analyzes content and extracts key information,” while the document analysis agent has a tool `analyze_document` described as “analyzes documents and extracts key information.” How should you fix the misrouting problem?

    How should you fix the misrouting problem?

    CorrectB) Rename the web-search tool to `extract_web_results` and update its description to “processes and returns information retrieved from web search and URLs.”

    Renaming the web-search tool to `extract_web_results` and updating its description to explicitly reference web search and URLs directly removes the root cause by eliminating semantic overlap between the two tool names and descriptions. This makes each tool’s purpose unambiguous, enabling the coordinator to reliably distinguish document analysis from web search.

  3. #157Customer Support Agent

    While testing, you notice the agent often calls `get_customer` when users ask about order status, even though `lookup_order` would be more appropriate. What should you check first to address this problem?

    What should you check first?

    CorrectD) Check the tool descriptions to ensure they clearly differentiate each tool’s purpose.

    Tool descriptions are the primary input the model uses to decide which tool to call. When an agent consistently picks the wrong tool, the first diagnostic step is to verify that tool descriptions clearly separate each tool’s purpose and usage boundaries.

  4. #160Customer Support Agent

    Production logs show the agent often calls `get_customer` when users ask about orders (e.g., “check my order #12345”) instead of calling `lookup_order`. Both tools have minimal descriptions (“Gets customer information” / “Gets order details”) and accept similar-looking identifier formats. What is the most effective first step to improve tool selection reliability?

    What is the most effective first step?

    CorrectD) Expand each tool’s description to include input formats, example queries, edge cases, and boundaries explaining when to use it versus similar tools.

    Expanding tool descriptions with input formats, example queries, edge cases, and clear boundaries directly fixes the root cause—minimal descriptions that don’t give the LLM enough information to distinguish similar tools. It’s a low-effort, high-impact first step that improves the primary mechanism the LLM uses for tool selection.

  5. #165

    After integrating a local MCP server providing code analysis tools (`analyze_dependencies`, `find_dead_code`, `calculate_complexity`), you verify the server is healthy and tools appear in the tools/list response. However, you observe that the agent consistently uses Grep to search for import statements instead of calling `analyze_dependencies`—even when users explicitly ask about "code dependencies." Examining tool definitions reveals: MCP: `analyze_dependencies` - "Analyzes dependency graph" Built-in: Grep - "Search file contents for a pattern using regular expressions. Returns matching lines with line numbers and surrounding context."

    What's the most effective approach to improve the agent's selection of MCP tools?

    CorrectD) Expand MCP tool descriptions to detail capabilities and outputs—e.g., "Builds dependency graph showing direct imports, transitive dependencies, and cycles."

    Correct. Tool selection is driven by the descriptions the model sees. A one-line description like 'Analyzes dependency graph' loses against Grep's rich description. Beef up the MCP tool's description.

  6. #167

    After adding an MCP server with specialized code refactoring tools (`extract_function`, `rename_variable`, `inline_function`), you notice the agent still uses basic text manipulation via Write and Bash sed commands for refactoring tasks. The MCP server is connected and healthy. Examining the configuration, you find each MCP tool has a minimal description like "`extract_function`: extracts a function from code."

    What's the most effective way to improve adoption of the MCP refactoring tools?

    CorrectD) Enhance the MCP tool descriptions to explain when each tool is preferable to text manipulation and clarify expected inputs and outputs.

    Correct. Tool selection is driven by the descriptions Claude sees. When the MCP tools say 'extracts a function from code' and Write/sed come with rich documentation, Claude picks Write/sed. Beef up the descriptions.

  7. #172

    Your pipeline uses a tool called `extract_metadata` with a JSON schema for paper details. You've also defined `lookup_citations` and `verify_doi` tools for enrichment. During testing, you notice that when users include requests like "extract the metadata and tell me how cited it is," Claude sometimes calls `lookup_citations` first, which fails because it needs the DOI that `extract_metadata` would provide.

    What's the most effective way to ensure structured metadata extraction happens first?

    CorrectC) Set `tool_choice` to {"type": "tool", "name": "`extract_metadata`"} and process the enrichment requests in subsequent turns after receiving the extracted metadata.

    Correct. `tool_choice`=specific-tool deterministically forces `extract_metadata` on the first turn. Then you hand control back to 'auto' to let the model use citations/DOI enrichment with the metadata in context.

  8. #377Customer Support Agent

    Production logs show a consistent pattern: when customers include the word “account” in their message (e.g., “I want to check my account for an order I made yesterday”), the agent calls `get_customer` first 78% of the time. When customers phrase similar requests without “account” (e.g., “I want to check an order I made yesterday”), it calls `lookup_order` first 93% of the time. Tool descriptions are clear and unambiguous. What is the most likely root cause of this discrepancy?

    What is the most likely root cause?

    CorrectA) The system prompt contains keyword-sensitive instructions that steer behavior based on terms like “account,” creating unintended tool-selection patterns.

    The systematic keyword-driven pattern (78% vs 93%) strongly indicates explicit routing logic in the system prompt reacting to the word “account” and steering the agent toward customer-related tools. Since tool descriptions are already clear, the discrepancy points to prompt-level instructions creating unintended behavioral steering.

05

Demonstrate with examples; don't just restate the instructions

D4CI/CDClaude CodeSupport
Correct

When instructions yield inconsistent output, add three to six concrete examples covering the tricky cases. Worked examples teach the desired behavior more reliably than additional prose rules.

Trap

Rewriting the same instructions at greater length and hoping for a different result.

11 exam questions that test this
  1. #11Customer Support Agent

    Your agent handles single-issue requests with 94% accuracy (e.g., “I need a refund for order #1234”). But when customers include multiple issues in one message (e.g., “I need a refund for order #1234 and also want to update the shipping address for order #5678”), tool selection accuracy drops to 58%. The agent usually solves only one issue or mixes parameters across requests. What approach most effectively improves reliability for multi-issue requests?

    What approach is most effective?

    CorrectC) Add few-shot examples to the prompt demonstrating correct reasoning and tool sequencing for multi-issue requests.

    Few-shot examples that demonstrate correct reasoning and tool sequencing for multi-issue requests are most effective because the agent already performs well on single issues—what it needs is guidance on the pattern for decomposing and routing multiple issues and keeping parameters separated.

  2. #364Claude Code for Continuous Integration

    Your automated reviews find real issues, but developers report the feedback is not actionable. Findings include phrases like “complex ticket routing logic” or “potential null pointer” without specifying what exactly to change. When you add detailed instructions like “always include concrete fix suggestions,” the model still produces inconsistent output—sometimes detailed, sometimes vague. Which prompting technique most reliably produces consistently actionable feedback?

    Which prompting technique is most reliable?

    CorrectD) Add 3–4 few-shot examples showing the exact required format: identified issue, location in code, concrete fix suggestion.

    Few-shot examples are the most effective technique for achieving consistent output format when instructions alone produce variable results. Providing 3–4 examples that show the exact desired structure (issue, location, concrete fix) gives the model a concrete pattern to follow, which is more reliable than abstract instructions.

  3. #374Code Generation with Claude Code

    You asked Claude Code to implement a function that transforms API responses into an internal normalized format. After two iterations, the output structure still doesn’t match expectations—some fields are nested differently and timestamps are formatted incorrectly. You described requirements in prose, but Claude interprets them differently each time.

    Which approach is most effective for the next iteration?

    CorrectB) Provide 2–3 concrete input-output examples showing the expected transformation for representative API responses.

    Concrete input-output examples remove ambiguity inherent in prose descriptions by showing Claude the exact expected transformation results. This directly addresses the root cause—misinterpretation of textual requirements—by providing unambiguous patterns for field nesting and timestamp formatting.

  4. #375Customer Support Agent

    Your agent achieves 55% first-contact resolution, well below the 80% target. Logs show it escalates simple cases (standard replacements for damaged goods with photo proof) while trying to handle complex situations requiring policy exceptions autonomously. What is the most effective way to improve escalation calibration?

    What is the most effective way to improve escalation calibration?

    CorrectC) Add explicit escalation criteria to the system prompt with few-shot examples showing when to escalate versus resolve autonomously.

    Explicit escalation criteria with few-shot examples directly address the root cause—unclear decision boundaries between simple and complex cases. It’s the most proportional, effective first intervention that teaches the agent when to escalate and when to resolve autonomously without extra infrastructure.

  5. #378Customer Support Agent

    Production logs show the agent sometimes chooses `get_customer` when `lookup_order` would be more appropriate, especially for ambiguous queries like “I need help with my recent purchase.” You decide to add few-shot examples to the system prompt to improve tool selection. Which approach most effectively addresses the problem?

    Which approach is most effective?

    CorrectC) Add 4–6 examples targeted at ambiguous scenarios, each with rationale for why one tool was chosen over plausible alternatives.

    Targeting few-shot examples at the specific ambiguous scenarios where errors occur, with explicit rationale for why one tool is preferable to alternatives, teaches the model the comparative decision process needed for edge cases. This is more effective than generic examples or declarative rules.

  6. #379Conversational AI Architecture Patterns

    During QA testing, Claude follows system prompt guidelines for the first 10–15 turns, but later responses deviate. The conversation is still within token limits.

    What is the best solution?

    CorrectC) Insert user-role messages reinforcing guidelines at conversation breakpoints.

    Periodic injection of behavioral reminders directly combats instruction drift by re-establishing constraints at regular intervals as conversation history accumulates. Moving guidelines to the first user message (A) reduces their authority. Starting a new conversation (B) destroys context. Post-response validation (D) is corrective rather than preventive and adds significant latency.

  7. #381Conversational AI Architecture Patterns

    Users report repetitive response openings like "Certainly!" and "I'd be happy to help!"

    What is the most effective approach?

    CorrectA) Append a partial assistant message with a direct response opening.

    Prefilling the assistant's response with the beginning of a direct answer prevents greeting patterns at the generation level—the model continues from the prefill rather than generating new opening phrases. System prompt instructions (D) can help but are less reliable since the model may still produce variants. Post-processing (C) is a fragile workaround. Temperature (B) controls randomness, not specific phrase patterns.

  8. #388

    Your schema includes a skills: string[] field. Production monitoring reveals three consistency issues: (1) compound phrases like "Python and SQL" are sometimes kept as one entry, sometimes split; (2) implied but unstated skills occasionally appear in extractions; (3) similar documents produce wildly different array lengths (5-10 vs 40+ entries). Your prompt currently says "Extract all skills mentioned."

    What's the most effective improvement?

    CorrectA) Add few-shot examples demonstrating compound phrase handling, explicit mention criteria, and appropriate entry granularity.

    Correct. All three issues are about the model's interpretation of what counts as 'a skill.' Few-shot examples teach the pattern concretely — split vs. not-split, mentioned vs. inferred, appropriate granularity.

  9. #394

    After implementing tool use with strict schema definitions, JSON syntax errors are eliminated, but 5% of extractions still have valid JSON with empty arrays or null values for required fields like citations and methodology. Spot-checking reveals that source documents contain this information, but in varied formats—inline citations vs. bibliographies, methodology sections vs. details embedded in introductions.

    What's the most effective way to address these failures?

    CorrectD) Add few-shot examples demonstrating extractions from documents with varied structures—showing how to identify citations in different formats and locate methodology details across section types.

    Correct. The failure mode is the model not recognizing varied formats. Concrete examples across the format distribution directly raise recall on the 5%.

  10. #397

    Your extraction system parses e-commerce product descriptions to extract specifications like dimensions, weight, and materials into JSON. Despite having a well-defined schema, the model inconsistently extracts the "materials" field—sometimes returning "cotton blend", other times "Cotton/Polyester mix", and occasionally omitting the field when material information is clearly present in the source.

    What's the most effective way to improve extraction consistency?

    CorrectD) Add few-shot examples showing 2-3 complete input-output pairs with standardized material description formats

    Correct. Few-shot examples demonstrate the exact canonical format you want ('cotton/polyester' as a normalized list), and they also raise recall on material info that was being skipped.

  11. #504Conversational AI Architecture Patterns

    Your AI tutor has a 2,800-token system prompt defining teaching methodology and adaptation rules. After 12 turns, the assistant starts ignoring proficiency levels.

    What is the most effective fix?

    CorrectB) Replace verbose rules with few-shot examples demonstrating proficiency-level adaptation.

    A 2,800-token system prompt with declarative rules is vulnerable to drift because abstract rules require the model to reason about them on every turn. Replacing verbose rules with concrete few-shot examples that demonstrate correct proficiency-level adaptation gives the model clear behavioral patterns to match—this is more reliably followed across many turns than abstract instructions. Reminder injection (A) helps but addresses symptoms; end-placement (C) helps initially but not with turn-level drift; regeneration (D) is expensive and corrective.

06

Define 'unacceptable' with precise, testable criteria

D4CI/CDSupport
Correct

Replace vague directives like 'flag bad comments' with an explicit rule, e.g. 'flag a comment only when it states the opposite of what the code does.' Make the boundary checkable.

Trap

Leaving the criterion subjective and hoping the model infers it, or correcting the mistakes downstream.

10 exam questions that test this
  1. #366Claude Code for Continuous Integration

    Your automated review analyzes comments and docstrings. The current prompt instructs Claude to “check that comments are accurate and up to date.” Findings often flag acceptable patterns (TODO markers, simple descriptions) while missing comments describing behavior the code no longer implements. What change addresses the root cause of this inconsistent analysis?

    What change addresses the root cause?

    CorrectD) Specify explicit criteria: flag comments only when the behavior they claim contradicts the code’s actual behavior.

    Explicit criteria—flagging comments only when claimed behavior contradicts actual code behavior—directly addresses the root cause by replacing a vague instruction with a precise definition of what constitutes a problem. This reduces false positives on acceptable patterns and misses of truly misleading comments.

  2. #367Claude Code for Continuous Integration

    Your automated code review system shows inconsistent severity ratings—similar issues like null pointer risks are rated “critical” in some PRs but only “medium” in others. Developer surveys show growing distrust—many start dismissing findings without reading because “half are wrong.” High-false-positive categories erode trust in accurate categories. Which approach best restores developer trust while improving the system?

    Which approach best restores developer trust?

    CorrectA) Temporarily disable high-false-positive categories (style, naming, documentation) and keep only high-precision categories while improving prompts.

    Temporarily disabling high-false-positive categories immediately stops trust erosion by removing noisy findings that cause developers to dismiss everything, while preserving value from high-precision categories like security and correctness. It also creates space to improve prompts for problematic categories before re-enabling them.

  3. #371Claude Code for Continuous Integration

    Your automated code review averages 15 findings per pull request, and developers report a 40% false-positive rate. The bottleneck is investigation time: developers must click into each finding to read Claude’s rationale before deciding whether to fix or dismiss it. Your CLAUDE.md already contains comprehensive rules for acceptable patterns, and stakeholders rejected any approach that filters findings before developers see them. What change best addresses investigation time?

    What change best addresses investigation time?

    CorrectA) Require Claude to include its rationale and confidence estimate directly in each finding.

    Including rationale and confidence directly in each finding reduces investigation time by letting developers quickly triage without opening each finding. It satisfies the “no filtering” constraint because all findings remain visible while accelerating developer decision-making.

  4. #372Claude Code for Continuous Integration

    Analysis of your automated code review shows large differences in false-positive rates by finding category: security/correctness findings have 8% false positives, performance findings 18%, style/naming findings 52%, and documentation findings 48%. Developer surveys show growing distrust—many start dismissing findings without reading because “half are wrong.” High-false-positive categories erode trust in accurate categories. Which approach best restores developer trust while improving the system?

    Which approach best restores developer trust?

    CorrectA) Temporarily disable high-false-positive categories (style, naming, documentation) and keep only high-precision categories while improving prompts.

    Temporarily disabling high-false-positive categories immediately stops trust erosion by removing noisy findings that cause developers to dismiss everything, while preserving value from high-precision categories like security and correctness. It also creates space to improve prompts for problematic categories before re-enabling them.

  5. #376Customer Support Agent

    Production metrics show that when resolving complex billing disputes or multi-order returns, customer satisfaction scores are 15% lower than for simple cases—even when the resolution is technically correct. Root-cause analysis shows the agent provides accurate solutions but inconsistently explains rationale: sometimes omitting relevant policy details, sometimes missing timeline info or next steps. The specific context gaps vary case by case. You want to improve solution quality without adding human oversight. What approach is most effective?

    What approach is most effective?

    CorrectA) Add a self-critique stage where the agent evaluates a draft response for completeness—ensuring it resolves the customer’s issue, includes relevant context, and anticipates follow-up questions.

    A self-critique stage (the evaluator-optimizer pattern) directly addresses inconsistent explanation completeness by forcing the agent to assess its own draft against concrete criteria—such as policy context, timelines, and next steps—before presenting it. This catches case-specific gaps without human oversight.

  6. #386

    A user is expanding the research system beyond its single web search agent by adding specialized data sources. They add a financial API agent that returns structured JSON with revenue, margins, and growth rates; a news monitoring agent that returns prose summaries of recent developments; and a patent analysis agent that returns structured lists of technology areas. The synthesis agent combines these into executive briefings. Currently, it converts everything to bullet points, causing financial comparisons to lose tabular clarity and news summaries to lose narrative flow.

    What change would most improve briefing quality?

    CorrectC) Update the synthesis agent to render each content type appropriately—financial data as tables, news as prose.

    Correct. Executive briefings need mixed rendering: tables for numbers, prose for narrative. Asking synthesis to preserve the native format of each input is the right abstraction.

  7. #389

    Your system has been operating with 100% human review for 3 months. Analysis shows that extractions with model confidence >90% have 97% accuracy overall. To reduce reviewer workload, you plan to automate high-confidence extractions. Before deploying, what validation step is most critical?

    CorrectA) Analyze accuracy by document type and field to verify high-confidence extractions perform consistently across all segments, not just in aggregate.

    Correct. Aggregate accuracy hides segment failures — one document type could be 70% while others are 99%. Auto-routing by overall number alone risks systemic errors in the weak segments.

  8. #392

    Your extraction pipeline processes restaurant menus and must output structured JSON with fields for item names, descriptions, prices, and dietary tags. Some menus use inconsistent formatting—prices as "$12" vs "12.00", dietary info as icons vs text.

    What's the most reliable approach?

    CorrectD) Define a strict output schema and include format normalization rules in your prompt.

    Correct. Strict schema + explicit normalization rules ('prices as decimal with two places', 'dietary as enumerated tags') lets the model both extract and normalize in one pass.

  9. #393

    Your system extracts event metadata (date, location, organizer, `attendee_count`) from news articles using a JSON schema with all nullable fields. During evaluation, you observe the model frequently generates plausible but incorrect values for fields not mentioned in the article—for example, outputting "500" for `attendee_count` when the source contains no attendance information.

    What's the most effective way to reduce these false extractions?

    CorrectB) Add prompt instructions to return null for any field where information is not directly stated in the source.

    Correct. The fields are already nullable; the model just needs an explicit instruction to prefer null over a plausible guess. This is the standard fix for schema-aware hallucination.

  10. #396

    Your extraction uses tool use with a JSON schema where `property_type` is defined as an enum: ['house', 'apartment', 'condo', 'townhouse']. After deployment, 8% of extractions fail schema validation. Investigation reveals listings mention many uncommon property types—"studio", "loft", "duplex", "mobile home", "tiny house", "converted warehouse"—and new types continue appearing regularly.

    What's the most effective long-term solution?

    CorrectB) Add an "other" value to your enum with a separate `property_type_detail` string field for specifics when "other" is selected.

    Correct. Keeps the strong enum for the common cases (clean downstream joins) while giving a well-typed escape hatch that preserves detail. Stable long-term.

07

Use the interactive API for urgent work and the Batch API for deferrable work

D5CI/CD
Correct

When a person is waiting (e.g., a pre-merge review), use the interactive, real-time path. For work that can finish later (e.g., overnight jobs), use the Batch API — slower but roughly half the cost.

Trap

Making users wait on batch jobs to save money, or paying interactive rates for work nobody is waiting on.

6 exam questions that test this
  1. #156Claude Code for Continuous Integration

    Your code review component is iterative: Claude analyzes the changed file, then may request related files (imports, base classes, tests) via tool calls to understand context before providing final feedback. Your application defines a tool that lets Claude request file contents; Claude calls the tool, gets results, and continues analysis. You’re evaluating batch processing to reduce API cost. What is the primary technical limitation when considering batch processing for this workflow?

    What is the primary technical limitation?

    CorrectB) The asynchronous model cannot execute tools mid-request and return results for Claude to continue analysis.

    A “fire-and-forget” asynchronous Batch API model has no mechanism to intercept a tool call during a request, execute the tool, and return results for Claude to continue analysis. This is fundamentally incompatible with iterative tool-calling workflows that require multiple tool request/response rounds within a single logical interaction.

  2. #363Claude Code for Continuous Integration

    Your CI/CD system runs three Claude-based analyses: (1) fast style checks on every PR that block merging until completion, (2) comprehensive weekly security audits of the entire codebase, and (3) nightly test-case generation for recently changed modules. The Message Batches API offers 50% savings but processing can take up to 24 hours. You want to optimize API cost while maintaining an acceptable developer experience. Which combination correctly matches each task to an API approach?

    Which combination is correct?

    CorrectB) Use synchronous calls for PR style checks; use the Message Batches API for weekly security audits and nightly test generation.

    PR style checks block developers and require immediate responses via synchronous calls, while weekly security audits and nightly test generation are scheduled tasks with flexible deadlines that can tolerate up to a 24-hour batch window—capturing 50% savings for both.

  3. #365Claude Code for Continuous Integration

    Your CI pipeline includes two Claude-based code review modes: a pre-merge-commit hook that blocks PR merge until completion, and a “deep analysis” that runs overnight, polls for batch completion, and posts detailed suggestions to the PR. You want to reduce API cost using the Message Batches API, which offers 50% savings but requires polling and can take up to 24 hours. Which mode should use batch processing?

    Which mode should use batch processing?

    CorrectB) Only the deep analysis.

    Deep analysis is an ideal candidate for batch processing because it already runs overnight, tolerates delay, and uses a polling model before publishing results—matching the asynchronous, polling-based architecture of the Message Batches API while capturing 50% savings.

  4. #373Claude Code for Continuous Integration

    Your team wants to reduce API costs for automated analysis. Currently, synchronous Claude calls support two workflows: (1) a blocking pre-merge check that must complete before developers can merge, and (2) a technical debt report generated overnight for review the next morning. Your manager proposes moving both to the Message Batches API to save 50%. How should you evaluate this proposal?

    How should you evaluate this proposal?

    CorrectC) Use batch processing only for technical debt reports; keep synchronous calls for pre-merge checks.

    Message Batches API processing can take up to 24 hours with no latency SLA, which is acceptable for overnight technical debt reports but unacceptable for blocking pre-merge checks where developers wait. This matches each workflow to the right API based on latency requirements.

  5. #387

    Your extraction system processes two document types: standard monthly reports (archived after processing) and urgent exception reports (must trigger business alerts within 30 minutes of receipt). Both use the same JSON schema. You want to minimize API costs while meeting latency requirements.

    How should you architect the processing pipeline?

    CorrectD) Route standard reports to the `Batch API` for 50% cost savings, and route urgent exception reports to the real-time Messages API.

    Correct. Match latency profile to document urgency: batch for the bulk (cheap), real-time for the latency-sensitive exceptions (fast). Minimizes cost while meeting SLA.

  6. #398

    Documents arrive continuously throughout business hours and need structured data extracted. To reduce costs, you want to use the `Message Batches API` (50% discount, up-to-24-hour processing window). Your SLA specifies that extraction results must be available within 30 hours of document arrival with 99.9% reliability.

    Which batching strategy is most appropriate?

    CorrectC) Submit batches every 4 hours containing documents from that window

    Correct. Max 4-hour wait + up to 24-hour batch SLO = 28-hour worst case, leaving a 2-hour cushion under the 30-hour SLA to absorb batch variance and hit 99.9%.

08

Run high-volume, noisy work in an isolated context

D5Claude Code
Correct

Large, verbose tasks consume the main context window and crowd out the primary objective. Execute them in a separate context (e.g., context: fork) and return only a concise summary.

Trap

Performing the entire noisy task in the main context, diluting it until the original goal is lost.

25 exam questions that test this
  1. #16

    Your multi-agent research pipeline crashed after processing 12 of 28 documents. The web search agent had identified relevant sources, the document analysis agent had partially completed extraction, and the synthesizer had begun pattern identification. You need to resume processing without repeating work or losing fidelity of prior findings.

    What state management approach best balances information fidelity with context efficiency when restoring agent state?

    CorrectC) Have each agent persist a structured report to a known location. On resume, the coordinator loads the reports and injects relevant state into agent prompts.

    Correct. Structured per-agent reports keep fidelity (the findings, with schema), let the coordinator stay in charge of orchestration, and keep each subagent's context focused. This is the orchestrator + compact artifact pattern.

  2. #25

    Production monitoring shows that follow-up queries like "summarize what we learned about market trends" consistently take 40+ seconds. Investigation reveals the coordinator spawns the synthesis subagent for each summarization request, passing 80K+ tokens of accumulated findings. The coordinator already has these findings in its context from orchestrating the research.

    What's the most effective way to improve response time for these follow-up summaries?

    CorrectB) Have the coordinator handle straightforward summarization requests directly using its existing context, reserving subagent spawning for complex analysis.

    Correct. If the coordinator already has the findings, spawning a subagent to re-ingest 80K tokens is pure overhead. Let the coordinator answer simple follow-ups itself.

  3. #29

    An engineer used the agent yesterday to analyze a legacy authentication module, identifying two distinct refactoring approaches: extracting a microservice versus refactoring in-place. Today, they want to explore both approaches in depth—having the agent propose specific code changes for each—before deciding which to implement.

    What's the most effective way to structure this exploration?

    CorrectD) Use `fork_session` to create two branches from yesterday's analysis, exploring one approach in each fork.

    Correct. Forking from yesterday's session gives each approach its own independent context starting from the same analysis baseline — clean, parallel, no contamination.

  4. #30

    An engineer asks your agent to identify untested code paths in a legacy payment processing module spanning 45 files. After reading the first 8 source files, the agent's responses are becoming noticeably less accurate—it's forgetting previously discussed code patterns and hasn't yet located all test files or traced critical payment flows.

    What's the most effective approach to complete this investigation?

    CorrectB) Spawn subagents to investigate specific questions (e.g., "find all test files for payment processing", "trace refund flow dependencies") while the main agent coordinates findings and preserves high-level understanding.

    Correct. Delegate well-scoped investigations to subagents with fresh context, while the main agent keeps the architectural overview. This is the pattern for scaling exploration beyond a single context window.

  5. #33

    Your agent has analyzed a complex service module—reading 23 source files, tracing request flows, and identifying error handling patterns. A developer wants to compare two testing strategies before committing to one: end-to-end tests with mocked external services vs. snapshot tests capturing expected outputs. They need to independently develop both approaches to evaluate trade-offs.

    How should you manage the sessions?

    CorrectB) Resume the analysis session with `fork_session` enabled, creating a separate branch for each testing strategy.

    Correct. Forking gives each strategy its own independent context starting from the exact analysis baseline — no cross-contamination, no re-analysis.

  6. #35

    A customer returns 4 hours after their initial session about the same billing dispute. The previous 32-turn session contains `lookup_order` results showing "Status: PENDING, Expected resolution: 24-48 hours." In testing, you observe that when resuming sessions with stale tool results, the agent often references the outdated data in responses (e.g., "I see your refund is still being processed") even after subsequent fresh tool calls return different information.

    What approach most reliably handles returning customers?

    CorrectB) Start a new session, inject a structured summary of the previous interaction (issue type, actions taken, resolution status), then make fresh tool calls before engaging.

    Correct. A clean session with a summary keeps the narrative continuity while guaranteeing the agent isn't reasoning over stale tool results.

  7. #255Code Generation with Claude Code

    Your team created a `/analyze-codebase` skill that performs deep code analysis—dependency scanning, test coverage counts, and code quality metrics. After running the command, team members report Claude becomes less responsive in the session and loses the context of the original task.

    How do you most effectively fix this while keeping full analysis capabilities?

    CorrectA) Add `context: fork` in the skill frontmatter to run the analysis in an isolated subagent context.

    `context: fork` runs the analysis in an isolated subagent context so the large output does not pollute the main session’s context window and Claude does not lose track of the original task. It preserves full analysis capability while keeping the main session responsive.

  8. #263Code Generation with Claude Code

    You create a custom skill `/explore-alternatives` that your team uses to brainstorm and evaluate implementation approaches before choosing one. Developers report that after running the skill, subsequent Claude responses are influenced by the alternatives discussion—sometimes referencing rejected approaches or retaining exploration context that interferes with actual implementation.

    How should you most effectively configure this skill?

    CorrectB) Add `context: fork` in the skill frontmatter.

    `context: fork` runs the skill in an isolated subagent context so exploration discussions do not pollute the main conversation history. This prevents rejected approaches and brainstorming context from influencing subsequent implementation work.

  9. #362Claude Code for Continuous Integration

    Your team uses Claude Code for generating code suggestions, but you notice a pattern: non-obvious issues—performance optimizations that break edge cases, cleanups that unexpectedly change behavior—are only caught when another team member reviews the PR. Claude’s reasoning during generation shows it considered these cases but concluded its approach was correct. Which approach directly addresses the root cause of this self-check limitation?

    Which approach directly addresses the root cause?

    CorrectA) Run a second independent instance of Claude Code to review the changes without access to the generator’s reasoning.

    A second independent Claude Code instance without access to the generator’s reasoning directly addresses the root cause by avoiding confirmation bias. This “fresh eyes” perspective mirrors human peer review, where another reviewer catches issues the author rationalized.

  10. #368Claude Code for Continuous Integration

    Your automated review generates test-case suggestions for each PR. Reviewing a PR that adds course completion tracking, Claude suggests 10 test cases, but developer feedback shows that 6 duplicate scenarios already covered by the existing test suite. What change most effectively reduces duplicate suggestions?

    What change is most effective?

    CorrectA) Include the existing test file in context so Claude can determine what scenarios are already covered.

    Including the existing test file fixes the root cause of duplication: Claude can only avoid suggesting already-covered scenarios if it knows what tests already exist. This gives Claude the information needed to propose genuinely new, valuable tests.

  11. #369Claude Code for Continuous Integration

    After an initial automated review identifies 12 findings, a developer pushes new commits to address issues. Re-running review produces 8 findings, but developers report that 5 duplicate previous comments on code that was already fixed in the new commits. What is the most effective way to eliminate this redundant feedback while maintaining thoroughness?

    What is the most effective way to eliminate redundant feedback?

    CorrectD) Include previous review findings in context and instruct Claude to report only new or still-unresolved issues.

    Including prior review findings in context lets Claude distinguish new problems from those already addressed in recent commits. This preserves review thoroughness while using Claude’s reasoning to avoid redundant feedback on fixed code.

  12. #494Multi-agent Research System

    Production monitoring shows inconsistent synthesis quality. When aggregated results are ~75K tokens, the synthesis agent reliably cites information from the first 15K tokens (web-search headlines/snippets) and the last 10K tokens (document analysis conclusions), but often misses critical findings in the middle 50K tokens—even when they directly answer the research question. How should you restructure the aggregated input?

    How should you restructure the aggregated input?

    CorrectC) Place a key-findings summary at the start of the aggregated input and organize detailed results with explicit section headings for easier navigation.

    Putting a key-findings summary at the start leverages primacy effects so critical information sits in the most reliably processed position. Adding explicit section headings throughout helps the model navigate and attend to mid-input content, directly mitigating the “lost in the middle” phenomenon.

  13. #495Multi-agent Research System

    In testing, the combined output of the web-search agent (85K tokens including page content) and the document analysis agent (70K tokens including chains of thought) totals 155K tokens, but the synthesis agent performs best with inputs under 50K tokens. Which solution is most effective?

    Which solution is most effective?

    CorrectA) Modify upstream agents to return structured data (key facts, quotes, relevance scores) instead of verbose content and reasoning.

    Modifying upstream agents to return structured data fixes the root cause by reducing token volume at the source while preserving essential information. It avoids passing bulky page content and reasoning traces that inflate tokens without improving the synthesis step.

  14. #496Code Generation with Claude Code

    You’re adding error-handling wrappers around external API calls across a 120-file codebase. The work has three phases: (1) discover all call sites and patterns, (2) collaboratively design the error-handling approach, and (3) implement wrappers consistently. In Phase 1, Claude generates large output listing hundreds of call sites with context, quickly filling the context window before discovery finishes.

    Which approach is most effective to complete the task while maintaining implementation consistency?

    CorrectA) Use an Explore subagent for Phase 1 to isolate verbose discovery output and return a summary, then continue Phases 2–3 in the main conversation.

    An Explore subagent isolates the verbose discovery output in a separate context and returns only a concise summary to the main conversation. This preserves the main context window for the collaborative design and consistent implementation phases where retained context is most valuable.

  15. #497Customer Support Agent

    Production logs show a pattern: customers reference specific amounts (e.g., “the 15% discount I mentioned”), but the agent responds with incorrect values. Investigation shows these details were mentioned 20+ turns ago and condensed into vague summaries like “promotional pricing was discussed.” What fix is most effective?

    What fix is most effective?

    CorrectC) Extract transactional facts (amounts, dates, order numbers) into a persistent “case facts” block included in every prompt outside the summarized history.

    Summarization inherently loses precise details. Extracting transactional facts into a structured “case facts” block outside the summarized history preserves critical information so it’s reliably available in every prompt regardless of how many turns have been summarized.

  16. #500Conversational AI Architecture Patterns

    After a 40-minute cooking session, the conversation reaches 78,000 tokens. History includes allergies, recipe scaling, clarified cooking terms, and general discussion. You must reduce tokens while preserving important information.

    What approach best balances preservation with token reduction?

    CorrectC) Extract critical structured data (allergies, quantities, preferences), summarize general discussion, and keep recent exchanges verbatim.

    The hybrid approach preserves the highest-value information at the lowest cost. Critical facts like allergies and recipe quantities are extracted into a compact structured block (preventing the precision loss that occurs during summarization), general discussion is summarized, and recent exchanges are kept verbatim for conversational coherence. Options A and B risk losing critical dietary information; D is architectural overkill for a single cooking session.

  17. #501Conversational AI Architecture Patterns

    Users report that during extended conversations the assistant loses track of earlier topics and preferences. Your current implementation keeps only the last 25 message pairs.

    What is the most effective solution?

    CorrectA) Hybrid approach: summarize older messages while keeping recent ones verbatim.

    The hybrid approach addresses both dimensions of the problem: retaining exact recent context (critical for conversational coherence) while maintaining a compressed representation of earlier preferences (preventing total loss when pairs are dropped). Increasing the window (C) simply delays the same problem. Vector search (B) may miss important context that isn't semantically similar to the current query. Full per-turn summarization (D) adds overhead and accumulates summarization errors.

  18. #502Conversational AI Architecture Patterns

    Users report that latency increases and costs rise when conversations exceed 50 turns.

    What is the primary cause?

    CorrectA) The entire conversation history is included with each API request.

    Claude's API is fully stateless—every request must include the complete conversation history in the `messages` array. As conversations grow, each request carries more tokens, which directly increases both processing latency and cost. The model does not maintain any internal state between calls (D is false), and response length is not inherently tied to conversation length (B).

  19. #503Conversational AI Architecture Patterns

    After three months of weekly sessions, conversation history grows to 85,000 tokens. When a user asks "What did we conclude about the theme of isolation?", the assistant gives generic answers instead of referencing previous discussions.

    What is the most effective approach?

    CorrectC) Semantic embeddings with retrieval of relevant exchanges.

    Semantic search over conversation history is the only approach that scales to three months of discussion while being able to surface specific relevant exchanges on demand. Rolling window (A) would discard most of the history. Progressive summarization (B) compresses discussions into abstractions that lose the specific conclusions users are asking about. XML tags (D) require restructuring all past content and don't solve the retrieval problem at this scale.

  20. #505Conversational AI Architecture Patterns

    Your assistant uses a contractor-persona system prompt. Early turns follow the rules, but by turn 7 the assistant gives generic advice. Conversation length is only 2,500 tokens.

    What is the most likely cause?

    CorrectC) Accumulated assistant responses dilute system prompt influence.

    As assistant responses accumulate in the conversation history, the proportion of text reflecting the system prompt's behavioral constraints decreases relative to the growing body of assistant-generated content. The model increasingly pattern-matches to its own prior outputs rather than the system prompt, compounding drift even at short token lengths. The system prompt is included in every API call (D is false as a standalone explanation), and model attention degradation (B) doesn't operate at 2,500 tokens.

  21. #506

    During testing, you observe that in extended exploration sessions (30+ minutes), the agent starts giving inconsistent answers about code structure it discussed earlier. Engineers report having to repeat context about modules they've already explored.

    What's the most effective approach to address this?

    CorrectA) Have the agent maintain a scratchpad file that records key findings, referencing it for subsequent questions.

    Correct. A scratchpad offloads findings to durable storage the agent can re-read on demand, giving it a stable 'memory' independent of how crowded the context window gets.

  22. #507

    Your agent has spent 25 minutes exploring a game engine's rendering subsystem—reading shader code, buffer management, and frame synchronization logic. An engineer now asks it to understand how the physics engine integrates with rendering for collision debug overlays. You notice recent responses reference "typical rendering patterns" rather than the specific VulkanPipeline and FrameGraph classes it discovered earlier.

    What's the most effective approach?

    CorrectC) Summarize key rendering findings, then spawn a sub-agent for physics exploration with that summary in its initial context.

    Correct. Condense what you've learned about rendering into a compact summary, then give a fresh subagent that summary plus the physics task — you preserve the important signal and escape the degraded context.

  23. #508

    An engineer asks the agent to understand how the caching layer works before adding a new cache invalidation trigger. After initial Grep searches, the agent has identified that caching logic spans 15 files including decorators, middleware, and service classes (~8,000 lines total).

    What's the most effective next step for building understanding while managing context constraints?

    CorrectB) Analyze imports and class hierarchies to identify the base cache class, Read that file to understand the interface, then trace specific invalidation implementations.

    Correct. Start from the architectural root (the interface), then navigate only the specific implementations that matter for invalidation — focused reading, low context cost.

  24. #510

    A customer raises three separate issues during one session: a refund inquiry (turns 1-15), a subscription question (turns 16-30), and a payment method update (turns 31-45). At turn 48, the customer asks "What happened with my refund?" The conversation is approaching context limits.

    What strategy best maintains the agent's ability to address all issues throughout the session?

    CorrectC) Summarize earlier turns into a narrative description, preserving full message history only for the active issue.

    Correct. Progressive summarization compresses stable resolved topics while keeping the active thread verbatim — the classic pattern for long multi-issue conversations near the context limit.

  25. #511

    Your agent has called `lookup_order` multiple times while investigating a customer's return requests. Each response includes 40+ fields (items, shipping details, payment info, status history). Tool outputs now represent the majority of the conversation's context. The customer mentions two more orders they want to discuss.

    What's the most effective approach before making additional lookups?

    CorrectA) Extract only return-relevant fields (items, purchase date, return window, status) from each existing order response, removing verbose details

    Correct. Keep the fields that matter for the task and drop the rest. This directly addresses the context-bloat problem before you add two more lookups.

09

Put each instruction in the configuration file built for it

D3Claude Code
Correct

Route guidance to the right place: CLAUDE.md for always-on rules, Skills for on-demand capabilities, .claude/rules/ for file-type rules, .claude/commands/ for shared team commands. Reference secrets securely — never hard-code them.

Trap

Dumping everything into a single file, or writing a credential directly into the code.

13 exam questions that test this
  1. #251Claude Code for Continuous Integration

    Your pipeline script runs `claude "Analyze this pull request for security issues"`, but the job hangs indefinitely. Logs show Claude Code is waiting for interactive input. What is the correct approach to run Claude Code in an automated pipeline?

    What is the correct approach?

    CorrectB) Add the `-p` flag: `claude -p "Analyze this pull request for security issues"`.

    The `-p` (or `--print`) flag is the documented way to run Claude Code non-interactively. It processes the prompt, prints the result to stdout, and exits without waiting for user input—ideal for CI/CD pipelines.

  2. #253Code Generation with Claude Code

    Your CLAUDE.md file has grown to 400+ lines containing coding standards, testing conventions, a detailed PR review checklist, deployment instructions, and database migration procedures. You want Claude to always follow coding standards and testing conventions, but apply PR review, deploy, and migration guidance only when doing those tasks.

    Which restructuring approach is most effective?

    CorrectD) Keep universal standards in CLAUDE.md and create Skills for workflow-specific guidance (PR review, deploy, migrations) with trigger keywords.

    CLAUDE.md content loads in every session, ensuring coding standards and testing conventions always apply, while Skills are invoked on demand when Claude detects trigger keywords—ideal for workflow-specific guidance like PR review, deployment, and migrations.

  3. #256Code Generation with Claude Code

    Your team uses a `/commit` skill in `.claude/skills/commit/SKILL.md`. A developer wants to customize it for their personal workflow (different commit message format, extra checks) without affecting teammates.

    What do you recommend?

    CorrectC) Create a personal version at `~/.claude/skills/commit/SKILL.md` with the same name.

    Personal skills take precedence over project skills with the same name. A personal skill at `~/.claude/skills/commit/SKILL.md` will override the team’s project skill, allowing the developer to customize their workflow while maintaining the familiar `/commit` command name for their personal use. This approach is better than option A because it preserves the original command name, improving the developer’s workflow without affecting teammates.

  4. #257Code Generation with Claude Code

    Your team has used Claude Code for months. Recently, three developers report Claude follows the guidance “always include comprehensive error handling,” but a fourth developer who just joined says Claude does not follow it. All four work in the same repo and have up-to-date code.

    What is the most likely cause and fix?

    CorrectA) The guidance lives in the original developers’ user-level `~/.claude/CLAUDE.md` files, not in the project `.claude/CLAUDE.md`. Move the instruction to the project-level file so all team members receive it.

    If the guidance was added only to the original developers’ user-level configs and not to the project-level `.claude/CLAUDE.md`, new team members won’t receive it. Moving it to the project-level configuration ensures all current and future team members automatically get the guidance.

  5. #258Code Generation with Claude Code

    You find that including 2–3 full endpoint implementation examples as context significantly improves consistency when generating new API endpoints. However, this context is useful only when creating new endpoints—not when debugging, reviewing code, or other work in the API directory.

    Which configuration approach is most effective?

    CorrectD) Create a skill that references the endpoint examples and contains pattern-following instructions, invoked on demand via a slash command.

    A skill invoked on demand loads the example context only when generating new endpoints, not during unrelated tasks like debugging or review. This keeps the main context clean while preserving high-quality generation when needed.

  6. #259Code Generation with Claude Code

    Your team created a `/migration` skill that generates database migration files. It takes the migration name via `$ARGUMENTS`. In production you observe three issues: (1) developers often run the skill without arguments, causing poorly named files, (2) the skill sometimes uses database schema details from unrelated prior conversations, and (3) a developer accidentally ran destructive test cleanup when the skill had broad tool access.

    Which configuration approach fixes all three problems?

    CorrectB) Add `argument-hint` in frontmatter to request required parameters, use `context: fork` to isolate execution, and restrict `allowed-tools` to file-write operations.

    This uses three separate configuration features to address each problem: `argument-hint` improves argument entry and reduces missing arguments, `context: fork` prevents context leakage from prior conversations, and `allowed-tools` constrains the skill to safe file-writing operations, preventing destructive actions.

  7. #260Code Generation with Claude Code

    Your codebase contains areas with different coding conventions: React components use functional style with hooks, API handlers use async/await with specific error handling, and database models follow the repository pattern. Test files are distributed across the codebase next to the code under test (e.g., `Button.test.tsx` next to `Button.tsx`), and you want all tests to follow the same conventions regardless of location.

    What is the most supported way to ensure Claude automatically applies the correct conventions when generating code?

    CorrectD) Create rule files under `.claude/rules/` with YAML frontmatter specifying glob patterns to conditionally apply conventions based on file paths.

    `.claude/rules/` files with YAML frontmatter and glob patterns (e.g., `**/*.test.tsx`, `src/api/**/*.ts`) enable deterministic, path-based convention application regardless of directory structure. This is the most supported approach for cross-cutting patterns like distributed test files.

  8. #261Code Generation with Claude Code

    You want to create a custom slash command `/review` that runs your team’s standard code review checklist. It should be available to every developer when they clone or update the repository.

    Where should you create the command file?

    CorrectB) In the project repository under `.claude/commands/`.

    Putting custom slash commands under `.claude/commands/` inside the project repository ensures they are version-controlled and automatically available to every developer who clones or updates the repo. This is the intended location for project-level custom commands in Claude Code.

  9. #262Code Generation with Claude Code

    Your team’s CLAUDE.md grew beyond 500 lines mixing TypeScript conventions, testing guidance, API patterns, and deployment procedures. Developers find it hard to locate and update the right sections.

    What approach does Claude Code support to organize project-level instructions into focused topical modules?

    CorrectB) Create separate Markdown files in `.claude/rules/`, each covering one topic (e.g., `testing.md`, `api-conventions.md`).

    Claude Code supports a `.claude/rules/` directory where you can create separate Markdown files for topical guidance (e.g., `testing.md`, `api-conventions.md`), allowing teams to organize large instruction sets into focused, maintainable modules.

  10. #264Code Generation with Claude Code

    Your team wants to add a GitHub MCP server for searching PRs and checking CI status via Claude Code. Each of six developers has their own personal GitHub access token. You want consistent tooling across the team without committing credentials to version control.

    Which configuration approach is most effective?

    CorrectC) Add the server to the project `.mcp.json` using environment variable substitution (`${GITHUB_TOKEN}`) for auth and document the required environment variable in the project README.

    A project `.mcp.json` with environment variable substitution is idiomatic: it provides a single version-controlled source of truth for MCP configuration while letting each developer supply credentials via environment variables. Documenting the variable makes onboarding easy without committing secrets.

  11. #265

    An engineer used `Claude Code` yesterday to investigate authentication flows in a legacy monolith, building up significant context over a 2-hour session. Today she wants to continue that specific investigation. She's worked on three other codebases since then and knows the session was named "auth-deep-dive".

    How should she resume?

    CorrectD) Use `--resume` auth-deep-dive to load that specific session by name

    Correct. `--resume` with the session name is designed for exactly this: pick a specific prior session out of many, by the name you gave it.

  12. #380Conversational AI Architecture Patterns

    Your assistant must maintain an enthusiastic tone, explain its reasoning, and ask clarifying questions. Where should these behavioral guidelines be defined?

    Where should these behavioral guidelines be defined?

    CorrectB) In the system prompt.

    The system prompt is specifically designed for persistent behavioral constraints and guidelines that apply throughout the entire conversation. Prepending to each user message (A) is redundant overhead. The first assistant message (C) is unreliable because the model can deviate from its own prior statements. Environment variables (D) have no effect on model behavior.

  13. #382Conversational AI Architecture Patterns

    A webhook notifies your system that a user's package has shipped while the user is actively chatting. You want the assistant to incorporate this naturally into the next response.

    What is the best approach?

    CorrectD) Append the status update as a prefix to the next user message.

    Prefixing the status update to the next user message injects real-time context at a natural conversation boundary without disrupting the flow. Modifying the system prompt (A) requires rebuilding the session or is architecturally cumbersome. A synthetic user message (B) can break the natural dialogue flow and confuse attribution. Forcing a tool call each turn (C) is wasteful when events are rare.

10

Enforce mandatory steps in code, not in prose

D4D1Support
Correct

When a step is non-negotiable — e.g., verifying a customer's identity before issuing a refund — enforce it programmatically so it cannot be skipped, rather than relying on the model to remember.

Trap

Adding 'always verify first' to the prompt and trusting the instruction to hold; a request is not a guarantee.

8 exam questions that test this
  1. #21

    The synthesis agent receives summarized findings from the web search and document analysis agents, then passes a consolidated summary to the report generator. During testing, you discover the generated reports make factual claims without proper citations—the report generator cannot attribute statements to their original sources because that metadata was lost during the summarization steps.

    What's the most effective approach to ensure proper source attribution in the final reports?

    CorrectA) Have each agent output structured data separating content summaries from source metadata (URLs, document names, page numbers).

    Correct. Structured content + separate source metadata preserves the mapping end-to-end, so the report generator receives both what was said and where it came from.

  2. #22

    In production, final reports frequently contain claims without proper source attribution. Investigation shows that while the web search and document analysis agents correctly attach citations to their outputs, the synthesis agent loses track of which sources support which conclusions when combining findings.

    What's the most effective architectural change?

    CorrectB) Require all subagents to output structured claim-source mappings that the synthesis agent must preserve and merge when combining findings from multiple sources.

    Correct. Explicit claim-source mappings are a first-class output the synthesis agent can merge deterministically — no attribution gets dropped during summarization.

  3. #38

    Compliance requires that refunds exceeding $500 must automatically escalate to a human agent—this rule cannot be left to model discretion. Despite clear system prompt instructions, production logs show the agent occasionally processes high-value refunds directly (3% failure rate).

    How should you achieve guaranteed compliance?

    CorrectC) Implement a hook to intercept tool calls; when the refund process amount exceeds $500, block it and invoke human escalation.

    Correct. Compliance-grade rules belong outside the model — a deterministic hook on the tool call is guaranteed to fire every time, independent of model behavior.

  4. #158Customer Support Agent

    Production logs show that in 12% of cases your agent skips `get_customer` and calls `lookup_order` directly using only the customer-provided name, sometimes leading to misidentified accounts and incorrect refunds. What change most effectively fixes this reliability problem?

    What change is most effective?

    CorrectC) Add a programmatic precondition that blocks `lookup_order` and `process_refund` until `get_customer` returns a verified customer identifier.

    A programmatic precondition provides a deterministic guarantee that required sequencing is followed. It’s the most effective approach because it eliminates the possibility of skipping verification, regardless of LLM behavior.

  5. #162Conversational AI Architecture Patterns

    Your `remove_team_member` tool uses a `dry_run: boolean` parameter for previewing impacts before execution. Production monitoring shows the agent bypasses the preview step by calling with `dry_run=false` directly. You need to ensure every removal is preceded by a preview that the user explicitly confirms.

    What is the most reliable approach?

    CorrectD) Replace with two tools: `preview_remove_member` returns impact details and a single-use confirmation token; `execute_remove_member` requires that token, binding execution to the preview.

    The two-tool token-binding approach makes it architecturally impossible to execute without a prior preview—the execute tool literally requires a token that only the preview tool can generate. This is the only approach that enforces the constraint at the code level rather than relying on LLM compliance with instructions (C), timing heuristics (A), or orchestration infrastructure (B).

  6. #250Claude Code for Continuous Integration

    Your CI pipeline runs the Claude Code CLI (in `--print` mode) using CLAUDE.md to provide project context for code review, and developers generally find the reviews substantive. However, they report that integrating findings into the workflow is difficult—Claude outputs narrative paragraphs that must be manually copied into PR comments. The team wants to automatically post each finding as a separate inline PR comment at the relevant place in code, which requires structured data with file path, line number, severity level, and suggested fix. Which approach is most effective?

    Which approach is most effective?

    CorrectB) Use the CLI flags `--output-format json` and `--json-schema` to enforce structured findings, then parse the output to post inline comments via the GitHub API.

    Using `--output-format json` with `--json-schema` enforces structured output at the CLI level, guaranteeing well-formed JSON with the required fields (file path, line number, severity, suggested fix) that can be reliably parsed and posted as inline PR comments via the GitHub API. It leverages built-in CLI capabilities designed specifically for structured output.

  7. #395

    Your extraction pipeline processes invoices and extracts line items, subtotals, tax amounts, and grand totals. During evaluation, you discover that in 18% of extractions, the sum of extracted line item amounts doesn't match the extracted grand total—sometimes due to OCR errors in the source document, sometimes due to extraction mistakes by the model. Downstream accounting systems reject records with mismatched totals.

    What's the most effective approach to improve extraction reliability?

    CorrectA) Add a "`calculated_total`" field where the model sums extracted line items alongside a "`stated_total`" field. Flag records for human review when values differ.

    Correct. Capturing both values makes the discrepancy a first-class signal — you catch OCR errors and extraction mistakes uniformly, and you can route only the mismatched 18% to humans.

  8. #399

    After deployment, you find that 12% of extractions contain semantic errors that pass JSON schema validation (e.g., a duration like "30 minutes" incorrectly placed in an ingredient quantity field). Human reviewers have capacity to check only 20% of extractions.

    Which approach most effectively allocates reviewer attention?

    CorrectA) Have the model output field-level confidence scores, then calibrate review thresholds using a labeled validation set.

    Correct. Field-level confidence lets you route the low-confidence 20% — which is where the 12% semantic errors concentrate — to humans. Calibration makes the threshold choice data-driven.

11

Plan before building when several viable approaches exist

D3D1Claude Code
Correct

For large or ambiguous tasks with multiple possible designs, investigate, weigh the options, and agree on an approach before implementation begins.

Trap

Starting to code immediately, before it is clear which approach is best.

6 exam questions that test this
  1. #24

    The coordinator provides detailed step-by-step instructions to the web search subagent, specifying exact search queries, source priorities, and date filters. Production monitoring reveals three issues: (1) the subagent reports "insufficient results" rather than trying alternative approaches when pre-specified searches fail, (2) research quality drops for emerging topics that don't match expected patterns, and (3) the subagent rarely surfaces valuable tangential sources.

    What's the most effective way to improve subagent adaptability?

    CorrectD) Specify research goals and quality criteria (coverage breadth, source diversity, recency) rather than procedural steps, letting the subagent determine its search strategy.

    Correct. Delegate intent and quality bars, not procedures. The subagent can then choose queries, follow promising tangents, and recover from dead ends on its own.

  2. #166

    An engineer asks the agent to find all callers of a function before removing it. The function is defined in a core library but is also exposed through wrapper modules that rename the function for domain-specific use (e.g., calculateTax in the library becomes computeOrderTax in the orders module).

    What exploration strategy will most reliably identify all callers?

    CorrectA) Read the library and wrapper modules to identify all exposed names for the function, then Grep for each name across the codebase.

    Correct. You have to enumerate every name the function is exposed under — otherwise renamed wrappers hide callers. Read the relevant modules, gather all aliases, then grep for each.

  3. #252Code Generation with Claude Code

    You need to add Slack as a new notification channel. The existing codebase has clear, established patterns for email, SMS, and push channels. However, Slack’s API offers fundamentally different integration approaches—incoming webhooks (simple, one-way), bot tokens (support delivery confirmation and programmatic control), or Slack Apps (two-way events, requires workspace approval). Your task says “add Slack support” without specifying integration method or requiring advanced features like delivery tracking.

    How should you approach this task?

    CorrectB) Switch to planning mode to explore integration options and architectural implications, then present a recommendation before implementation.

    Slack integration has multiple valid approaches with significantly different architectural implications, and requirements are ambiguous. Planning mode lets you evaluate trade-offs among webhooks, bot tokens, and Slack Apps and align on an approach before implementation.

  4. #254Code Generation with Claude Code

    You’re tasked with restructuring your team’s monolithic application into microservices. This impacts changes across dozens of files and requires decisions about service boundaries and module dependencies.

    Which approach should you choose?

    CorrectA) Switch to planning mode to explore the codebase, understand dependencies, and design the implementation approach before making changes.

    Planning mode is the right strategy for complex architectural restructuring like splitting a monolith: it allows safe exploration and informed decisions about boundaries before committing to potentially expensive changes across many files.

  5. #383Conversational AI Architecture Patterns

    Users frequently send requests like "Book a venue for the party." The assistant asks 4+ clarifying questions, causing 35% abandonment.

    What approach best improves the trade-off?

    CorrectC) State assumptions explicitly and proceed while inviting corrections.

    Stating assumptions explicitly and proceeding gives the user an immediate, useful response while preserving their ability to correct wrong assumptions. Hidden defaults (A) leave the user unaware of what was assumed. A compound question list (B) still demands upfront effort from the user. A structured form (D) adds more friction, not less—contradicting the goal of reducing abandonment.

  6. #384Conversational AI Architecture Patterns

    Users ask vague requests like "Can you help with the report?" The assistant responds by asking multiple questions (which report? what help? deadline?), causing 40% abandonment.

    What is the best solution?

    CorrectA) Make reasonable assumptions, state them explicitly, and offer to adjust.

    Proceeding with reasonable stated assumptions eliminates the back-and-forth entirely while keeping the user informed and in control. Predefined silent interpretations (C) leave users confused when the response doesn't match their intent. A single-question limit (D) still requires turns of back-and-forth. A smaller classification model (B) adds latency and infrastructure complexity without solving the core UX problem.

12

Decompose large requests, parallelize, and share context

D1Multi-agentSupport
Correct

When a request contains several independent parts, split them and process them concurrently against shared context, rather than handling them one slow step at a time and re-deriving the same information.

Trap

Processing sequentially and repeatedly looking up the same information.

14 exam questions that test this
  1. #2Multi-agent Research System

    The web-search and document-analysis agents have completed their tasks and returned results to the coordinator. What is the next step for creating an integrated research report?

    Which next step is most appropriate?

    CorrectC) The coordinator passes both sets of results to the synthesis agent for a unified integration.

    In a coordinator–subagent architecture, the coordinator forwards both result sets to the synthesis agent for centralized integration, preserving control and ensuring high-quality merging.

  2. #4Multi-agent Research System

    After running the system on “AI impact on creative industries,” you observe that every subagent completes successfully: the web-search agent finds relevant articles, the document analysis agent summarizes them correctly, and the synthesis agent produces coherent text. However, final reports cover only visual art and completely miss music, literature, and film. In the coordinator logs, you see it decomposed the topic into three subtasks: “AI in digital art,” “AI in graphic design,” and “AI in photography.” What is the most likely root cause?

    What is the most likely root cause?

    CorrectC) The coordinator’s task decomposition is too narrow, assigning subagents work that does not cover all relevant areas.

    The coordinator decomposed a broad topic only into visual-art subtasks, missing music, literature, and film entirely. Since subagents executed their assignments correctly, the narrow decomposition is the obvious root cause.

  3. #7Multi-agent Research System

    A colleague proposes that the document analysis agent should send its results directly to the synthesis agent, bypassing the coordinator. What is the main advantage of keeping the coordinator as the central hub for all communication between subagents?

    What is the main advantage of keeping the coordinator as the central hub?

    CorrectA) The coordinator can observe all interactions, handle errors uniformly, and decide what information each subagent should receive.

    The coordinator pattern provides centralized visibility into all interactions, uniform error handling across the system, and fine-grained control over what information each subagent receives—these are the primary advantages of a star-shaped communication topology.

  4. #9Multi-agent Research System

    While researching a broad topic, you observe that the web-search agent and the document analysis agent investigate the same subtopics, leading to substantial duplication in their outputs. Token usage nearly doubles without a proportional increase in research breadth or depth. What is the most effective way to address this?

    What is the most effective way to address this?

    CorrectB) The coordinator explicitly partitions the research space before delegating, assigning each agent distinct subtopics or source types.

    Having the coordinator explicitly partition the research space before delegating is most effective because it addresses the root cause—unclear task boundaries—before any work begins. It preserves parallelism while preventing duplicated effort and wasted tokens.

  5. #12Customer Support Agent

    Production logs show that for simple requests like “refund for order #1234,” your agent resolves the issue in 3–4 tool calls with 91% success. But for complex requests like “I was billed twice, my discount didn’t apply, and I want to cancel,” the agent averages 12+ tool calls with only 54% success—often investigating issues sequentially and fetching redundant customer data for each. What change most effectively improves handling of complex requests?

    What change is most effective?

    CorrectC) Decompose the request into separate issues, then investigate each in parallel using shared customer context before synthesizing a final resolution.

    Decomposing into separate issues and investigating in parallel with shared customer context fixes both key problems: it eliminates redundant data retrieval by reusing shared context across issues and reduces total tool-call loops by parallelizing investigation before synthesizing a single resolution.

  6. #14Customer Support Agent

    Production metrics show your agent averages 4+ API loops per resolution. Analysis reveals Claude often requests `get_customer` and `lookup_order` in separate sequential turns even when both are needed initially. What is the most effective way to reduce the number of loops?

    What is the most effective way to reduce loops?

    CorrectD) Instruct Claude in the prompt to bundle tool requests into one turn and return all results together before the next API call.

    Prompting Claude to bundle related tool requests into a single turn leverages its native ability to request multiple tools at once. It directly fixes the sequential-call pattern with minimal architectural change.

  7. #17

    After the web search agent finds 25 sources (120K tokens of raw content), the document analysis agent extracts key insights (15K tokens), and the synthesis agent produces a coherent narrative draft (3K tokens), the coordinator must pass context to the report generation agent for the final output with proper source citations.

    What context-passing strategy provides the best balance of completeness and efficiency?

    CorrectB) Pass the synthesis draft along with a structured source index that maps key claims to their source URLs and relevant excerpts.

    Correct. The synthesis gives the narrative; the source index gives the report generator exactly the binding it needs to cite without re-reading 120K tokens of raw content.

  8. #18

    The web search agent has gathered several relevant sources for a research topic. The document analysis agent now needs to examine these sources.

    How does information typically flow between these two specialized subagents?

    CorrectC) The coordinator agent receives the web search agent's output and includes relevant findings in the prompt when invoking the document analysis agent.

    Correct. In an orchestrator-worker pattern the coordinator is the hub. It collects each subagent's output and explicitly forwards the relevant parts into the next subagent's prompt.

  9. #19

    In production, you observe that simple fact-checking queries (e.g., "What year was the Paris Climate Agreement signed?") traverse all four subagents sequentially, consuming 40+ seconds and significant tokens per query. Complex comparative research benefits from the full pipeline. Your query distribution is diverse and evolving as users discover new applications.

    What's the most effective approach to optimize for varying query complexity?

    CorrectC) Have the coordinator analyze each query and dynamically decide which subagents to invoke based on its assessment of query requirements.

    Correct. Letting the coordinator LLM reason about each query and pick only the subagents it needs adapts naturally to an evolving, diverse query distribution — this is the strength of the orchestrator pattern.

  10. #23

    After the web search agent and document analysis agent complete their tasks, the coordinator invokes the synthesis agent. However, the synthesis agent responds that it cannot complete the task because no research findings were provided.

    What is the most likely cause of this issue?

    CorrectB) The coordinator did not include the outputs from the previous agents in the synthesis agent's prompt.

    Correct. Subagent invocations are isolated — nothing flows between them unless the coordinator explicitly puts it in the prompt. The message 'no findings provided' is exactly what you'd see.

  11. #26

    When analyzing complex legal cases that cite multiple precedents, the document analysis subagent processes each sequentially. A landmark case citing 12 precedents takes over 3 minutes to analyze completely.

    What's the most effective way to reduce this latency while preserving the coordinator's ability to monitor and debug the system?

    CorrectC) Have the coordinator spawn parallel document analysis subagents, each handling a subset of precedents, then aggregate results before synthesis.

    Correct. Coordinator-managed parallelism fans out the work, keeps each subagent's scope tight, and preserves a single hub for monitoring and aggregation.

  12. #31

    A developer asks the agent to investigate why a specific API endpoint intermittently returns 500 errors. The codebase has 200+ files and the developer doesn't know which components are involved. The agent must trace the error through routing, middleware, business logic, and database layers.

    What task decomposition approach would be most effective?

    CorrectB) Have the agent dynamically generate investigation subtasks based on what it discovers at each step, adapting its exploration plan as new information about the error path emerges.

    Correct. Debugging is adaptive by nature — each file you read changes the most useful next step. Let the agent follow the evidence.

  13. #34

    An engineer who just joined the team asks the agent to help them understand the authentication and authorization architecture before making security improvements. The codebase has 800+ files across multiple services.

    What exploration strategy will most effectively build understanding, given Claude built-in tools and context limits?

    CorrectC) Use Grep to find authentication entry points, read those files, then follow imports and function calls to map the auth flow incrementally.

    Correct. Start at entry points (login, token verify, middleware), then trace outward following real code edges. Incremental, grounded, fits within context limits.

  14. #370Claude Code for Continuous Integration

    A pull request changes 14 files in an inventory tracking module. A single-pass review that analyzes all files together produces inconsistent results: detailed feedback on some files but shallow comments on others, missed obvious bugs, and contradictory feedback (a pattern is flagged in one file but identical code is approved in another file in the same PR). How should you restructure the review?

    How should you restructure the review?

    CorrectB) Split into focused passes: review each file individually for local issues, then run a separate integration-oriented pass to examine cross-file data flows.

    Focused per-file passes address the root cause—attention dilution—by ensuring consistent depth and reliable local issue detection. A separate integration-oriented pass then covers cross-file concerns such as dependency and data-flow interactions.

How to read any question

When two answers both look right, ask these five questions. The correct answer usually passes all five.

  1. Does it address the root cause, or only the symptom? The correct answer fixes why the failure occurred; traps merely tidy up the visible effect.
  2. Does it preserve all the information? Strong answers retain the facts and their provenance; weak ones discard or obscure information.
  3. Does it resolve the issue at the smallest, simplest level? Recover locally and scope tools to least privilege — neither escalate everything nor rely on luck.
  4. Is it the right mechanism for the job? Enforce in code vs. request in prose · interactive vs. batch · the correct file for each instruction.
  5. Is it the smallest change that works? The simplest effective fix beats rebuilding everything — and beats doing nothing.

What's on the exam

How the 136 practice questions break down by domain, and why you shouldn't guess by the answer letter.

Domains · weight = official blueprint; count = items in this bank

D1 · Agent Architecture and Orchestration27% · 43 q
D2 · Tool Design and MCP Integration18% · 20 q
D3 · Claude Code Configuration and Workflows20% · 16 q
D4 · Prompt Engineering and Structured Output20% · 38 q
D5 · Context Management and Reliability15% · 19 q

Correct-answer letter (of 136)

A
39
B
32
C
36
D
29

Fairly even — so don't guess by letter. A longer option that 'does the work AND preserves the information AND escalates appropriately' is often correct, but confirm it against the five questions above rather than choosing it for its length.