[{"id": "g1", "domain": 1, "scenario": "Multi-agent Research System", "situation": "A document analysis agent discovers that two credible sources contain directly contradictory statistics for a key metric: a government report states 40% growth, while an industry analysis states 12%. Both sources look credible, and the discrepancy could materially affect the research conclusions. How should the document analysis agent handle this situation most effectively?", "question": "Which approach is most effective?", "options": [{"letter": "A", "text": "Apply credibility heuristics to pick the most likely correct number, finish analysis with that value, and add a footnote mentioning the discrepancy.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Include both numbers in the analysis output without marking them as conflicting, letting the synthesis agent decide which to use based on broader context.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Stop analysis and immediately escalate to the coordinator, asking it to decide which source is more authoritative before continuing.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Complete analysis with both numbers, explicitly annotate the conflict with source attribution, and let the coordinator decide how to reconcile the data before passing to synthesis.", "correct": true, "explanation": "This approach preserves separation of responsibilities: the analysis agent completes its core work without blocking, preserves both conflicting values with clear attribution, and correctly passes reconciliation to the coordinator, which has broader context."}], "correct": "D", "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "group": "B"}, {"id": "g2", "domain": 1, "scenario": "Multi-agent Research System", "situation": "The web-search and document-analysis agents have completed their tasks and returned results to the coordinator. What is the next step for creating an integrated research report?", "question": "Which next step is most appropriate?", "options": [{"letter": "A", "text": "Each agent sends its results directly to the report-writing agent, bypassing the coordinator.", "correct": false, "explanation": ""}, {"letter": "B", "text": "The document analysis agent requests web-search results and merges them internally.", "correct": false, "explanation": ""}, {"letter": "C", "text": "The coordinator passes both sets of results to the synthesis agent for a unified integration.", "correct": true, "explanation": "In a coordinator–subagent architecture, the coordinator forwards both result sets to the synthesis agent for centralized integration, preserving control and ensuring high-quality merging."}, {"letter": "D", "text": "The coordinator concatenates the raw outputs from both agents and returns them as the final result.", "correct": false, "explanation": ""}], "correct": "C", "task_id": "1.2", "objective": "Orchestrate multi-agent systems with coordinator-subagent patterns", "group": "A"}, {"id": "g3", "domain": 1, "scenario": "Multi-agent Research System", "situation": "A document analysis subagent frequently fails when processing PDF files: some have corrupted sections that trigger parsing exceptions, others are password-protected, and sometimes the parsing library hangs on large files. Currently, any exception immediately terminates the subagent and returns an error to the coordinator, which must decide whether to retry, skip, or fail the whole task. This causes excessive coordinator involvement in routine error handling. What architectural improvement is most effective?", "question": "Which improvement is most effective?", "options": [{"letter": "A", "text": "Create a dedicated error-handling agent that monitors all failures via a shared queue and decides recovery actions, sending restart commands directly to subagents.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Configure the subagent to always return partial results with a success status, embedding error details in metadata; the coordinator treats all responses as successful.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Make the coordinator validate all documents before sending them to the subagent, rejecting documents that might cause failures.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Implement local recovery in the subagent for transient failures and escalate to the coordinator only errors it cannot resolve, including attempted steps and partial results.", "correct": true, "explanation": "Handle errors at the lowest level capable of resolving them. Local recovery reduces coordinator workload while still escalating truly unrecoverable issues with full context and partial progress."}], "correct": "D", "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "group": "J"}, {"id": "g4", "domain": 1, "scenario": "Multi-agent Research System", "situation": "After running the system on “AI impact on creative industries,” you observe that every subagent completes successfully: the web-search agent finds relevant articles, the document analysis agent summarizes them correctly, and the synthesis agent produces coherent text. However, final reports cover only visual art and completely miss music, literature, and film. In the coordinator logs, you see it decomposed the topic into three subtasks: “AI in digital art,” “AI in graphic design,” and “AI in photography.” What is the most likely root cause?", "question": "What is the most likely root cause?", "options": [{"letter": "A", "text": "The synthesis agent lacks instructions to detect coverage gaps.", "correct": false, "explanation": ""}, {"letter": "B", "text": "The document analysis agent filters out non-visual sources due to overly strict relevance criteria.", "correct": false, "explanation": ""}, {"letter": "C", "text": "The coordinator’s task decomposition is too narrow, assigning subagents work that does not cover all relevant areas.", "correct": true, "explanation": "The coordinator decomposed a broad topic only into visual-art subtasks, missing music, literature, and film entirely. Since subagents executed their assignments correctly, the narrow decomposition is the obvious root cause."}, {"letter": "D", "text": "The web-search agent’s queries are insufficient and should be broadened to cover more sectors.", "correct": false, "explanation": ""}], "correct": "C", "task_id": "1.6", "objective": "Design task decomposition strategies for complex workflows", "group": "A"}, {"id": "g5", "domain": 1, "scenario": "Multi-agent Research System", "situation": "The web-search subagent returns results for only 3 of 5 requested source categories (competitor sites and industry reports succeed, but news archives and social feeds time out). The document analysis subagent successfully processes all provided documents. The synthesis subagent must produce a summary from mixed-quality upstream inputs. Which error-propagation strategy is most effective?", "question": "Which error-propagation strategy is most effective?", "options": [{"letter": "A", "text": "Continue synthesis using only successful sources and produce an output without mentioning which data was unavailable.", "correct": false, "explanation": ""}, {"letter": "B", "text": "The synthesis subagent returns an error to the coordinator, triggering a full retry or task failure due to incomplete data.", "correct": false, "explanation": ""}, {"letter": "C", "text": "The synthesis subagent asks the coordinator to retry timed-out sources with a longer timeout before starting synthesis.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Structure the synthesis output with coverage annotations that indicate which conclusions are well-supported and where gaps exist due to unavailable sources.", "correct": true, "explanation": "Coverage annotations implement graceful degradation with transparency, preserving value from completed work while propagating uncertainty to enable informed decisions about confidence."}], "correct": "D", "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "group": "B"}, {"id": "g6", "domain": 1, "scenario": "Multi-agent Research System", "situation": "The document analysis subagent encounters a corrupted PDF file that it cannot parse. When designing the system’s error handling, what is the most effective way to handle this failure?", "question": "Which approach is most effective?", "options": [{"letter": "A", "text": "Return an error with context to the coordinator agent, allowing it to decide how to proceed.", "correct": true, "explanation": "Returning an error with context to the coordinator is the most effective approach because it lets the coordinator make an informed decision—skip the file, try an alternative parsing method, or notify the user—while maintaining visibility into the failure."}, {"letter": "B", "text": "Silently skip the corrupted document and continue processing the remaining files to avoid interrupting the workflow.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Automatically retry parsing the document three times with exponential backoff before reporting a failure.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Throw an exception that terminates the entire research workflow.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "group": "J"}, {"id": "g8", "domain": 1, "scenario": "Multi-agent Research System", "situation": "A colleague proposes that the document analysis agent should send its results directly to the synthesis agent, bypassing the coordinator. What is the main advantage of keeping the coordinator as the central hub for all communication between subagents?", "question": "What is the main advantage of keeping the coordinator as the central hub?", "options": [{"letter": "A", "text": "The coordinator can observe all interactions, handle errors uniformly, and decide what information each subagent should receive.", "correct": true, "explanation": "The coordinator pattern provides centralized visibility into all interactions, uniform error handling across the system, and fine-grained control over what information each subagent receives—these are the primary advantages of a star-shaped communication topology."}, {"letter": "B", "text": "The coordinator batches multiple requests to subagents, reducing total API calls and overall latency.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Routing through the coordinator enables automatic retry logic that direct inter-agent calls cannot support.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Subagents use isolated memory, and direct communication would require complex serialization that only the coordinator can perform.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "1.2", "objective": "Orchestrate multi-agent systems with coordinator-subagent patterns", "group": "A"}, {"id": "g9", "domain": 1, "scenario": "Multi-agent Research System", "situation": "The web-search subagent times out while researching a complex topic. You need to design how information about this failure is returned to the coordinator. Which error-propagation approach best enables intelligent recovery?", "question": "Which error-propagation approach best enables intelligent recovery?", "options": [{"letter": "A", "text": "Return structured error context to the coordinator including the failure type, the query executed, any partial results, and potential alternative approaches.", "correct": true, "explanation": "Returning structured error context—including failure type, executed query, partial results, and alternative approaches—gives the coordinator everything needed to make intelligent recovery decisions (e.g., retry with a modified query or continue with partial results). It preserves maximum context for informed coordination-level decision-making."}, {"letter": "B", "text": "Catch the timeout within the subagent and return an empty result set marked as successful.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Implement automatic exponential-backoff retries inside the subagent, only returning a generic “search unavailable” status after exhausting retries.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Propagate the timeout exception directly to the top-level handler, terminating the entire research workflow.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "group": "J"}, {"id": "g11", "domain": 1, "scenario": "Multi-agent Research System", "situation": "While researching a broad topic, you observe that the web-search agent and the document analysis agent investigate the same subtopics, leading to substantial duplication in their outputs. Token usage nearly doubles without a proportional increase in research breadth or depth. What is the most effective way to address this?", "question": "What is the most effective way to address this?", "options": [{"letter": "A", "text": "Allow both agents to finish in parallel, then have the coordinator deduplicate overlapping results before passing them to the synthesis agent.", "correct": false, "explanation": ""}, {"letter": "B", "text": "The coordinator explicitly partitions the research space before delegating, assigning each agent distinct subtopics or source types.", "correct": true, "explanation": "Having the coordinator explicitly partition the research space before delegating is most effective because it addresses the root cause—unclear task boundaries—before any work begins. It preserves parallelism while preventing duplicated effort and wasted tokens."}, {"letter": "C", "text": "Implement a shared-state mechanism where agents log their current focus area so other agents can dynamically avoid duplication during execution.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Switch to sequential execution where document analysis runs only after web search completes, using web-search results as context to avoid duplication.", "correct": false, "explanation": ""}], "correct": "B", "task_id": "1.6", "objective": "Design task decomposition strategies for complex workflows", "group": "A"}, {"id": "g12", "domain": 1, "scenario": "Multi-agent Research System", "situation": "During research, the web-search subagent queries three source categories with different outcomes: academic databases return 15 relevant papers, industry reports return “0 results,” and patent databases return “Connection timeout.” When designing error propagation to the coordinator, which approach enables the best recovery decisions?", "question": "Which approach enables the best recovery decisions?", "options": [{"letter": "A", "text": "Aggregate the results into a single success-percentage metric (e.g., “67% source coverage”) with detailed logs available on demand.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Report both “timeout” and “0 results” as failures requiring coordinator intervention.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Retry transient failures internally and report only persistent errors.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Distinguish access failures (timeout) that require a retry decision from valid empty results (“0 results”) that represent successful queries.", "correct": true, "explanation": "A timeout (access failure) and “0 results” (valid empty result) are semantically different outcomes requiring different responses. Distinguishing them allows the coordinator to retry the patent database while accepting the industry reports “0 results” as a valid, informative finding."}], "correct": "D", "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "group": "J"}, {"id": "g47", "domain": 1, "scenario": "Customer Support Agent", "situation": "Your agent handles single-issue requests with 94% accuracy (e.g., “I need a refund for order #1234”). But when customers include multiple issues in one message (e.g., “I need a refund for order #1234 and also want to update the shipping address for order #5678”), tool selection accuracy drops to 58%. The agent usually solves only one issue or mixes parameters across requests. What approach most effectively improves reliability for multi-issue requests?", "question": "What approach is most effective?", "options": [{"letter": "A", "text": "Implement a preprocessing layer that uses a separate model call to decompose multi-issue messages into separate requests, handle each independently, and merge results.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Combine related tools into fewer universal tools.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Add few-shot examples to the prompt demonstrating correct reasoning and tool sequencing for multi-issue requests.", "correct": true, "explanation": "Few-shot examples that demonstrate correct reasoning and tool sequencing for multi-issue requests are most effective because the agent already performs well on single issues—what it needs is guidance on the pattern for decomposing and routing multiple issues and keeping parameters separated."}, {"letter": "D", "text": "Implement response validation that detects incomplete answers and automatically reprompts the agent to resolve missed issues.", "correct": false, "explanation": ""}], "correct": "C", "task_id": "3.5", "objective": "Apply iterative refinement techniques for progressive improvement", "group": "I"}, {"id": "g48", "domain": 1, "scenario": "Customer Support Agent", "situation": "Production logs show that for simple requests like “refund for order #1234,” your agent resolves the issue in 3–4 tool calls with 91% success. But for complex requests like “I was billed twice, my discount didn’t apply, and I want to cancel,” the agent averages 12+ tool calls with only 54% success—often investigating issues sequentially and fetching redundant customer data for each. What change most effectively improves handling of complex requests?", "question": "What change is most effective?", "options": [{"letter": "A", "text": "Add explicit verification checkpoints between stages, requiring the agent to record progress after resolving each issue before moving to the next.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Reduce the number of tools by combining `get_customer`, `lookup_order`, and billing-related tools into a single `investigate_issue` tool.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Decompose the request into separate issues, then investigate each in parallel using shared customer context before synthesizing a final resolution.", "correct": true, "explanation": "Decomposing into separate issues and investigating in parallel with shared customer context fixes both key problems: it eliminates redundant data retrieval by reusing shared context across issues and reduces total tool-call loops by parallelizing investigation before synthesizing a single resolution."}, {"letter": "D", "text": "Add few-shot examples to the system prompt demonstrating ideal tool-call sequences for various multi-faceted billing scenarios.", "correct": false, "explanation": ""}], "correct": "C", "task_id": "1.1", "objective": "Design and implement agentic loops for autonomous task execution", "group": "A"}, {"id": "g50", "domain": 1, "scenario": "Customer Support Agent", "situation": "After calling `get_customer` and `lookup_order`, the agent has all available system data but still faces uncertainty. Which situation is the most justified trigger for calling `escalate_to_human`?", "question": "Which situation is most justified for escalation?", "options": [{"letter": "A", "text": "A customer wants to cancel an order shipped yesterday and arriving tomorrow. The agent should escalate because the customer might change their mind after receiving the package.", "correct": false, "explanation": ""}, {"letter": "B", "text": "A customer claims they didn’t receive an order, but tracking shows it was delivered and signed for at their address three days ago. The agent should escalate because presenting contradictory evidence could harm the customer relationship.", "correct": false, "explanation": ""}, {"letter": "C", "text": "A customer requests competitor price matching. Your policies allow price adjustments for price drops on your own site within 14 days, but say nothing about competitor prices. The agent should escalate for policy interpretation.", "correct": true, "explanation": "This is a genuine policy gap: company rules cover price drops on your own site but do not address competitor price matching. The agent must not invent policy and should escalate for human judgment on how to interpret or extend existing rules."}, {"letter": "D", "text": "A customer message contains both a billing question and a product return. The agent should escalate so a human can coordinate both issues in one interaction.", "correct": false, "explanation": ""}], "correct": "C", "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "group": "C"}, {"id": "g53", "domain": 1, "scenario": "Customer Support Agent", "situation": "Production metrics show your agent averages 4+ API loops per resolution. Analysis reveals Claude often requests `get_customer` and `lookup_order` in separate sequential turns even when both are needed initially. What is the most effective way to reduce the number of loops?", "question": "What is the most effective way to reduce loops?", "options": [{"letter": "A", "text": "Implement speculative execution that automatically calls likely-needed tools in parallel with any requested tool and returns all results regardless of what was requested.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Increase `max_tokens` to give Claude more room to plan and naturally combine tool requests.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Create composite tools like `get_customer_with_orders` that bundle common lookup combinations into single calls.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Instruct Claude in the prompt to bundle tool requests into one turn and return all results together before the next API call.", "correct": true, "explanation": "Prompting Claude to bundle related tool requests into a single turn leverages its native ability to request multiple tools at once. It directly fixes the sequential-call pattern with minimal architectural change."}], "correct": "D", "task_id": "1.1", "objective": "Design and implement agentic loops for autonomous task execution", "group": "A"}, {"id": "g58", "domain": 1, "scenario": "Customer Support Agent", "situation": "You are implementing the agent loop for your support agent. After each Claude API call, you must decide whether to continue the loop (run requested tools and call Claude again) or stop (present the final answer to the customer). What determines this decision?", "question": "What determines this decision?", "options": [{"letter": "A", "text": "Check the `stop_reason` field in Claude’s response—continue if it is `tool_use` and stop if it is `end_turn`.", "correct": true, "explanation": "`stop_reason` is Claude’s explicit structured signal for loop control: `tool_use` indicates Claude wants to run a tool and receive results back, while `end_turn` indicates Claude has completed its response and the loop should end."}, {"letter": "B", "text": "Parse Claude’s text for phrases like “I’m done” or “Can I help with anything else?”—natural language signals indicate task completion.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Set a maximum iteration count (e.g., 10 calls) and stop when reached, regardless of whether Claude indicates more work is needed.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Check whether the response contains assistant text content—if Claude generated explanatory text, the loop should terminate.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "1.1", "objective": "Design and implement agentic loops for autonomous task execution", "group": "A"}, {"id": "m1", "domain": 1, "scenario": null, "situation": "Your multi-agent research pipeline crashed after processing 12 of 28 documents. The web search agent had identified relevant sources, the document analysis agent had partially completed extraction, and the synthesizer had begun pattern identification. You need to resume processing without repeating work or losing fidelity of prior findings.", "question": "What state management approach best balances information fidelity with context efficiency when restoring agent state?", "options": [{"letter": "A", "text": "Have each agent maintain its own persistent state file and reload it independently at the start of each session.", "correct": false, "explanation": "Fragmented per-agent state breaks coordinator visibility and makes cross-agent reasoning brittle. Each agent reloading unrelated internal state also bloats their context with noise they don't need."}, {"letter": "B", "text": "Persist the coordinator's conversation log containing all task delegations and responses, providing this to agents when resuming.", "correct": false, "explanation": "A full conversation log is the highest-fidelity but least context-efficient option. You'd replay everything into each subagent and blow past context budgets."}, {"letter": "C", "text": "Have each agent persist a structured report to a known location. On resume, the coordinator loads the reports and injects relevant state into agent prompts.", "correct": true, "explanation": "Correct. Structured per-agent reports keep fidelity (the findings, with schema), let the coordinator stay in charge of orchestration, and keep each subagent's context focused. This is the orchestrator + compact artifact pattern."}, {"letter": "D", "text": "Index all agent outputs in a shared vector store. When resuming, each agent queries the store using semantic search to retrieve relevant prior findings.", "correct": false, "explanation": "Semantic retrieval over agent outputs is overkill for resuming a crashed pipeline — it adds a retrieval failure mode and can miss precise state that a structured report preserves exactly."}], "correct": "C", "task_id": "1.7", "objective": "Manage session state, resumption, and forking", "group": "A"}, {"id": "m2", "domain": 1, "scenario": null, "situation": "After the web search agent finds 25 sources (120K tokens of raw content), the document analysis agent extracts key insights (15K tokens), and the synthesis agent produces a coherent narrative draft (3K tokens), the coordinator must pass context to the report generation agent for the final output with proper source citations.", "question": "What context-passing strategy provides the best balance of completeness and efficiency?", "options": [{"letter": "A", "text": "Pass only the synthesis draft and have a separate post-processing pipeline match claims to sources and insert citations after the report is generated.", "correct": false, "explanation": "Post-hoc matching is fragile — without the mapping the model used, you can't reliably bind claims to the right source. Hallucinated or misattributed citations are the usual failure mode."}, {"letter": "B", "text": "Pass the synthesis draft along with a structured source index that maps key claims to their source URLs and relevant excerpts.", "correct": true, "explanation": "Correct. The synthesis gives the narrative; the source index gives the report generator exactly the binding it needs to cite without re-reading 120K tokens of raw content."}, {"letter": "C", "text": "Pass a condensed summary of all prior stages that preserves the main findings and attributes them to sources by name only.", "correct": false, "explanation": "Name-only attribution loses URLs and excerpts, so the report generator can't quote or verify — and 'by name only' tends to drift into vague citations."}, {"letter": "D", "text": "Pass the full accumulated context from all prior agents.", "correct": false, "explanation": "Maximum completeness but wasteful — 120K+ tokens of raw search content is mostly irrelevant noise at the report-generation stage."}], "correct": "B", "task_id": "1.4", "objective": "Implement multi-step workflows with enforcement and handoff patterns", "group": "A"}, {"id": "m4", "domain": 1, "scenario": null, "situation": "The web search agent has gathered several relevant sources for a research topic. The document analysis agent now needs to examine these sources.", "question": "How does information typically flow between these two specialized subagents?", "options": [{"letter": "A", "text": "The agents communicate through an event-driven message queue, with the document analysis agent subscribing to web search completion events.", "correct": false, "explanation": "Event buses aren't part of the Claude subagent model. Subagents don't publish/subscribe to each other directly."}, {"letter": "B", "text": "The web search agent directly invokes the document analysis agent, passing the discovered sources as parameters.", "correct": false, "explanation": "Subagents are isolated — they can't directly call sibling subagents. That would also tightly couple them and defeat the orchestrator pattern."}, {"letter": "C", "text": "The coordinator agent receives the web search agent's output and includes relevant findings in the prompt when invoking the document analysis agent.", "correct": true, "explanation": "Correct. In an orchestrator-worker pattern the coordinator is the hub. It collects each subagent's output and explicitly forwards the relevant parts into the next subagent's prompt."}, {"letter": "D", "text": "Both agents access a shared memory store where the web search agent writes findings and the document analysis agent reads them.", "correct": false, "explanation": "A shared store can be layered in as an optimization, but it's not how subagents typically communicate — and it introduces stale-read and consistency problems."}], "correct": "C", "task_id": "1.2", "objective": "Orchestrate multi-agent systems with coordinator-subagent patterns", "group": "A"}, {"id": "m5", "domain": 1, "scenario": null, "situation": "In production, you observe that simple fact-checking queries (e.g., \"What year was the Paris Climate Agreement signed?\") traverse all four subagents sequentially, consuming 40+ seconds and significant tokens per query. Complex comparative research benefits from the full pipeline. Your query distribution is diverse and evolving as users discover new applications.", "question": "What's the most effective approach to optimize for varying query complexity?", "options": [{"letter": "A", "text": "Implement pattern-based routing that categorizes queries by structure (single-fact vs. comparative vs. analytical) and maps each category to a predefined subagent combination.", "correct": false, "explanation": "Pattern routing calcifies as the query distribution evolves — new intents break the patterns and silently land on the wrong pipeline."}, {"letter": "B", "text": "Create a fast-path for factual questions that bypasses subagents entirely, routing all other queries through the complete pipeline to ensure research thoroughness.", "correct": false, "explanation": "A binary fast-path fixes the worst case but wastes the full pipeline on every moderately complex query that could have used a subset."}, {"letter": "C", "text": "Have the coordinator analyze each query and dynamically decide which subagents to invoke based on its assessment of query requirements.", "correct": true, "explanation": "Correct. Letting the coordinator LLM reason about each query and pick only the subagents it needs adapts naturally to an evolving, diverse query distribution — this is the strength of the orchestrator pattern."}, {"letter": "D", "text": "Train a query complexity classifier on labeled historical data to predict optimal subagent combinations, retraining periodically as query patterns evolve.", "correct": false, "explanation": "A trained classifier needs labels, retraining, and drift monitoring. That's over-engineered when the coordinator can make the same call at runtime from the query itself."}], "correct": "C", "task_id": "1.2", "objective": "Orchestrate multi-agent systems with coordinator-subagent patterns", "group": "A"}, {"id": "m6", "domain": 1, "scenario": null, "situation": "When researching \"renewable energy adoption,\" the web search agent returns recent statistics (2024: 35% adoption) while the document analysis agent extracts data from internal reports (2022: 18% adoption). The synthesis agent incorrectly flags these as contradictory sources rather than recognizing the data shows growth over time.", "question": "What change would best enable the synthesis agent to correctly interpret such temporal differences?", "options": [{"letter": "A", "text": "Require subagents to include publication or data collection dates in their structured outputs.", "correct": true, "explanation": "Correct. The synthesis agent misreads the data because it never sees the dates. Making each data point carry its own timestamp in the structured output lets synthesis reason about trends instead of contradictions."}, {"letter": "B", "text": "Add a conflict resolution agent that automatically discards older data when newer data exists for the same metric.", "correct": false, "explanation": "Silently discarding older data destroys trend information — the very thing the question is asking about."}, {"letter": "C", "text": "Configure the web search agent to only return results from the past 6 months.", "correct": false, "explanation": "Shrinking the window throws away historical context and doesn't fix the architectural gap that metadata isn't being passed to synthesis."}, {"letter": "D", "text": "Instruct the synthesis agent to always treat the most recent data as authoritative and place older findings in a separate historical appendix.", "correct": false, "explanation": "Prompt instructions alone are unreliable and still hide temporal reasoning behind a rule. Structuring the data with dates is the systematic fix."}], "correct": "A", "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "group": "B"}, {"id": "m7", "domain": 1, "scenario": null, "situation": "The synthesis agent receives summarized findings from the web search and document analysis agents, then passes a consolidated summary to the report generator. During testing, you discover the generated reports make factual claims without proper citations—the report generator cannot attribute statements to their original sources because that metadata was lost during the summarization steps.", "question": "What's the most effective approach to ensure proper source attribution in the final reports?", "options": [{"letter": "A", "text": "Have each agent output structured data separating content summaries from source metadata (URLs, document names, page numbers).", "correct": true, "explanation": "Correct. Structured content + separate source metadata preserves the mapping end-to-end, so the report generator receives both what was said and where it came from."}, {"letter": "B", "text": "Have the report generator query the web search agent to re-locate sources for claims in the final report.", "correct": false, "explanation": "Re-querying is slow, lossy, and invites misattribution — the web search agent doesn't know which claim came from which source."}, {"letter": "C", "text": "Instruct the synthesis agent to embed source references inline within its summary text using a consistent citation format.", "correct": false, "explanation": "Inline citations in free text drift and get dropped during later summarization. Structured metadata survives transformations better."}, {"letter": "D", "text": "Skip summarization and pass full raw outputs from web search and document analysis directly to the report generator.", "correct": false, "explanation": "You'd blow the context budget and make the report generator's job harder. Summarization is useful; losing metadata during summarization is the actual bug."}], "correct": "A", "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "group": "B"}, {"id": "m9", "domain": 1, "scenario": null, "situation": "In production, final reports frequently contain claims without proper source attribution. Investigation shows that while the web search and document analysis agents correctly attach citations to their outputs, the synthesis agent loses track of which sources support which conclusions when combining findings.", "question": "What's the most effective architectural change?", "options": [{"letter": "A", "text": "Maintain complete transcripts of all subagent interactions and add a citation-resolution agent to analyze logs and determine attributions before report generation.", "correct": false, "explanation": "Post-hoc log analysis is fragile and expensive — and still loses bindings when synthesis merges points from multiple sources."}, {"letter": "B", "text": "Require all subagents to output structured claim-source mappings that the synthesis agent must preserve and merge when combining findings from multiple sources.", "correct": true, "explanation": "Correct. Explicit claim-source mappings are a first-class output the synthesis agent can merge deterministically — no attribution gets dropped during summarization."}, {"letter": "C", "text": "Add a verification step where the report generator uses semantic similarity matching against original sources to reconstruct which claims came from which documents.", "correct": false, "explanation": "Similarity search can attribute the wrong source when two sources say similar things. You want the model's original mapping, not a reconstruction."}, {"letter": "D", "text": "Have the coordinator inject source identifier prefixes into text before each handoff, then parse these prefixes at report generation to reconstruct citations.", "correct": false, "explanation": "Prefix tokens inside prose get dropped, paraphrased, or hallucinated. Structured mappings outside the prose are more robust."}], "correct": "B", "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "group": "B"}, {"id": "m10", "domain": 1, "scenario": null, "situation": "After the web search agent and document analysis agent complete their tasks, the coordinator invokes the synthesis agent. However, the synthesis agent responds that it cannot complete the task because no research findings were provided.", "question": "What is the most likely cause of this issue?", "options": [{"letter": "A", "text": "The synthesis agent's context window is not large enough to hold the combined outputs from both previous agents.", "correct": false, "explanation": "A context-window overflow typically surfaces as a truncation or API error, not as the agent saying 'no findings provided.'"}, {"letter": "B", "text": "The coordinator did not include the outputs from the previous agents in the synthesis agent's prompt.", "correct": true, "explanation": "Correct. Subagent invocations are isolated — nothing flows between them unless the coordinator explicitly puts it in the prompt. The message 'no findings provided' is exactly what you'd see."}, {"letter": "C", "text": "The subagents need to share a single API connection to enable automatic context sharing between invocations.", "correct": false, "explanation": "There's no 'automatic context sharing' over a shared connection. Context isolation is by design."}, {"letter": "D", "text": "The synthesis agent needs tools that can fetch results directly from the other agents' conversation histories.", "correct": false, "explanation": "Cross-agent history fetching isn't a standard capability and isn't the right fix — the coordinator should be forwarding findings explicitly."}], "correct": "B", "task_id": "1.3", "objective": "Configure subagent invocation, context passing, and spawning", "group": "B"}, {"id": "m12", "domain": 1, "scenario": null, "situation": "The coordinator provides detailed step-by-step instructions to the web search subagent, specifying exact search queries, source priorities, and date filters. Production monitoring reveals three issues: (1) the subagent reports \"insufficient results\" rather than trying alternative approaches when pre-specified searches fail, (2) research quality drops for emerging topics that don't match expected patterns, and (3) the subagent rarely surfaces valuable tangential sources.", "question": "What's the most effective way to improve subagent adaptability?", "options": [{"letter": "A", "text": "Remove procedural details entirely, delegating with simple goals like \"research X thoroughly\" and relying on the subagent's general capabilities.", "correct": false, "explanation": "Too far in the other direction — 'research X thoroughly' loses the guardrails (source quality, recency) that make delegation reliable."}, {"letter": "B", "text": "Add explicit fallback directives to the detailed instructions: \"If specified searches yield fewer than N results, attempt alternative query formulations before reporting failure.\"", "correct": false, "explanation": "Patches one failure mode but keeps the agent locked into procedural thinking — still brittle for emerging topics and tangential sources."}, {"letter": "C", "text": "Implement a topic classification step where the coordinator categorizes requests as \"well-defined\" or \"exploratory\" and uses different instruction styles for each category.", "correct": false, "explanation": "A classifier with two buckets is fragile, and you're still shipping rigid instructions in the well-defined branch."}, {"letter": "D", "text": "Specify research goals and quality criteria (coverage breadth, source diversity, recency) rather than procedural steps, letting the subagent determine its search strategy.", "correct": true, "explanation": "Correct. Delegate intent and quality bars, not procedures. The subagent can then choose queries, follow promising tangents, and recover from dead ends on its own."}], "correct": "D", "task_id": "1.3", "objective": "Configure subagent invocation, context passing, and spawning", "group": "B"}, {"id": "m13", "domain": 1, "scenario": null, "situation": "Production monitoring shows that follow-up queries like \"summarize what we learned about market trends\" consistently take 40+ seconds. Investigation reveals the coordinator spawns the synthesis subagent for each summarization request, passing 80K+ tokens of accumulated findings. The coordinator already has these findings in its context from orchestrating the research.", "question": "What's the most effective way to improve response time for these follow-up summaries?", "options": [{"letter": "A", "text": "Pre-generate and cache summaries at multiple granularities whenever new findings accumulate.", "correct": false, "explanation": "Speculative generation wastes tokens for summaries the user may never ask for, and cache granularities never quite match what's asked."}, {"letter": "B", "text": "Have the coordinator handle straightforward summarization requests directly using its existing context, reserving subagent spawning for complex analysis.", "correct": true, "explanation": "Correct. If the coordinator already has the findings, spawning a subagent to re-ingest 80K tokens is pure overhead. Let the coordinator answer simple follow-ups itself."}, {"letter": "C", "text": "Enable prompt caching on the synthesis subagent to reduce the overhead of repeatedly transferring the same research findings.", "correct": false, "explanation": "Prompt caching trims cost on the repeated prefix but the architectural waste — spawning an entire subagent for a summary — remains."}, {"letter": "D", "text": "Spawn the synthesis subagent with reduced context and have it request specific findings from the coordinator on-demand.", "correct": false, "explanation": "Adds round-trips and complexity for the same result the coordinator could produce directly."}], "correct": "B", "task_id": "1.2", "objective": "Orchestrate multi-agent systems with coordinator-subagent patterns", "group": "A"}, {"id": "m14", "domain": 1, "scenario": null, "situation": "When analyzing complex legal cases that cite multiple precedents, the document analysis subagent processes each sequentially. A landmark case citing 12 precedents takes over 3 minutes to analyze completely.", "question": "What's the most effective way to reduce this latency while preserving the coordinator's ability to monitor and debug the system?", "options": [{"letter": "A", "text": "Implement a message queue where precedent analysis tasks are processed asynchronously by a pool of worker agents.", "correct": false, "explanation": "External queues complicate observability — the coordinator loses direct visibility into which tasks succeeded and what they produced."}, {"letter": "B", "text": "Create a recursive agent hierarchy where analysis agents subdivide work among child agents until reaching single-precedent granularity.", "correct": false, "explanation": "Recursion adds levels of indirection that make debugging and monitoring harder, for no real speedup beyond the first fan-out."}, {"letter": "C", "text": "Have the coordinator spawn parallel document analysis subagents, each handling a subset of precedents, then aggregate results before synthesis.", "correct": true, "explanation": "Correct. Coordinator-managed parallelism fans out the work, keeps each subagent's scope tight, and preserves a single hub for monitoring and aggregation."}, {"letter": "D", "text": "Enable the document analysis subagent to spawn its own specialized subagents dynamically when it encounters cases with many citations.", "correct": false, "explanation": "Nested spawning hides execution inside subagents and makes the coordinator's debug view incomplete."}], "correct": "C", "task_id": "1.2", "objective": "Orchestrate multi-agent systems with coordinator-subagent patterns", "group": "A"}, {"id": "m15", "domain": 1, "scenario": null, "situation": "The coordinator agent has `AgentDefinitions` configured for all four specialized subagents, each with appropriate descriptions, prompts, and tool restrictions. During testing, you notice the coordinator correctly reasons about when to delegate—it generates messages like \"I'll ask the web search agent to find sources on this topic\"—but no subagent execution ever occurs. The coordinator then proceeds as if the delegation happened and continues with incomplete information. Logs show no errors.", "question": "What is the most likely cause?", "options": [{"letter": "A", "text": "The coordinator's `max_tokens` setting is too low, causing the Task tool invocation to be truncated before the subagent type parameter can be specified.", "correct": false, "explanation": "A truncation would show up in logs and usually leave partial `tool_use` blocks — not silent no-ops."}, {"letter": "B", "text": "The `AgentDefinitions` are configured correctly, but the coordinator's system prompt doesn't explicitly list the available subagent types, preventing the model from knowing they can be invoked.", "correct": false, "explanation": "Tool/agent schemas are surfaced to the model automatically — you don't need to re-list them in the system prompt for the model to see them."}, {"letter": "C", "text": "The coordinator's allowedTools configuration doesn't include \"Task\", so while it can reason about delegation, it cannot invoke the tool required to spawn subagents.", "correct": true, "explanation": "Correct. Without the Task tool in allowedTools, the coordinator can talk about delegating but has no way to actually call a subagent — which matches the 'reasons about it, no execution, no errors' symptom."}, {"letter": "D", "text": "Subagent context isolation means task descriptions from the coordinator don't automatically reach subagents; you need to configure explicit context forwarding in ClaudeAgentOptions.", "correct": false, "explanation": "Context isolation is real, but it affects what the subagent sees once spawned — not whether the coordinator can spawn it at all."}], "correct": "C", "task_id": "1.3", "objective": "Configure subagent invocation, context passing, and spawning", "group": "B"}, {"id": "m21", "domain": 1, "scenario": null, "situation": "Your codebase exploration tool stores session IDs to allow engineers to continue investigations across work sessions. An engineer spent an hour yesterday analyzing a legacy authentication module, building context about its architecture and dependencies. They want to continue today. The session ID is valid, but version control shows 3 of the 12 files the agent previously read were modified overnight by a teammate's merge.", "question": "What approach best balances efficiency and accuracy?", "options": [{"letter": "A", "text": "Resume the session without informing the agent about the changed files", "correct": false, "explanation": "Silent resume leaves the agent reasoning on stale content for 3 of 12 files — exactly the source of bad recommendations."}, {"letter": "B", "text": "Start a fresh session to ensure the agent works with current codebase state without stale assumptions", "correct": false, "explanation": "Throws away an hour of valid context about 9 files that didn't change."}, {"letter": "C", "text": "Resume the session and inform the agent which specific files changed for targeted re-analysis", "correct": true, "explanation": "Correct. Keeps the expensive context you already built, while telling the agent exactly which 3 files to re-read — minimum waste, maximum accuracy."}, {"letter": "D", "text": "Resume the session and immediately have the agent re-read all 12 previously analyzed files", "correct": false, "explanation": "Unnecessary for the 9 unchanged files. Just adds tokens without improving accuracy."}], "correct": "C", "task_id": "1.7", "objective": "Manage session state, resumption, and forking", "group": "A"}, {"id": "m22", "domain": 1, "scenario": null, "situation": "An engineer used the agent yesterday to analyze a legacy authentication module, identifying two distinct refactoring approaches: extracting a microservice versus refactoring in-place. Today, they want to explore both approaches in depth—having the agent propose specific code changes for each—before deciding which to implement.", "question": "What's the most effective way to structure this exploration?", "options": [{"letter": "A", "text": "Resume yesterday's session to explore the first approach, then start a new session for the second, manually recreating the original context.", "correct": false, "explanation": "Manual recreation is error-prone and loses the exact working state of yesterday's analysis."}, {"letter": "B", "text": "Start two fresh sessions, manually providing a summary of yesterday's analysis findings to establish context.", "correct": false, "explanation": "Redoes work and risks the two sessions diverging from the same baseline you established yesterday."}, {"letter": "C", "text": "Resume yesterday's session and explore both approaches sequentially within the same conversation thread.", "correct": false, "explanation": "Sequential exploration in one thread lets each approach contaminate the other's context."}, {"letter": "D", "text": "Use `fork_session` to create two branches from yesterday's analysis, exploring one approach in each fork.", "correct": true, "explanation": "Correct. Forking from yesterday's session gives each approach its own independent context starting from the same analysis baseline — clean, parallel, no contamination."}], "correct": "D", "task_id": "1.7", "objective": "Manage session state, resumption, and forking", "group": "A"}, {"id": "m24", "domain": 1, "scenario": null, "situation": "An engineer asks your agent to identify untested code paths in a legacy payment processing module spanning 45 files. After reading the first 8 source files, the agent's responses are becoming noticeably less accurate—it's forgetting previously discussed code patterns and hasn't yet located all test files or traced critical payment flows.", "question": "What's the most effective approach to complete this investigation?", "options": [{"letter": "A", "text": "Document all current findings in a summary report, clear context completely, then use that report as the sole reference for continuing the investigation.", "correct": false, "explanation": "A single report becomes the only source of truth and tends to compress away the specific code patterns you'd need later."}, {"letter": "B", "text": "Spawn subagents to investigate specific questions (e.g., \"find all test files for payment processing\", \"trace refund flow dependencies\") while the main agent coordinates findings and preserves high-level understanding.", "correct": true, "explanation": "Correct. Delegate well-scoped investigations to subagents with fresh context, while the main agent keeps the architectural overview. This is the pattern for scaling exploration beyond a single context window."}, {"letter": "C", "text": "Clear context with /clear, then selectively re-read only the most critical files discovered so far, writing key findings to a scratchpad file that persists between context resets.", "correct": false, "explanation": "Scratchpads help, but clearing + re-reading discards the understanding you already built across 8 files."}, {"letter": "D", "text": "Switch to using Grep to search for specific function names instead of reading full files, reducing the content loaded into context for remaining exploration.", "correct": false, "explanation": "You can't identify untested paths just by grepping names — you need to read enough of each path to know what branches the tests don't cover."}], "correct": "B", "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "group": "G"}, {"id": "m25", "domain": 1, "scenario": null, "situation": "A developer asks the agent to investigate why a specific API endpoint intermittently returns 500 errors. The codebase has 200+ files and the developer doesn't know which components are involved. The agent must trace the error through routing, middleware, business logic, and database layers.", "question": "What task decomposition approach would be most effective?", "options": [{"letter": "A", "text": "Have the agent first create a comprehensive plan mapping all code paths through the endpoint before beginning any file exploration or code reading.", "correct": false, "explanation": "You can't build a correct plan for an unknown error without any exploration. Planning blind wastes time and misses the actual failure path."}, {"letter": "B", "text": "Have the agent dynamically generate investigation subtasks based on what it discovers at each step, adapting its exploration plan as new information about the error path emerges.", "correct": true, "explanation": "Correct. Debugging is adaptive by nature — each file you read changes the most useful next step. Let the agent follow the evidence."}, {"letter": "C", "text": "Define a fixed sequence of investigation steps upfront—grep for error patterns, then read error handlers, then check database queries, then examine middleware—executing each step regardless of intermediate findings.", "correct": false, "explanation": "A fixed pipeline wastes work on layers that aren't involved and can lock the agent out of the actual root cause path."}, {"letter": "D", "text": "Run parallel worker agents that simultaneously investigate all four layers, then synthesize their findings to identify where the error originates.", "correct": false, "explanation": "Fan-out is useful once you have well-scoped subtasks. Here, you don't — you'd pay 4× the cost to look in three layers that aren't the problem."}], "correct": "B", "task_id": "1.6", "objective": "Design task decomposition strategies for complex workflows", "group": "A"}, {"id": "m26", "domain": 1, "scenario": null, "situation": "An engineer's exploration subagent spent 30 minutes analyzing a legacy payment system, reading 47 files and documenting data flows. The session was interrupted when the engineer's connection dropped. While away, a teammate merged a PR that renamed two utility functions. The engineer wants to continue the same exploration.", "question": "What's the most effective approach?", "options": [{"letter": "A", "text": "Resume the subagent from its previous transcript without mentioning the changes—the architecture understanding remains valid.", "correct": false, "explanation": "Silently resuming lets the agent keep referencing the old function names in its recommendations."}, {"letter": "B", "text": "Launch a fresh subagent and include the prior transcript in the initial prompt for context.", "correct": false, "explanation": "Loading 30 minutes of transcript into a new subagent's prompt is wasteful and pollutes its starting context."}, {"letter": "C", "text": "Launch a fresh subagent with a summary of prior findings.", "correct": false, "explanation": "Re-summarization throws away the fine-grained data flows the agent was already tracking across 47 files."}, {"letter": "D", "text": "Resume the subagent from its previous transcript and inform it about the renamed functions.", "correct": true, "explanation": "Correct. Keep the accumulated understanding, and give it a targeted delta about the renames so it can update its mental model — minimum waste, maximum accuracy."}], "correct": "D", "task_id": "1.7", "objective": "Manage session state, resumption, and forking", "group": "A"}, {"id": "m28", "domain": 1, "scenario": null, "situation": "Your agent has analyzed a complex service module—reading 23 source files, tracing request flows, and identifying error handling patterns. A developer wants to compare two testing strategies before committing to one: end-to-end tests with mocked external services vs. snapshot tests capturing expected outputs. They need to independently develop both approaches to evaluate trade-offs.", "question": "How should you manage the sessions?", "options": [{"letter": "A", "text": "Export the analysis session's key findings to a file, then create two new sessions that reference this file.", "correct": false, "explanation": "An exported summary is lossier than the live session state the agent built up while reading 23 files."}, {"letter": "B", "text": "Resume the analysis session with `fork_session` enabled, creating a separate branch for each testing strategy.", "correct": true, "explanation": "Correct. Forking gives each strategy its own independent context starting from the exact analysis baseline — no cross-contamination, no re-analysis."}, {"letter": "C", "text": "Start two fresh sessions, having each re-read the relevant source files before beginning.", "correct": false, "explanation": "Burns tokens redoing work the original session already did, and risks each fresh session reaching different conclusions from the same code."}, {"letter": "D", "text": "Continue in the original session, developing end-to-end tests first, then snapshot tests sequentially.", "correct": false, "explanation": "Sequential development in one thread lets the first strategy's implementation bias reasoning about the second."}], "correct": "B", "task_id": "1.7", "objective": "Manage session state, resumption, and forking", "group": "A"}, {"id": "m30", "domain": 1, "scenario": null, "situation": "An engineer who just joined the team asks the agent to help them understand the authentication and authorization architecture before making security improvements. The codebase has 800+ files across multiple services.", "question": "What exploration strategy will most effectively build understanding, given Claude built-in tools and context limits?", "options": [{"letter": "A", "text": "Read any CLAUDE.md and README files first, then ask the engineer to specify which 10-15 files are most important for understanding the auth system.", "correct": false, "explanation": "The engineer just joined — they probably don't know which files matter. That's exactly what the agent is supposed to help with."}, {"letter": "B", "text": "Launch parallel subagents to explore different services simultaneously, then synthesize their findings into an architectural overview.", "correct": false, "explanation": "Without knowing where auth lives, fan-out explores too broadly and subagents duplicate and miss cross-service flows."}, {"letter": "C", "text": "Use Grep to find authentication entry points, read those files, then follow imports and function calls to map the auth flow incrementally.", "correct": true, "explanation": "Correct. Start at entry points (login, token verify, middleware), then trace outward following real code edges. Incremental, grounded, fits within context limits."}, {"letter": "D", "text": "Read all files containing \"auth\", \"login\", \"permission\", or \"token\" in their content or filename.", "correct": false, "explanation": "Those keywords hit a huge amount of unrelated code across 800+ files and drown context in noise."}], "correct": "C", "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "group": "G"}, {"id": "m31", "domain": 1, "scenario": null, "situation": "A customer returns 4 hours after their initial session about the same billing dispute. The previous 32-turn session contains `lookup_order` results showing \"Status: PENDING, Expected resolution: 24-48 hours.\" In testing, you observe that when resuming sessions with stale tool results, the agent often references the outdated data in responses (e.g., \"I see your refund is still being processed\") even after subsequent fresh tool calls return different information.", "question": "What approach most reliably handles returning customers?", "options": [{"letter": "A", "text": "Resume with full history but filter out previous `tool_result` messages before resuming, keeping only the human/assistant turns so the agent must re-fetch needed data.", "correct": false, "explanation": "Stripping `tool_results` from the middle of a conversation can leave assistant messages referencing nonexistent results — the transcript becomes internally inconsistent."}, {"letter": "B", "text": "Start a new session, inject a structured summary of the previous interaction (issue type, actions taken, resolution status), then make fresh tool calls before engaging.", "correct": true, "explanation": "Correct. A clean session with a summary keeps the narrative continuity while guaranteeing the agent isn't reasoning over stale tool results."}, {"letter": "C", "text": "Resume with full history and add a system prompt instruction telling the agent to always prefer the most recent tool results when multiple calls to the same tool exist in context.", "correct": false, "explanation": "Prompt instructions are suggestions. You observed that exact failure mode in testing — the model still references old results."}, {"letter": "D", "text": "Resume with full history and configure the agent to automatically re-call all previously-used tools at session start to ensure data freshness.", "correct": false, "explanation": "Blanket re-calling is wasteful, slow, and still leaves the old results sitting in context to confuse the model."}], "correct": "B", "task_id": "1.4", "objective": "Implement multi-step workflows with enforcement and handoff patterns", "group": "A"}, {"id": "m32", "domain": 1, "scenario": null, "situation": "You're implementing the escalation logic for when the agent should call `escalate_to_human`. Your team proposes four different approaches for triggering escalation.", "question": "Which approach will most reliably identify cases that genuinely require human intervention?", "options": [{"letter": "A", "text": "Instruct the agent to escalate when the customer requests a human, when the issue requires policy exceptions, or when the agent cannot make meaningful progress.", "correct": true, "explanation": "Correct. Escalation decisions are judgment calls about intent and progress — exactly what LLMs are good at. Clear criteria in natural language outperform rigid rules for the long tail."}, {"letter": "B", "text": "Configure the agent to escalate after three consecutive tool calls that fail to resolve the customer's stated issue, ensuring a reasonable attempt before involving a human.", "correct": false, "explanation": "A hard retry count fires both too early (legitimate retries) and too late (obvious policy issues on the first call)."}, {"letter": "C", "text": "Implement sentiment analysis that monitors for frustration indicators (negative language, repeated questions, exclamation marks) and trigger escalation when the frustration score exceeds a configured threshold.", "correct": false, "explanation": "Sentiment can catch frustration but misses calm customers who simply need a human for a policy exception, and over-escalates on stylistic language."}, {"letter": "D", "text": "Build a rules engine that maps specific issue types, customer segments, and product categories to escalation decisions, removing the need for model judgment calls.", "correct": false, "explanation": "Rules engines break on the cases they weren't designed for — and customer support is full of those."}], "correct": "A", "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "group": "C"}, {"id": "m33", "domain": 1, "scenario": null, "situation": "After investigating a billing dispute over 25+ turns, you've identified that duplicate charges occurred due to a payment gateway timeout triggering retry logic. The required refund ($847) exceeds your $500 authorization limit. You need to call `escalate_to_human`, and the human agent won't have access to your conversation transcript.", "question": "What context should you pass to enable effective resolution?", "options": [{"letter": "A", "text": "The customer's original complaint verbatim plus the tool result excerpts showing duplicate transactions.", "correct": false, "explanation": "Raw artifacts without synthesis force the human to re-do the 25 turns of investigation you just finished."}, {"letter": "B", "text": "A structured summary: customer ID, root cause, refund amount, and recommended action.", "correct": true, "explanation": "Correct. A structured handoff with identifiers, cause, amount, and recommended action is what a human agent needs to pick up the case instantly without re-investigating."}, {"letter": "C", "text": "The complete conversation transcript with all tool results.", "correct": false, "explanation": "Dumping the whole transcript forces the human to wade through 25 turns instead of reading a one-screen brief."}, {"letter": "D", "text": "Your diagnosis and the refund amount only.", "correct": false, "explanation": "Missing customer identifiers and recommended action — the human can't act without them."}], "correct": "B", "task_id": "1.4", "objective": "Implement multi-step workflows with enforcement and handoff patterns", "group": "A"}, {"id": "m34", "domain": 1, "scenario": null, "situation": "Compliance requires that refunds exceeding $500 must automatically escalate to a human agent—this rule cannot be left to model discretion. Despite clear system prompt instructions, production logs show the agent occasionally processes high-value refunds directly (3% failure rate).", "question": "How should you achieve guaranteed compliance?", "options": [{"letter": "A", "text": "Modify the refund tool to return an error with message \"Amount exceeds policy limit—please escalate\" when threshold is exceeded.", "correct": false, "explanation": "This helps, but depends on the agent interpreting the error correctly and escalating — still model discretion at the decision point."}, {"letter": "B", "text": "Add few-shot examples to the prompt showing correct escalation behavior at various refund amounts ($400, $500, $600).", "correct": false, "explanation": "Few-shots shift the distribution but don't remove the 3% failure. Compliance says the rule can't be left to model discretion."}, {"letter": "C", "text": "Implement a hook to intercept tool calls; when the refund process amount exceeds $500, block it and invoke human escalation.", "correct": true, "explanation": "Correct. Compliance-grade rules belong outside the model — a deterministic hook on the tool call is guaranteed to fire every time, independent of model behavior."}, {"letter": "D", "text": "Strengthen the system prompt with emphatic language: \"CRITICAL POLICY: Refunds over $500 MUST trigger human escalation. NEVER process these directly.\"", "correct": false, "explanation": "Emphatic prompts reduce but don't eliminate misuse. 'Cannot be left to model discretion' rules out any prompt-only solution."}], "correct": "C", "task_id": "1.5", "objective": "Apply Agent SDK hooks for tool call interception and data normalization", "group": "E"}, {"id": "m35", "domain": 1, "scenario": null, "situation": "During a billing dispute resolution, your agent successfully retrieves customer info via `get_customer` and order details via `lookup_order`, but when attempting to call `process_refund`, the tool returns a timeout error. The agent has enough information to explain the charges and verify refund eligibility, but cannot actually process the refund due to the backend failure.", "question": "What approach best balances first-contact resolution with appropriate error handling?", "options": [{"letter": "A", "text": "Escalate immediately to a human agent since the refund action cannot be completed", "correct": false, "explanation": "Punts a transient backend issue to a human when the agent could still explain the situation and offer options."}, {"letter": "B", "text": "Implement automatic retries with exponential backoff for `process_refund`, keeping the conversation open until the refund is successfully processed", "correct": false, "explanation": "Forcing the customer to wait on indefinite retries is poor UX and may never succeed if the outage is long-lived."}, {"letter": "C", "text": "Explain the billing, confirm refund eligibility, acknowledge the system issue preventing immediate processing, and offer escalation or retry later", "correct": true, "explanation": "Correct. Deliver the partial value you can (explanation + eligibility), be honest about the failure, and let the customer choose between human escalation or a retry. Classic graceful degradation."}, {"letter": "D", "text": "Confirm the refund will be processed and close the conversation, since the system has all necessary information to complete it automatically", "correct": false, "explanation": "Committing to an outcome that didn't happen is misleading the customer — a bigger failure than the timeout."}], "correct": "C", "task_id": "1.4", "objective": "Implement multi-step workflows with enforcement and handoff patterns", "group": "A"}, {"id": "m36", "domain": 1, "scenario": null, "situation": "A customer writes: \"I've been going back and forth on this return for days. I just want to speak to someone who can actually help me.\" The agent has confirmed via `lookup_order` that the return is straightforward—within policy and eligible for immediate processing.", "question": "What should the agent do?", "options": [{"letter": "A", "text": "Acknowledge frustration, inform them this is resolvable now, and offer to complete it or escalate", "correct": true, "explanation": "Correct. Honor the feeling, give them the fast resolution path in writing, and preserve their choice. That's the customer-respect move that still leverages the agent's capability."}, {"letter": "B", "text": "Call `escalate_to_human` immediately to honor the customer's request", "correct": false, "explanation": "Unnecessary queueing when the issue is one tool call away. Frustrates the customer further by adding waiting to an already-simple case."}, {"letter": "C", "text": "Process the refund via `process_refund` to resolve the underlying issue, then inform them it's complete", "correct": false, "explanation": "Takes unilateral action after the customer explicitly asked to speak to someone — overrides their stated preference."}, {"letter": "D", "text": "Ask what specifically hasn't worked in previous attempts before deciding whether to escalate or resolve automatically", "correct": false, "explanation": "Interrogating a frustrated customer about past failures is the opposite of what they asked for."}], "correct": "A", "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "group": "C"}, {"id": "m39", "domain": 1, "scenario": null, "situation": null, "question": "When the agent calls `lookup_order` and receives order details showing the item was purchased 45 days ago, how does the agentic loop determine whether to call `process_refund` or `escalate_to_human` next?", "options": [{"letter": "A", "text": "The orchestration layer automatically routes to the next tool based on the order's status field.", "correct": false, "explanation": "There's no implicit orchestration layer picking tools from a field. The model drives tool selection."}, {"letter": "B", "text": "The agent follows a pre-configured decision tree mapping order attributes to specific tool calls.", "correct": false, "explanation": "Agentic loops are model-driven, not decision-tree-driven. Hardcoded trees are the opposite of what the agent pattern is for."}, {"letter": "C", "text": "The order details are added to the conversation and the model reasons about which action to take.", "correct": true, "explanation": "Correct. The agentic loop works by appending `tool_result` messages to the conversation and letting the model decide the next step on each turn. That's how 45 days → refund vs. escalate gets resolved."}, {"letter": "D", "text": "The agent executes the remaining steps in a tool sequence planned at the start of the request.", "correct": false, "explanation": "There's no upfront, committed plan — the agent chooses each next step based on the latest context."}], "correct": "C", "task_id": "1.1", "objective": "Design and implement agentic loops for autonomous task execution", "group": "A"}, {"id": "m40", "domain": 1, "scenario": null, "situation": "A customer sends: \"This is frustrating. I've explained my issue twice and nothing is being resolved. I want to talk to a real person NOW.\" The agent has not yet called any tools to investigate their account.", "question": "What should the agent do?", "options": [{"letter": "A", "text": "Acknowledge the frustration and ask one targeted question to understand the specific issue before escalating.", "correct": true, "explanation": "Correct. The customer has said 'twice' but you have no context yet. One acknowledging, focused question gives you a shot at first-contact resolution without dismissing the frustration or delaying a potential handoff."}, {"letter": "B", "text": "Briefly explain what the agent can help with and offer to resolve the issue quickly, escalating only if the customer repeats their request.", "correct": false, "explanation": "Launching into capability-listing at a frustrated customer who asked for a human reads as dismissive."}, {"letter": "C", "text": "Immediately call `escalate_to_human` with the conversation history.", "correct": false, "explanation": "Escalating with zero tool context creates a cold handoff where the human also starts from scratch."}, {"letter": "D", "text": "First call `get_customer` and `lookup_order` to gather account context, then escalate to a human agent.", "correct": false, "explanation": "Investigating without asking adds latency and doesn't respect the customer's request. A single acknowledging question is faster and more respectful."}], "correct": "A", "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "group": "C"}, {"id": "m41", "domain": 1, "scenario": null, "situation": "Your agent is handling a billing dispute. After calling `get_customer` and `lookup_order`, it identifies that the dispute involves a promotional pricing error requiring manager approval—beyond the agent's authorization level.", "question": "How should the workflow handle this mid-process escalation?", "options": [{"letter": "A", "text": "Call `escalate_to_human` passing only the customer's original message.", "correct": false, "explanation": "Drops the tool-derived context the agent just gathered, forcing the human to re-do the investigation."}, {"letter": "B", "text": "Compile a structured handoff with customer details, order info, and the identified issue before calling `escalate_to_human`.", "correct": true, "explanation": "Correct. A structured brief (who, what order, what issue, why it exceeds auth) lets the human agent pick up instantly. That's the mid-process escalation pattern."}, {"letter": "C", "text": "Attempt the refund with `process_refund` anyway, escalating only if the system rejects the transaction.", "correct": false, "explanation": "Knowingly exceeding authorization is a policy violation — not something to try and hope the system catches."}, {"letter": "D", "text": "Persist the complete conversation and tool response history to a database, then call `escalate_to_human` with a reference ID.", "correct": false, "explanation": "Adds infrastructure and an extra lookup step when an inline structured brief is simpler and faster."}], "correct": "B", "task_id": "1.4", "objective": "Implement multi-step workflows with enforcement and handoff patterns", "group": "A"}, {"id": "f1-001", "domain": 1, "task_id": "1.5", "objective": "Apply Agent SDK hooks for tool call interception and data normalization", "situation": "An architect is consolidating tool-output normalization into a single PostToolUse hook that must work for both MCP tools and built-in tools. Currently, a legacy MCP tool uses updatedMCPToolOutput and a newer built-in tool uses updatedToolOutput.", "question": "Which field should the shared hook use to replace the output for both tool types?", "options": [{"letter": "A", "text": "systemMessage, because it can be used to display the normalized output to the user for all tools.", "correct": false, "explanation": "systemMessage is used to set system-level instructions and is not a mechanism for replacing tool output within a PostToolUse hook. The correct approach for output replacement is using updatedToolOutput or updatedMCPToolOutput."}, {"letter": "B", "text": "additionalContext, because it provides a way to add extra information that overrides the original tool output.", "correct": false, "explanation": "additionalContext merely appends extra information to the existing tool output without replacing it. To fully override and normalize the output, you must use updatedToolOutput (or updatedMCPToolOutput for MCP-specific hooks), which completely replaces the original result."}, {"letter": "C", "text": "updatedMCPToolOutput, because it is the field intended for cross‑tool output replacement.", "correct": false, "explanation": "updatedMCPToolOutput is specifically designed to replace output only for MCP (Model Context Protocol) tools. It does not affect built-in tools like WebFetch, Bash, or Read, making it unsuitable for a shared hook that must cover both MCP and built-in tools."}, {"letter": "D", "text": "updatedToolOutput, because it replaces output for all tools (built‑in and MCP) in the PostToolUse hook, while updatedMCPToolOutput only works for MCP tools.", "correct": true, "explanation": "According to Anthropic's Agent SDK documentation, updatedToolOutput is the recommended field for true data normalization within PostToolUse hooks. It completely replaces the original tool result for both MCP and built-in tools (e.g., Bash, WebFetch, Read), ensuring a consistent, canonical shape. This approach saves tokens, reduces errors, and is key for unifying output normalization across all tool types."}], "correct": "D", "select": 1, "group": "E"}, {"id": "f1-002", "domain": 1, "task_id": "1.6", "objective": "Design task decomposition strategies for complex workflows", "situation": "During a security audit, an agent is asked to determine whether a reported vulnerability in one library is exploitable anywhere in a large application. At the outset, the architect does not yet know which call sites are affected, whether those call sites are reachable, or what mitigations, if any, already exist in the codebase.", "question": "Which task decomposition pattern should the architect choose, and why?", "options": [{"letter": "A", "text": "A single-pass prompt, because the model can always determine exploitability directly from the vulnerability's CVE description without ever inspecting the codebase", "correct": false, "explanation": "Determining exploitability requires inspecting the actual codebase for reachable call sites and existing mitigations; the CVE description alone does not establish whether this specific application is affected."}, {"letter": "B", "text": "Prompt chaining, because vulnerability audits always follow the same three fixed steps of scan, patch, and verify regardless of the application involved", "correct": false, "explanation": "Assuming a fixed scan-patch-verify sequence works for every application ignores that the number and nature of investigation steps vary by codebase and cannot be predetermined."}, {"letter": "C", "text": "Prompt chaining, because breaking the audit into a fixed number of stages guarantees that every call site will be found before the process finally ends", "correct": false, "explanation": "A fixed number of stages does not guarantee completeness for an open-ended search problem, since the number of call sites needing investigation is not known ahead of time."}, {"letter": "D", "text": "Orchestrator-workers, because the required investigation steps cannot be predicted upfront and must be generated from what each search for call sites reveals", "correct": true, "explanation": "Since the scope of call sites, reachability, and mitigations is unknown in advance and depends on what is discovered during the search, a dynamically decomposed orchestrator-workers approach that adapts its subtasks is appropriate."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-003", "domain": 1, "task_id": "1.2", "objective": "Orchestrate multi-agent systems with coordinator-subagent patterns", "situation": "A coordinator delegates a task to a 'documentation-reviewer' subagent that should only read and comment on files, never modify them. During testing, the subagent unexpectedly edits a file it was reviewing.", "question": "What configuration change prevents this?", "options": [{"letter": "A", "text": "Move the file-editing logic into the coordinator so the subagent never needs to call any tools at all", "correct": false, "explanation": "The subagent still needs Read access to review files, so removing all tools would prevent it from doing its job rather than just preventing edits."}, {"letter": "B", "text": "Add a system prompt instruction telling the subagent not to modify files, without changing its tool access", "correct": false, "explanation": "A prompt instruction alone is a soft constraint; if Edit and Write remain in the subagent's tool set, it can still invoke them despite the instruction."}, {"letter": "C", "text": "Lower the subagent's model tier so it is less capable of generating file-modification tool calls", "correct": false, "explanation": "A lower-tier model can still call Edit or Write if those tools remain available; capability level doesn't enforce a hard restriction on which tools can be invoked."}, {"letter": "D", "text": "Restrict the subagent's tools field to read-only tools like Read and Grep, omitting Edit and Write", "correct": true, "explanation": "Restricting the tools field to read-only tools such as Read and Grep, and omitting Edit and Write, enforces at the configuration level that the subagent cannot modify files."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-004", "domain": 1, "task_id": "1.5", "objective": "Apply Agent SDK hooks for tool call interception and data normalization", "situation": "An architect's agent calls three MCP tools that each return timestamps in a different format: Unix epoch integers, ISO 8601 strings, and numeric status codes mixed with dates. The model frequently misreads these inconsistent formats when reasoning about order history.", "question": "Which hook design correctly normalizes the data before the model ever sees it?", "options": [{"letter": "A", "text": "Register a PreToolUse hook matched to the three MCP tools that intercepts the call, rewrites any timestamp field in tool_input to ISO 8601 format, ensuring each MCP server receives a pre-normalized request payload.", "correct": false, "explanation": "PreToolUse hooks intercept the tool call before execution, allowing modification of the tool's input, not its output. Since the issue is inconsistent timestamp formats in the responses from the MCP tools, normalizing the input does not address the output inconsistency. The correct approach is to process the tool’s response after execution, which is what PostToolUse hooks handle."}, {"letter": "B", "text": "Register a PostToolUse hook matched to the three MCP tools that parses each tool's response and returns hookSpecificOutput.updatedToolOutput with values rewritten into one consistent format.", "correct": true, "explanation": "PostToolUse hooks are an official feature in Claude Code designed to run immediately after a tool completes successfully, and they can be configured to match specific tools via matchers. They can parse the tool's response and return updatedToolOutput, which replaces the original output with a normalized version before the model processes it. This directly addresses the requirement to normalize timestamps before the model sees the data. This is the recommended approach per Anthropic's documentation for post-processing and quality control."}, {"letter": "C", "text": "Register a Notification hook for the three MCP tools that parses each tool invocation's returned message field for date-like substrings, converts them into ISO 8601 format, and republishes a normalized summary to the transcript.", "correct": false, "explanation": "Notification hooks are not documented for intercepting and modifying tool outputs; they are intended for event-driven notifications on tool events. The option describes parsing the returned message field and republishing a normalized summary to the transcript, but this does not prevent the model from seeing the original inconsistent output. Thus, this approach fails to normalize the data before the model reads it."}, {"letter": "D", "text": "Add a system prompt instruction directing Claude to convert all timestamps and status codes from the three MCP tools into a single consistent ISO 8601 format before reasoning about order history.", "correct": false, "explanation": "A system prompt instruction relies on the model’s own processing, but Claude does not natively or reliably normalize timestamps automatically; it operates in a 'time vacuum' without explicit temporal context. Inconsistent timestamp formats may still cause misreads, and the instruction is not a hook that ensures deterministic normalization. The research confirms that explicit provision of timestamps is necessary, not reliance on the model's interpretation."}], "correct": "B", "select": 1, "group": "E"}, {"id": "f1-005", "domain": 1, "task_id": "1.2", "objective": "Orchestrate multi-agent systems with coordinator-subagent patterns", "situation": "A platform team is building a research assistant using Claude Agent SDK subagents. During code review, they notice that two subagents pass results directly to each other through a shared file that the coordinator never inspects, and when one subagent fails, the coordinator has no visibility into what happened.", "question": "Which architectural change should the team make to align with the hub-and-spoke coordinator pattern?", "options": [{"letter": "A", "text": "Merge the two subagents into one combined subagent that also owns the coordinator's error-handling responsibilities", "correct": false, "explanation": "Merging the subagents removes the specialization benefit of separate subagents and doesn't address the underlying issue of bypassing the coordinator for communication."}, {"letter": "B", "text": "Configure the subagents to poll a shared task queue directly and notify the coordinator only after completion", "correct": false, "explanation": "Polling a shared queue directly still routes communication outside the coordinator during execution, only surfacing results after the fact rather than giving the coordinator real-time control."}, {"letter": "C", "text": "Give each subagent direct write access to a shared database so they can exchange results outside the coordinator's view", "correct": false, "explanation": "A shared database still lets subagents exchange results without the coordinator observing or mediating the exchange, so failures remain invisible to the hub."}, {"letter": "D", "text": "Route all inter-subagent communication through the coordinator so it can log outcomes and handle failures consistently", "correct": true, "explanation": "Hub-and-spoke requires the coordinator to mediate all inter-subagent communication and information routing, giving it visibility to log outcomes and handle errors centrally."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-006", "domain": 1, "task_id": "1.3", "objective": "Configure subagent invocation, context passing, and spawning", "situation": "A coordinator needs to run \"style-checker,\" \"security-scanner,\" and \"test-coverage\" subagents against the same pull request, and the team wants all three to run concurrently rather than one after another. The coordinator currently emits one Task call, waits for the result, then emits the next Task call in a following turn.", "question": "What change achieves true parallel execution?", "options": [{"letter": "A", "text": "Emit all three Task calls for the checker, scanner, and coverage subagents within one coordinator response instead of spreading them across separate turns", "correct": true, "explanation": "Subagents run concurrently when their invocation calls are emitted together in a single coordinator response; issuing them across separate turns forces the coordinator to wait for one result before starting the next."}, {"letter": "B", "text": "Merge style-checker, security-scanner, and test-coverage into a single AgentDefinition so that one Task call now covers all three concerns at once", "correct": false, "explanation": "Merging distinct concerns into one subagent removes the ability to reason about them independently and does not itself produce concurrent execution; it just changes the task decomposition."}, {"letter": "C", "text": "Set persistSession to false on each subagent call so none of them block on writing a transcript to disk before the next one in line can start", "correct": false, "explanation": "persistSession controls whether a session is written to disk versus kept in memory; it has no effect on whether multiple subagent calls execute concurrently."}, {"letter": "D", "text": "Increase maxTurns on each of the three subagent definitions so they can each finish faster and thereby appear to overlap more in wall-clock time", "correct": false, "explanation": "maxTurns caps how many turns a subagent can take internally; it does not change whether multiple subagents are launched at the same time."}], "correct": "A", "select": 1, "group": "B"}, {"id": "f1-007", "domain": 1, "task_id": "1.6", "objective": "Design task decomposition strategies for complex workflows", "situation": "A content pipeline generates a product description, then always runs it through the same brand-tone checker, then always runs it through the same profanity filter, regardless of what the description says. A designer proposes replacing this with an orchestrator that dynamically decides, based on the description's content, whether to run the tone checker or the profanity filter first.", "question": "What is the strongest critique of this proposal?", "options": [{"letter": "A", "text": "Dynamic orchestration should be adopted anyway, since it always produces better brand-tone results than any fixed check order can ever achieve.", "correct": false, "explanation": "This is overstated and unsupported by the research. Dynamic orchestration is useful for complex tasks where subtasks cannot be predicted in advance, but it does not automatically produce better brand-tone results. For a fixed, predictable sequence, prompt chaining can improve accuracy by allowing Claude to focus on one subtask at a time, and a dynamic orchestrator would add unnecessary complexity."}, {"letter": "B", "text": "Because both checks always run in the same fixed order regardless of content, a prompt chain already handles this predictably; introducing a dynamic orchestrator adds unnecessary complexity and latency without improving the outcome.", "correct": true, "explanation": "Anthropic's guidance recommends prompt chaining for tasks that can be cleanly decomposed into fixed, sequential subtasks. In this pipeline, the brand-tone check and profanity filter always run in a predetermined order, so the sequence is predictable and does not require runtime decision-making. Replacing that with an orchestrator-workers pattern would add orchestration overhead and latency without meaningful benefit, since the subtasks are not dynamic or path-dependent."}, {"letter": "C", "text": "The profanity filter cannot function correctly unless it always runs before the tone checker in every possible pipeline design imaginable.", "correct": false, "explanation": "There is no documentation-backed universal rule requiring the profanity filter to run before the tone checker in every pipeline. The current pipeline runs the tone checker first and then the profanity filter, and a fixed order can be perfectly valid. This option invents an absolute dependency that is not supported by Anthropic's guidance."}, {"letter": "D", "text": "A dynamic orchestrator is necessary because product descriptions vary in length, something a fixed chain is fundamentally unable to accommodate.", "correct": false, "explanation": "Prompt chaining passes outputs from one step to the next and is not limited by variable input length. Variation in description length does not make the subtasks unpredictable or require dynamic orchestration. Dynamic orchestration is justified when the subtasks themselves cannot be determined in advance, not simply because input text length varies."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f1-008", "domain": 1, "task_id": "1.1", "objective": "Design and implement agentic loops for autonomous task execution", "situation": "A single response from Claude contains two tool_use blocks in the same turn: one requesting a currency-conversion tool and one requesting a tax-lookup tool, both needed to finish a pricing calculation.", "question": "How should the loop handle this before sending the next request?", "options": [{"letter": "A", "text": "Merge both tool requests into a single tool_result block with one combined tool_use_id chosen arbitrarily from the two requests", "correct": false, "explanation": "Each tool_use_id needs its own matching tool_result; combining two distinct calls under one arbitrary id breaks the pairing Claude expects."}, {"letter": "B", "text": "Execute both requested tools and append a separate tool_result block for each, matching each result to its own tool_use_id, before continuing", "correct": true, "explanation": "When a response includes multiple tool_use blocks, the loop must execute each one and return a matching tool_result for every tool_use_id before the conversation can continue coherently."}, {"letter": "C", "text": "Execute only the first tool_use block encountered and skip the second one, since the API always processes one tool request per turn", "correct": false, "explanation": "The API can return multiple tool_use blocks in a single turn when Claude decides several tools are needed; skipping one leaves an unanswered request."}, {"letter": "D", "text": "Send two entirely separate follow-up conversations to Claude, one addressing each tool_use block in isolation from the other", "correct": false, "explanation": "Splitting the two related tool_use blocks into separate conversations discards the shared context needed to resolve the original pricing calculation."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f1-009", "domain": 1, "task_id": "1.5", "objective": "Apply Agent SDK hooks for tool call interception and data normalization", "situation": "An Agent SDK application uses a PreToolUse hook to enforce a hard policy: an issue_refund call above $500 must never execute, and no approval prompt may be shown at call time. The hook returns permissionDecision: 'deny' with the generic permissionDecisionReason \"Refund denied\", and the agent responds by retrying the identical issue_refund call instead of using the available escalate_refund tool.", "question": "Which change best stops the retry loop while keeping the hook the deterministic gate?", "options": [{"letter": "A", "text": "Keep permissionDecision: 'deny' and rewrite permissionDecisionReason to state the $500 limit and name escalate_refund as the next step for the agent.", "correct": true, "explanation": "Anthropic's Agent SDK hooks documentation pairs the two fields exactly this way: permissionDecision: 'deny' stops the tool call, and permissionDecisionReason \"tells the model why, so it avoids retrying.\" The loop in this scenario comes from a reason that carries no usable information — a bare \"Refund denied\" leaves the model nothing to act on, so it repeats the identical call. Stating the $500 limit and naming escalate_refund keeps the hook itself the deterministic gate (the call is still blocked every time, with no prompt at call time) while giving the model a documented alternative to retrying."}, {"letter": "B", "text": "Move the amount check to a PostToolUse hook that lets the refund execute and then issues a compensating reversal transaction whenever the amount exceeded $500.", "correct": false, "explanation": "PostToolUse fires after the tool has already run, so the refund would be paid out before any amount check happened, and a compensating reversal is a second irreversible money movement that can itself fail or be delayed. A limit that must never be exceeded has to be enforced before execution, which is what PreToolUse exists for."}, {"letter": "C", "text": "Return permissionDecision: 'ask' so that a reviewer approves or rejects each refund above $500 interactively at the moment the agent makes the tool call.", "correct": false, "explanation": "Anthropic's Agent SDK hooks documentation describes permissionDecision: 'ask' as showing the call to the user for approval, which moves the final decision out of the hook. A reviewer could then approve a refund above the $500 limit, and the scenario requires that those refunds never execute and that no human prompt appears at call time. It may end the loop, but it trades deterministic enforcement for per-call human judgment and permission fatigue."}, {"letter": "D", "text": "Return permissionDecision: 'allow' together with an updatedInput that swaps the issue_refund call for a harmless no-op call so the tool call reports success.", "correct": false, "explanation": "Anthropic's Agent SDK hooks documentation shows updatedInput doing exactly one thing: rewriting the arguments of the tool Claude already called — its worked example intercepts Write calls and rewrites the file_path argument to prepend a sandbox directory. No hook field substitutes one tool for another, so swapping issue_refund for a different call cannot be built. Even if it could, reporting success for a refund that never happened misleads the model into telling the customer the money was returned, and hides the blocked action from the user and the audit trail."}], "correct": "A", "select": 1, "group": "E"}, {"id": "f1-010", "domain": 1, "task_id": "1.4", "objective": "Implement multi-step workflows with enforcement and handoff patterns", "situation": "A team wants to enforce that get_customer must run before process_refund, and registers two separate PreToolUse hooks: one matched to get_customer that writes a \"verified\" marker to a session file, and one matched to process_refund that reads that same file. A reviewer worries this design assumes the hooks run in a guaranteed order relative to each other.", "question": "Is that assumption safe, and why?", "options": [{"letter": "A", "text": "It is safe, because the runtime automatically orders hooks alphabetically by their tool name before executing them, which guarantees get_customer's hook always runs first.", "correct": false, "explanation": "There is no alphabetical ordering of hook execution based on tool names. Hook fire order is determined solely by the sequence in which the matched tools are invoked by the model, not by any naming convention. This claim is not supported by Anthropic's documentation."}, {"letter": "B", "text": "It is unsafe, because each hook runs only when its matched tool is invoked, but nothing guarantees that get_customer is invoked before process_refund. The model could skip the prerequisite tool entirely, causing the process_refund hook to read a file that may not exist or contain valid data.", "correct": true, "explanation": "The hook mechanism only ensures that a hook fires when its tool is called; it does not enforce that get_customer is called first. Without a programmatic gate, the model might directly call process_refund, leading to a missing or stale file. Anthropic's certification guidance emphasizes implementing programmatic prerequisites or hooks that block downstream calls until prerequisite steps complete, rather than relying on prompt engineering or implicit ordering."}, {"letter": "C", "text": "It is unsafe, because hooks matched to different tools share no session state with each other at all, so the process_refund hook can never see a file written by the get_customer hook.", "correct": false, "explanation": "Hooks run with the user's full permissions and can read/write the filesystem. They can share state through files, environment variables, or other IPC mechanisms. The real issue is not the lack of shared state but the absence of a guarantee that get_customer will execute before process_refund."}, {"letter": "D", "text": "It is unsafe, because every registered PreToolUse hook always executes in parallel for every tool call in the session regardless of its matcher, so the file could be read before it is ever written.", "correct": false, "explanation": "Hooks are not executed in parallel for every tool call; they only run when their matchers (e.g., a regex on tool_name) match the invoked tool. This behavior is documented in Anthropic Claude Code hook lifecycle rules, where hooks are triggered precisely by matching events, not indiscriminately."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f1-011", "domain": 1, "task_id": "1.5", "objective": "Apply Agent SDK hooks for tool call interception and data normalization", "situation": "A team wants to send every tool call's arguments to an external audit service without slowing down the agent's response time, and this audit step has no bearing on whether the call is allowed to proceed.", "question": "Which hook output pattern fits this requirement?", "options": [{"letter": "A", "text": "A PreToolUse hook that fires the audit request and returns {\"async\": true, \"asyncTimeout\": 30000} so the agent proceeds without waiting for the request to finish", "correct": true, "explanation": "Async output is meant precisely for side effects like audit logging that don't need to influence the call's outcome; the agent proceeds immediately while the background request completes, avoiding added latency."}, {"letter": "B", "text": "A PostToolUse hook that returns permissionDecision \"ask\" so the user is prompted to confirm the audit request was sent before the next tool call runs", "correct": false, "explanation": "permissionDecision fields belong to hookSpecificOutput for permission-gating hooks like PreToolUse; PostToolUse doesn't gate execution this way, and prompting the user for an audit confirmation adds unwanted friction and latency."}, {"letter": "C", "text": "A PreToolUse hook that returns permissionDecision \"allow\" together with updatedInput containing the audit payload appended to the original arguments", "correct": false, "explanation": "Appending an audit payload to updatedInput would alter the actual arguments sent to the tool, corrupting the real operation instead of simply logging it out-of-band."}, {"letter": "D", "text": "A PreToolUse hook that returns permissionDecision \"defer\" so the session pauses until the audit service confirms receipt, then resumes automatically", "correct": false, "explanation": "\"defer\" ends the query entirely so it can be resumed later; it introduces a hard pause rather than letting the agent continue immediately, which is the opposite of what a non-blocking audit step needs."}], "correct": "A", "select": 1, "group": "E"}, {"id": "f1-012", "domain": 1, "task_id": "1.6", "objective": "Design task decomposition strategies for complex workflows", "situation": "A team built a fixed five-step prompt chain to migrate a database schema: extract schema, generate migration script, validate syntax, apply migration, and confirm row counts. During testing, some migrations require an unplanned sixth step to backfill a newly discovered column that has no default value, which the chain cannot accommodate.", "question": "What is the fundamental flaw in the team's approach?", "options": [{"letter": "A", "text": "The fixed chain was correct, and the team should add a sixth hardcoded backfill step to the pipeline for every future migration regardless of schema.", "correct": false, "explanation": "Making the backfill step mandatory for all migrations would be wasteful and error-prone for migrations that don't need it. It does not address the inflexibility of the chain; unexpected requirements beyond backfilling would still break the pipeline."}, {"letter": "B", "text": "The team should not execute a migration without having a deterministic, complete understanding of the database schema beforehand; the design flaw is relying on runtime schema extraction that can miss columns.", "correct": false, "explanation": "While having a complete schema understanding is a best practice, the described failure occurs because the fixed pipeline cannot incorporate an additional backfill step when a new column is discovered. The root cause is not the schema extraction method but the rigidity of the chain. Even with an exhaustive schema upfront, migrations may still require dynamic adjustments, which a fixed chain cannot handle."}, {"letter": "C", "text": "The fixed chain only ever lacked a programmatic checkpoint right after the validate-syntax step, and adding that single checkpoint would resolve the backfill issue.", "correct": false, "explanation": "Adding a checkpoint would merely pause the pipeline, not generate the missing backfill step. The backfill logic would still need to be introduced, which a fixed chain cannot do. The issue is the chain's inability to expand its own subtasks, not the absence of a validation gate."}, {"letter": "D", "text": "The task has variable structure depending on discovered schema details, so it should use adaptive decomposition that adds backfill subtasks when they are found.", "correct": true, "explanation": "The fundamental flaw is that the team used a fixed chain for a task with variable structure. Migration workflows often require unplanned steps (like backfills) based on actual schema details. Using adaptive decomposition—such as dynamic tool calling or plan-and-execute patterns recommended in Anthropic's agent design documentation—allows the chain to insert necessary subtasks at runtime, overcoming the limitation of a hardcoded sequence."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-013", "domain": 1, "task_id": "1.2", "objective": "Orchestrate multi-agent systems with coordinator-subagent patterns", "situation": "A coordinator needs three independent code-quality checks (style, security, test coverage) performed on a pull request before it can synthesize a review. Currently the coordinator invokes each subagent one after another, and the review takes the sum of all three durations.", "question": "How should the coordinator change its invocation strategy?", "options": [{"letter": "A", "text": "Increase the maxTurns setting on each subagent so that each one finishes its individual check faster", "correct": false, "explanation": "maxTurns caps how many agentic turns a subagent can take before stopping; it doesn't make an individual check run faster or address the sequential invocation pattern."}, {"letter": "B", "text": "Invoke only the security subagent first, then decide whether to skip the other two remaining checks based on that outcome", "correct": false, "explanation": "The three checks are independent and all required for the review, so skipping style or test-coverage checks based on the security result would leave the review incomplete."}, {"letter": "C", "text": "Merge all three checks into the coordinator's own logic so no subagents are needed for this review", "correct": false, "explanation": "Folding the checks into the coordinator removes the specialization and tool-restriction benefits of dedicated subagents and doesn't address the sequential-latency problem directly."}, {"letter": "D", "text": "Invoke the style, security, and test-coverage subagents concurrently so the review finishes in the time of the slowest", "correct": true, "explanation": "Since the three checks are independent, invoking them concurrently lets the review finish in the time of the slowest check rather than the sum of all three durations."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-014", "domain": 1, "task_id": "1.2", "objective": "Orchestrate multi-agent systems with coordinator-subagent patterns", "situation": "An architect reviews a multi-agent research assistant and finds that each subagent independently implements its own retry logic, logging format, and rate-limit backoff, leading to inconsistent behavior across the system.", "question": "Which change best aligns this design with the hub-and-spoke coordinator pattern's intended benefits?", "options": [{"letter": "A", "text": "Centralize error handling, logging, and retry policy in the coordinator so all subagent communication stays consistent", "correct": true, "explanation": "The coordinator manages all inter-subagent communication and error handling in the hub-and-spoke pattern, so centralizing retry, logging, and backoff policy there produces consistent, observable behavior across the system."}, {"letter": "B", "text": "Assign one subagent the role of monitoring the others' retries and logs so the coordinator need not track failures", "correct": false, "explanation": "Delegating monitoring to another subagent adds an extra hop and still requires the coordinator to route information, rather than handling error management where communication naturally converges."}, {"letter": "C", "text": "Remove logging and retry logic from subagents entirely, since coordinator-based systems should never retry failed subtasks", "correct": false, "explanation": "Retries are still a valid and often necessary response to transient failures; the issue is inconsistency across subagents, not that retrying should never happen."}, {"letter": "D", "text": "Standardize the retry logic across subagents by copying the identical code into each subagent's own system prompt", "correct": false, "explanation": "Duplicating the same retry code into every subagent's prompt still leaves the logic scattered across subagents rather than centralized, and prompts are an unreliable place to enforce consistent behavior."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f1-015", "domain": 1, "task_id": "1.5", "objective": "Apply Agent SDK hooks for tool call interception and data normalization", "situation": "", "question": "A developer writes a PreToolUse hook that intercepts calls to a destructive shell command and returns the following JSON: { \"hookSpecificOutput\": { \"hookEventName\": \"PreToolUse\", \"permissionDecision\": \"deny\", \"permissionDecisionReason\": \"Use the sandboxed low-privilege wrapper script instead\" } } What is the runtime effect?", "options": [{"letter": "A", "text": "The tool call is allowed, and the permissionDecisionReason is logged as a warning.", "correct": false, "explanation": "A permissionDecision of \"deny\" explicitly cancels the tool call; it does not allow execution with a warning. The reason is not merely logged, but is fed back to Claude to explain the denial."}, {"letter": "B", "text": "The user is shown an interactive approval prompt asking to allow or deny the tool call.", "correct": false, "explanation": "An explicit permissionDecision of \"deny\" skips the interactive prompt and cancels the call directly. The native permission prompt would appear only if no explicit decision were returned, or if the decision were \"ask\"."}, {"letter": "C", "text": "The session terminates with an error because permissionDecisionReason is only valid with permissionDecision: \"ask\".", "correct": false, "explanation": "permissionDecisionReason is valid with \"deny\" and \"ask\", and returning a deny decision does not cause a session error. The hook response is processed normally and the tool call is cancelled."}, {"letter": "D", "text": "The tool call is cancelled, and the reason is provided to Claude to inform subsequent actions.", "correct": true, "explanation": "According to the Claude Code hooks reference, a PreToolUse hook can return permissionDecision: \"deny\" to cancel the tool call. The permissionDecisionReason is shown to Claude and included in the transcript, allowing the model to adjust its approach. This is a documented mechanism for blocking dangerous commands and enforcing low-privilege alternatives."}], "correct": "D", "select": 1, "group": "E"}, {"id": "f1-016", "domain": 1, "task_id": "1.7", "objective": "Manage session state, resumption, and forking", "situation": "While a session is resumed via --continue --fork-session to try a riskier approach, the architect notices that a permission the original session had approved with 'allow for this session' is being re-prompted in the new branch.", "question": "Why does this happen?", "options": [{"letter": "A", "text": "Permission approvals expire automatically after a fixed number of turns, independent of forking", "correct": false, "explanation": "permission approvals are not turn-limited; they are simply scoped to the session in which they were granted and do not propagate to new branches."}, {"letter": "B", "text": "The --fork-session flag was combined incorrectly with --continue and should have been used with --resume instead", "correct": false, "explanation": "--fork-session combines validly with either --continue or --resume; the flag pairing is not the cause of the re-prompt."}, {"letter": "C", "text": "Forking always resets the working directory, which invalidates any previously granted tool permissions", "correct": false, "explanation": "forking branches conversation history, not the working directory, and the directory itself is unaffected by the fork."}, {"letter": "D", "text": "Session-scoped permission approvals do not carry over from the original session into a newly forked branch", "correct": true, "explanation": "Documented forking behavior states that permissions approved with 'allow for this session' do not carry over to the new branch, so the branch re-prompts even though the conversation history is copied."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-017", "domain": 1, "task_id": "1.2", "objective": "Orchestrate multi-agent systems with coordinator-subagent patterns", "situation": "A coordinator delegates a literature-review task to three research subagents, giving each the identical instruction 'research recent advances in transformer architectures.' The subagents return heavily overlapping summaries covering the same handful of well-known papers, wasting effort and leaving other subtopics unexplored.", "question": "What should the coordinator do differently?", "options": [{"letter": "A", "text": "Give each subagent the same search tool but a different model, such as one on haiku and two on sonnet", "correct": false, "explanation": "Changing the underlying model for each subagent doesn't change what each is instructed to research, so all three would still target the same overlapping subject matter."}, {"letter": "B", "text": "Instruct all three subagents to run in parallel instead of sequentially so their results arrive together", "correct": false, "explanation": "Parallel execution addresses latency, not overlapping scope; the subagents would still research the same identical instruction concurrently instead of sequentially."}, {"letter": "C", "text": "Reduce the total number of research subagents from three down to just one so there is no possibility of any overlap", "correct": false, "explanation": "Dropping to one subagent removes the possibility of parallel coverage entirely and loses the throughput benefit of using multiple subagents to cover more ground."}, {"letter": "D", "text": "Assign each subagent a distinct subtopic or source type, such as papers, industry blog posts, and open-source code", "correct": true, "explanation": "Partitioning research scope across subagents by assigning distinct subtopics or source types minimizes duplication and expands overall coverage."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-018", "domain": 1, "task_id": "1.2", "objective": "Orchestrate multi-agent systems with coordinator-subagent patterns", "situation": "A support-ticket triage system always invokes a full pipeline of five subagents (classifier, sentiment analyzer, knowledge-base search, summarizer, and escalation checker) for every incoming ticket, including simple one-line requests that only need classification. Response times have become unacceptable.", "question": "How should the coordinator be redesigned?", "options": [{"letter": "A", "text": "Convert all five subagents into a single subagent that runs every step sequentially without coordinator involvement.", "correct": false, "explanation": "Collapsing all subagents into one sequential process removes the coordinator but still runs every step for every ticket, including steps a simple ticket does not need. This does not fix the overprocessing problem and loses the separation of concerns and specialized tool/context isolation that Anthropic's multi-agent orchestration supports."}, {"letter": "B", "text": "Have the coordinator assess each ticket's complexity by running a dedicated lightweight classification prompt that returns a structured JSON object with fields like complexity (e.g., low, medium, high) and required_subagents, then dynamically invoke only the subagents listed in that output.", "correct": true, "explanation": "This directly addresses how to measure complexity on the coordinator: use a structured prompt contract that scores complexity and lists required subagents. Anthropic guidance recommends defining explicit output formats (e.g., JSON fields for severity, recommended owner, and similar) and using lightweight models like Claude Haiku for triage. The coordinator consumes this JSON and invokes only the listed subagents, avoiding the full five-subagent pipeline for simple tickets. Add few-shot examples and confidence thresholds, routing low-confidence classifications to a human review queue."}, {"letter": "C", "text": "Remove the coordinator entirely and let the classifier subagent directly invoke the remaining subagents it deems necessary.", "correct": false, "explanation": "Removing the coordinator does not follow the recommended orchestrator-subagent pattern. Anthropic documents a lead agent or coordinator as responsible for task decomposition and delegation, which keeps routing logic separate and manageable. Embedding that responsibility inside the classifier can reduce clarity and may introduce coordination challenges."}, {"letter": "D", "text": "Run all five subagents in parallel for every ticket so total latency matches the slowest subagent instead of the sum.", "correct": false, "explanation": "Parallelism can reduce wall-clock latency, but it still invokes all five subagents for every ticket, wasting compute and tokens on unnecessary steps. Anthropic's guidance emphasizes complexity-aware delegation and dynamic invocation rather than running the full pipeline in parallel for all requests."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f1-019", "domain": 1, "task_id": "1.3", "objective": "Configure subagent invocation, context passing, and spawning", "situation": "During a design review, an engineer says: \"Let's just resume the analysis session twice, once for the caching approach and once for the queueing approach, so we get two independent explorations.\" A colleague objects.", "question": "What is the correct concern with this plan, given how resume and fork differ?", "options": [{"letter": "A", "text": "Resume can only ever be called once per session id, so the second resume attempt would fail outright and the queueing exploration could never start at all", "correct": false, "explanation": "There is no such limitation; a session can be resumed multiple times. Each resume appends to the same session history rather than failing. The recommended way to create independent branches is forking, not a second resume."}, {"letter": "B", "text": "Resuming quietly discards all prior tool results before continuing, so neither the caching nor the queueing exploration would retain the original analysis", "correct": false, "explanation": "Resuming fully restores the session state, including all prior messages and tool results; it does not discard anything. Only commands like /clear or /compact intentionally reset or summarize context."}, {"letter": "C", "text": "Resume and fork behave identically in this scenario, so the colleague's objection is unfounded and either call sequence produces two independent explorations", "correct": false, "explanation": "Resume and fork are fundamentally different: resume continues the same session linearly, while fork creates a divergent copy. Using resume twice would produce serial, not independent, explorations, making the colleague's concern valid."}, {"letter": "D", "text": "Resuming the same session twice appends both explorations to one shared history in sequence, so the second exploration sees the first; forking gives two independent branches", "correct": true, "explanation": "Resuming a session (via claude --continue or claude --resume <session-id>) restores the full context and appends new interactions to the same linear history, causing later explorations to inherit earlier context. Forking (via /branch, --fork-session, or --continue --fork-session) creates a new session that starts with a copy of the original history and then diverges independently, as confirmed in Anthropic's Claude Code documentation and best practices."}], "correct": "D", "select": 1, "group": "B"}, {"id": "f1-020", "domain": 1, "task_id": "1.3", "objective": "Configure subagent invocation, context passing, and spawning", "situation": "An architect is defining a \"database-migration\" AgentDefinition with a long prompt field describing SQL best practices, rollback strategy, and data integrity checks. Separately, each time the coordinator invokes this subagent it supplies a different per-call prompt describing the specific table and migration at hand.", "question": "What is the correct relationship between these two prompt sources?", "options": [{"letter": "A", "text": "Both prompt sources are concatenated and then truncated to the shorter of the two, so only whichever prompt is more concise is guaranteed to reach the subagent intact", "correct": false, "explanation": "There is no truncation-to-the-shorter-prompt behavior; both the definition's system prompt and the invocation prompt are used in full."}, {"letter": "B", "text": "AgentDefinition.prompt is ignored at runtime once a per-call prompt is supplied, so only the invocation-time prompt actually reaches the subagent's model", "correct": false, "explanation": "AgentDefinition.prompt is not discarded; it continues to shape the subagent's behavior as its system prompt on every invocation, alongside the per-call prompt."}, {"letter": "C", "text": "The per-call prompt is merged into the coordinator's own system prompt instead of the subagent's, so the subagent only ever sees AgentDefinition.prompt", "correct": false, "explanation": "The per-call prompt is delivered to the subagent being invoked, not folded into the coordinator's own system prompt."}, {"letter": "D", "text": "AgentDefinition.prompt sets the subagent's persistent system prompt and expertise, while the per-call prompt passed at invocation supplies the specific task details for that run", "correct": true, "explanation": "AgentDefinition.prompt defines the subagent's standing role, expertise, and constraints as its system prompt, while the prompt string passed at each invocation carries the specific task and data for that particular run; both reach the subagent together."}], "correct": "D", "select": 1, "group": "B"}, {"id": "f1-021", "domain": 1, "task_id": "1.6", "objective": "Design task decomposition strategies for complex workflows", "situation": "A researcher tasks an agent with answering an open-ended question: \"What is causing the 15% increase in checkout abandonment over the last month?\" The relevant data sources, whether logs, analytics dashboards, or recent deploys, are not specified in advance and depend on what early findings suggest.", "question": "How should this task be decomposed?", "options": [{"letter": "A", "text": "Chain together a fixed sequence of exactly three steps: analytics dashboard, then server logs, then recent deploys, always in that exact same order", "correct": false, "explanation": "Open-ended investigations require dynamic decomposition because the relevant data sources and next steps depend on early findings. A fixed, three-step sequence contradicts Anthropic's recommended approach for open-ended problems where the required steps cannot be predicted in advance."}, {"letter": "B", "text": "Have the model state the cause immediately from only the phrase describing the abandonment increase, without consulting any data source", "correct": false, "explanation": "The model should investigate available evidence. Anthropic's multi-agent research system delegates subtasks to gather and evaluate data; stating a cause without consulting any data source would be speculation rather than a decomposed research process."}, {"letter": "C", "text": "Generate an initial investigation subtask, evaluate its findings, then dynamically generate follow-up subtasks that pursue the most promising leads", "correct": true, "explanation": "Anthropic's multi-agent research system is designed for open-ended problems where it is difficult to predict the required steps in advance. A lead agent decomposes the query into an initial subtask, evaluates results, and adapts by generating follow-up subtasks that pursue promising leads rather than relying on a fixed sequence."}, {"letter": "D", "text": "Split the investigation into per-file local analysis passes across the codebase, followed by a single cross-file integration pass", "correct": false, "explanation": "This is a rigid, codebase-centric decomposition. The task is an open-ended investigation that may span logs, analytics, and deploys, and the appropriate decomposition should emerge dynamically from findings rather than a fixed per-file code review pattern."}], "correct": "C", "select": 1, "group": "A"}, {"id": "f1-022", "domain": 1, "task_id": "1.1", "objective": "Design and implement agentic loops for autonomous task execution", "situation": "While implementing an agent loop, an engineer is unsure where a tool_result block belongs once a tool finishes executing.", "question": "Which placement correctly continues the conversation so Claude can incorporate the result into its next round of reasoning?", "options": [{"letter": "A", "text": "Store the tool_result in a separate system-role message positioned before the original user prompt at the start of history", "correct": false, "explanation": "Placing the result in a system message ahead of the original prompt misrepresents it as background instruction rather than a response to a specific tool call."}, {"letter": "B", "text": "Attach the tool_result as metadata on the HTTP request headers rather than as part of the messages array sent to the API", "correct": false, "explanation": "Tool results are part of the conversation content the model reads, not transport-level metadata carried in request headers."}, {"letter": "C", "text": "Add the tool_result as content inside a new user-role message appended after the assistant message that contained the matching tool_use block", "correct": true, "explanation": "Tool results are returned as content in a subsequent user-role message, appended after the assistant turn that requested the tool, so Claude can read them on the next call."}, {"letter": "D", "text": "Insert the tool_result directly inside that same assistant-role message that originally contained the matching tool_use block, replacing it fully", "correct": false, "explanation": "Editing the prior assistant message to replace the tool_use block would erase the record of what Claude actually requested, breaking the tool_use/tool_result pairing."}], "correct": "C", "select": 1, "group": "A"}, {"id": "f1-023", "domain": 1, "task_id": "1.2", "objective": "Orchestrate multi-agent systems with coordinator-subagent patterns", "situation": "A coordinator invokes a web-search subagent that hits a transient network error partway through its task.", "question": "Under the hub-and-spoke pattern, where should the responsibility for deciding whether to retry or fall back lie?", "options": [{"letter": "A", "text": "In each individual subagent alone, because only the failed subagent has enough local detail to decide final recovery and should handle retry/fallback without escalating to the coordinator.", "correct": false, "explanation": "While the subagent may retry transient errors locally a few times, it must not own the final retry/fallback decision. In the hub-and-spoke pattern, subagents communicate exclusively through the coordinator; after bounded local retries, the subagent should stop and return with a retryable flag, allowing the coordinator to decide. Giving subagents independent final recovery control can lead to inconsistent behavior and cascading failures."}, {"letter": "B", "text": "In a separate monitoring subagent invoked solely to watch for and log errors raised by all of the other subagents.", "correct": false, "explanation": "A separate monitoring-only subagent is not part of the hub-and-spoke error-handling path for retry/fallback decisions. The coordinator itself maintains observability and audit trails; adding an extra spoke to watch other spokes adds complexity without moving the decision out of the coordinator."}, {"letter": "C", "text": "In the end user's client application, since it is the only component that ever observes the final aggregated results.", "correct": false, "explanation": "The client application receives only final aggregated results and has no visibility into subagent execution details; it is not responsible for transient error recovery. Under Anthropic’s shared responsibility and hub-and-spoke guidance, the coordinator remains responsible for orchestrating subagent error recovery after local retries."}, {"letter": "D", "text": "In the coordinator, but only after the web-search subagent has performed its own bounded local retries; if the transient error persists, the subagent returns control to the coordinator with a retryable flag, and the coordinator then makes the final retry or fallback decision.", "correct": true, "explanation": "Hub-and-spoke centralizes final error-handling responsibility in the coordinator, but transient errors are best retried locally first. The web-search subagent should make a bounded number of retries (e.g., 2-3) because it has the immediate context; if the error persists, it stops and returns control to the coordinator with a retryable flag set to true, letting the orchestrator decide on further retries or fallback. This keeps recovery consistent and auditable.\n\nAnthropic’s guidance notes that web-search API requests can return HTTP 200 while embedding failures in web_search_tool_result, so the coordinator may also need content-aware handling after the local retries are exhausted."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-024", "domain": 1, "task_id": "1.6", "objective": "Design task decomposition strategies for complex workflows", "situation": "A platform team is designing an agent that performs quarterly compliance reviews of Terraform configuration files. Every quarter, the exact same checklist of eleven rules is applied to every file, and no rule's outcome changes which other rules are checked. A junior engineer suggests using a dynamic orchestrator that generates new compliance subtasks at runtime.", "question": "What should the architect recommend instead, and why?", "options": [{"letter": "A", "text": "Keep the dynamic orchestrator, since generating compliance subtasks at runtime always produces higher-quality findings than any fixed checklist ever would", "correct": false, "explanation": "Dynamic subtask generation is valuable for tasks whose structure cannot be predicted, not a universal quality improvement; here it adds complexity without benefit since the checklist never changes."}, {"letter": "B", "text": "Use a fixed prompt chain, but let the model silently skip whichever of the eleven rules it judges unnecessary for a given file", "correct": false, "explanation": "Letting the model skip rules without a checkpoint undermines the reliability of a compliance review, where every one of the eleven fixed checks should be verified rather than left to unmonitored judgment."}, {"letter": "C", "text": "Use a fixed prompt chain that walks the eleven-rule checklist in the same order every quarter, since the subtasks never depend on intermediate findings", "correct": true, "explanation": "Since the eleven rules are always the same and independent of each other's outcomes, this is a textbook fixed and predictable workflow, so a prompt chain is the recommended pattern over a dynamic orchestrator that adds unneeded runtime decision-making."}, {"letter": "D", "text": "Switch to a dynamic orchestrator, since Terraform files vary too much in size for any fixed review process to ever be applied consistently", "correct": false, "explanation": "File size variation does not change the fact that the same eleven rules apply uniformly to every file, so it does not justify abandoning a fixed, predictable chain."}], "correct": "C", "select": 1, "group": "A"}, {"id": "f1-025", "domain": 1, "task_id": "1.6", "objective": "Design task decomposition strategies for complex workflows", "situation": "A team is reviewing a 60-file pull request for a payment processing service. They plan to run one agent pass per file to flag local issues, then feed all sixty per-file summaries into a second pass that looks specifically for mismatches between how modules call each other. During implementation, the second pass keeps missing a bug where one file's function signature changed but three callers in other files were not updated.", "question": "What is the most likely root cause?", "options": [{"letter": "A", "text": "Cross-file integration passes are fundamentally unable to detect function signature mismatches under any circumstances, regardless of the inputs given", "correct": false, "explanation": "Cross-file integration passes are specifically the pattern recommended for catching signature and call-site mismatches; the failure here is a lack of the necessary input data, not a fundamental limitation of the pattern itself."}, {"letter": "B", "text": "The per-file passes should be removed entirely and replaced with one prompt simply asking whether any callers across the service are broken", "correct": false, "explanation": "Removing per-file passes eliminates the detailed local analysis needed to first identify signature changes, leaving the single remaining prompt with even less information to detect broken callers."}, {"letter": "C", "text": "The per-file summaries likely omit the exact signatures and call sites the cross-file pass needs, so per-file output should also capture those details", "correct": true, "explanation": "If the per-file pass only reports local issues without recording exported signatures and where they are called, the cross-file pass has nothing concrete to compare, so enriching per-file output with that information is the fix that lets the integration pass actually detect the mismatch."}, {"letter": "D", "text": "Sixty files is too small a set for a two-pass strategy to be worthwhile, so all sixty files should instead be reviewed in one unified combined pass", "correct": false, "explanation": "Sixty files is a large enough review that a single combined pass would risk the same attention dilution the two-pass strategy was designed to avoid; the fix is refining pass content, not abandoning the two-pass structure."}], "correct": "C", "select": 1, "group": "A"}, {"id": "f1-026", "domain": 1, "task_id": "1.4", "objective": "Implement multi-step workflows with enforcement and handoff patterns", "situation": "A billing MCP server exposes several tools: issue_refund, void_authorization, and apply_credit, all of which move money and should be gated behind the same identity-verification check. Writing three nearly identical PreToolUse hooks, one per tool name, would work but creates duplication the architect wants to avoid.", "question": "What matcher approach covers all three with one hook registration, without also matching unrelated read-only tools from the same server?", "options": [{"letter": "A", "text": "No matcher at all, relying on the assumption that hooks without a matcher only fire for the subset of tools that historically move money", "correct": false, "explanation": "Without an explicit matcher, a PreToolUse hook typically fires for every tool invocation unless constrained by the framework. There is no documented behavior in Anthropic's Agent SDK where hooks without a matcher selectively fire based on tool history or category. This would lead to the identity-verification check being triggered for all tools, causing unnecessary friction and potentially breaking read-only workflows. A precise matcher is required."}, {"letter": "B", "text": "A regex matcher such as \"mcp__billing__(issue_refund|void_authorization|apply_credit)\" naming the three money-moving tools within the billing server's namespace", "correct": true, "explanation": "According to Anthropic's Agent SDK documentation, the naming convention for tool routing is mcp__<server-name>__<tool-name>. Using a regex that lists the specific tool names (issue_refund, void_authorization, apply_credit) ensures the PreToolUse hook fires only for these money-moving operations. This approach avoids duplication, keeps the configuration DRY, and excludes unrelated read-only tools from the same server. The 64-character limit for tool names is also respected."}, {"letter": "C", "text": "A matcher of \"*\" restricted afterward by giving the hook a five-second timeout, since a shorter timeout value narrows down which tools the hook actually applies to", "correct": false, "explanation": "A matcher of \"*\" would match every tool call across all MCP servers, not just the billing server. The timeout does not filter or restrict which tools the hook applies to; it merely sets an execution deadline. Even with a short timeout, the hook would still intercept unrelated read-only tools, causing unnecessary latency and potential failures. Matching must be based on tool names, not on arbitrary constraints like timeouts."}, {"letter": "D", "text": "The matcher \"mcp__billing\" by itself, since any matcher beginning with the server's name automatically covers every single tool that server exposes", "correct": false, "explanation": "Claude Code's PreToolUse hook matcher is evaluated as an exact-string match whenever it is made up only of letters, digits, underscores, dashes, spaces, commas or pipes — a bare server-name prefix like \"mcp__billing\" falls into this category, so it is compared as the literal tool name \"mcp__billing\" and matches no tool at all, since the server's actual tools are named mcp__billing__issue_refund, mcp__billing__void_authorization, and so on. It does not, as claimed here, automatically cover every tool the server exposes — reaching every tool under a server requires appending \".*\" to the prefix (for example mcp__billing__.*), and that broader pattern would still catch read-only tools like mcp__billing__get_invoice. Either way, this matcher fails to isolate the three money-moving tools."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f1-027", "domain": 1, "task_id": "1.4", "objective": "Implement multi-step workflows with enforcement and handoff patterns", "situation": "An architect is choosing where to place enforcement for a workflow where refund_tool must never run before verify_tool succeeds. One option is a Notification hook that logs a warning message whenever refund_tool is called without prior verification.", "question": "Why does this option fail to meet the deterministic-compliance requirement?", "options": [{"letter": "A", "text": "Notification hooks are only available in the Python SDK, so a TypeScript-based agent could not use this approach even if it wanted to log a warning, as the TypeScript runtime lacks hook registration.", "correct": false, "explanation": "The issue is not about SDK availability; Notification hooks are available across languages. The design fails because this hook type cannot block a tool call, regardless of the SDK used."}, {"letter": "B", "text": "Notification hooks execute before the tool call but require a human to manually click through every notification, adding non-deterministic latency because the workflow stalls until the human responds.", "correct": false, "explanation": "Notification hooks do not require human interaction to proceed; they simply log information. The fundamental flaw is that they lack any mechanism to block the tool call, not that they introduce latency."}, {"letter": "C", "text": "Notification hooks cannot be matched to a specific tool name at all, so the warning would fire for every tool call in the session, including routine operations like search, rather than only refund_tool.", "correct": false, "explanation": "Notification hooks can be matched to specific tool names using matchers. The real problem is that they are not designed to gate tool execution; they cannot prevent a tool call from proceeding."}, {"letter": "D", "text": "Notification hooks only carry status messages and cannot set a permissionDecision that blocks the call, so the refund would already have executed before the warning is logged.", "correct": true, "explanation": "Notification hooks are designed to carry status messages, not to make permission decisions. They cannot set a permissionDecision to block the tool call, so the refund tool would have already executed before the warning is logged, failing to prevent it."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-028", "domain": 1, "task_id": "1.5", "objective": "Apply Agent SDK hooks for tool call interception and data normalization", "situation": "A finance-operations agent uses a process_payment MCP tool. An architect wants to normalize the amount field, which some upstream integrations send as a string like \"$1,250.00\" and others send as a float like 1250.0, into a single float type, and then enforce a compliance threshold on the normalized value before the tool call proceeds.", "question": "Which design best achieves this, according to Anthropic's hook system behavior?", "options": [{"letter": "A", "text": "Perform both normalization and threshold checking in a single PostToolUse hook, using the tool's output to derive the normalized amount and then compare against the threshold.", "correct": false, "explanation": "This approach only acts after the tool has already executed, meaning a non-compliant tool call would still go through. Compliance enforcement must happen before execution to block unauthorized actions. A PostToolUse hook is also ill-suited for input validation because it does not receive the original tool_input by default and would need to reconstruct the normalized value from results, which may not be possible or accurate."}, {"letter": "B", "text": "Register two PreToolUse hooks in any order, as updatedInput automatically propagates between PreToolUse hooks so the normalization hook's changes are always visible to the threshold-enforcement hook.", "correct": false, "explanation": "updatedInput does not propagate between PreToolUse hooks. Each hook evaluates the same call independently using the original tool_input. Hooks cannot see or depend on modifications made by other PreToolUse hooks, and the hook execution order is unspecified. This makes chaining unreliable."}, {"letter": "C", "text": "Use a PreToolUse hook for normalization to convert the amount and update the input, then a PostToolUse hook for threshold enforcement to verify the converted amount after the tool executes.", "correct": false, "explanation": "While the normalization could occur in PreToolUse, relying on a PostToolUse hook for threshold enforcement does not prevent the tool from executing with a non-compliant input. The enforcement should happen before the tool call, not after. Additionally, hooks do not share state, so the PostToolUse hook would need to re-derive the normalized value from the tool result, adding unnecessary complexity."}, {"letter": "D", "text": "Implement a single PreToolUse hook that normalizes the amount to a float and then checks the threshold, returning an updatedInput with the normalized value if allowed, or denying the call otherwise.", "correct": true, "explanation": "This is the recommended and reliable approach. Multiple PreToolUse hooks do not share updatedInput – each hook receives the original tool_input independently, and the execution order is not guaranteed. Consolidating the logic into one hook avoids these issues and a known bug where updatedInput is ignored when multiple PreToolUse hooks fire for the same tool call (Issue #15897). It also ensures the threshold check uses the normalized value and runs before tool execution."}], "correct": "D", "select": 1, "group": "E"}, {"id": "f1-029", "domain": 1, "task_id": "1.1", "objective": "Design and implement agentic loops for autonomous task execution", "situation": "An agent loop generating a long report hits a response where stop_reason comes back as \"max_tokens\" rather than \"tool_use\" or \"end_turn\", because the output was truncated before Claude could finish. The loop's control flow only branches on those two familiar values and falls through to the \"end_turn\" branch by default.", "question": "What is the risk of that fallback behavior?", "options": [{"letter": "A", "text": "The loop treats a truncated, incomplete response as if the task were finished, so it stops the agent before Claude has actually completed the work", "correct": true, "explanation": "Falling through to the \"end_turn\" branch conflates truncation with genuine completion, so the loop ends the task while Claude's response was actually cut off mid-generation."}, {"letter": "B", "text": "The loop automatically increases the max_tokens parameter on the very next outgoing request without any code change, resolving the truncation entirely", "correct": false, "explanation": "Nothing in the API automatically adjusts max_tokens on a later request; any such adjustment would have to be implemented explicitly by the developer."}, {"letter": "C", "text": "The loop crashes immediately with an unhandled exception, since \"max_tokens\" is not a value the Messages API is permitted to return", "correct": false, "explanation": "\"max_tokens\" is a valid, documented stop_reason value; encountering it does not itself cause an exception unless the loop's own code is broken."}, {"letter": "D", "text": "The loop discards the truncated response entirely and silently resends the very first request in the conversation from scratch", "correct": false, "explanation": "The described fallback treats the response as complete rather than discarding it and restarting; resending the original request from scratch is not what a same-value fallback to \"end_turn\" does."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f1-030", "domain": 1, "task_id": "1.4", "objective": "Implement multi-step workflows with enforcement and handoff patterns", "situation": "A refund workflow requires that process_refund always be called with the exact customer_id captured by an earlier verified get_customer call, never a value the model retypes from the conversation. A PreToolUse hook already blocks the call when no verified ID exists in session state.", "question": "What should the hook do once a verified ID is present, to prevent the model from substituting a different ID string?", "options": [{"letter": "A", "text": "Return permissionDecision \"ask\" so a human reviewer retypes the same ID manually before every refund, even when verification already succeeded", "correct": false, "explanation": "Escalating every already-verified refund to manual retyping adds friction without fixing the substitution risk, and does not correct the argument itself."}, {"letter": "B", "text": "Return an empty object so the call proceeds unchanged, since the presence of a verified ID in session state is enough evidence that the model used it", "correct": false, "explanation": "An empty object allows the call with whatever customer_id the model supplied, which is exactly the untrusted value the workflow needs to prevent from reaching the tool."}, {"letter": "C", "text": "Return permissionDecision \"defer\" so the query pauses indefinitely until an operator resumes it with the corrected customer_id argument", "correct": false, "explanation": "Defer ends the query for later resumption; it does not correct the argument and is unnecessary once the ID is already verified in session state."}, {"letter": "D", "text": "Return permissionDecision \"allow\" together with updatedInput that overwrites the tool's customer_id argument with the verified ID stored in session state", "correct": true, "explanation": "Combining \"allow\" with updatedInput lets the hook substitute the trusted, verified value for whatever the model passed, closing the gap between what was verified and what actually reaches the tool."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-031", "domain": 1, "task_id": "1.5", "objective": "Apply Agent SDK hooks for tool call interception and data normalization", "situation": "An architect registers three independent PreToolUse hooks for the same charge_card tool: one checks fraud signals, one checks the daily spending cap, and one checks account status. During a live call, the fraud-signal hook returns permissionDecision \"deny\" while the other two both return \"allow\".", "question": "What happens to the tool call?", "options": [{"letter": "A", "text": "The call proceeds, because a majority of the registered hooks returned \"allow\" and the SDK resolves conflicting decisions by simple vote", "correct": false, "explanation": "There is no majority-vote resolution; the SDK does not count \"allow\" versus \"deny\" responses, it always applies the most restrictive decision present."}, {"letter": "B", "text": "The SDK raises a configuration error and halts the session, because hooks matched to the same tool are not permitted to return conflicting decisions", "correct": false, "explanation": "Conflicting decisions across hooks are an expected and supported pattern, not an error condition; the SDK resolves them by priority rather than failing the session."}, {"letter": "C", "text": "The call proceeds, because only the first hook registered in the PreToolUse array is evaluated and the remaining two hooks are skipped entirely", "correct": false, "explanation": "All matching hooks for an event run, typically in parallel, rather than short-circuiting after the first one; the fraud, spending-cap, and account-status checks are all evaluated."}, {"letter": "D", "text": "The call is blocked, because when multiple hooks disagree the most restrictive result applies and any single \"deny\" overrides the other hooks' \"allow\" decisions", "correct": true, "explanation": "When multiple hooks apply to the same event, the most restrictive outcome wins: deny takes priority over defer, which takes priority over ask, which takes priority over allow. A single deny blocks the operation no matter how many other hooks allowed it."}], "correct": "D", "select": 1, "group": "E"}, {"id": "f1-032", "domain": 1, "task_id": "1.4", "objective": "Implement multi-step workflows with enforcement and handoff patterns", "situation": "An architect is deciding whether identity verification before a refund needs a programmatic gate or can rely on prompt instructions alone. A colleague argues the system prompt already tells Claude to verify first, so a hook is redundant engineering effort.", "question": "Why is the colleague's reasoning wrong for this specific workflow?", "options": [{"letter": "A", "text": "System prompts are truncated by the runtime after a fixed number of characters, so verification instructions placed near the end of a long prompt are silently dropped, making enforcement unreliable.", "correct": false, "explanation": "Runtime truncation of system prompts at a fixed character limit is not a standard behavior, and instructions are not silently dropped in typical LLM platforms. The real problem is the probabilistic nature of prompt following, not a fictional truncation issue."}, {"letter": "B", "text": "Refunds are financial and hard to reverse, and prompt instructions have a non-zero failure rate, so a single skipped verification becomes an unrecoverable compliance gap.", "correct": true, "explanation": "Refunds are high-stakes and hard to reverse, making any failure to verify identity a serious compliance gap. Since prompt instructions have a non-zero failure rate, only a programmatic gate can deterministically enforce the verification before the refund."}, {"letter": "C", "text": "The refund tool's schema does not expose a customer_id parameter, so the model has no field in which to store verification outcomes, preventing any automated check that verification was performed.", "correct": false, "explanation": "A missing customer_id parameter in the tool schema is a fixable design detail, not the root cause of unreliability. Even with proper schema design, prompt-only ordering remains probabilistic, so it cannot ensure verification is always performed."}, {"letter": "D", "text": "Claude cannot call two different tools within the same conversation turn, so verification and refund must be separated by a human-reviewed pause to execute each in its own turn.", "correct": false, "explanation": "Claude can call multiple tools within a conversation, and there is no requirement for a human-reviewed pause between tool calls. The colleague's reasoning is wrong not because of tool-calling limitations, but because prompt instructions alone cannot guarantee execution of compliance-critical steps."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f1-033", "domain": 1, "task_id": "1.5", "objective": "Apply Agent SDK hooks for tool call interception and data normalization", "situation": "A PreToolUse hook needs to block refunds above $500 only when they target a specific merchant category, based on a field named category inside the refund tool's arguments. The hook is registered with matcher=\"refund_customer\".", "question": "Where should the category check be implemented?", "options": [{"letter": "A", "text": "Inside the callback function itself, by reading input_data[\"tool_input\"][\"category\"] and applying the conditional logic there, since matchers only filter by tool name", "correct": true, "explanation": "Matchers only filter on the event's target field, which for tool-based hooks is the tool name, not its arguments. Any argument-level condition, like a category field, must be evaluated inside the callback by reading tool_input directly."}, {"letter": "B", "text": "In the tool's own MCP server definition, since PreToolUse hooks cannot inspect individual tool arguments under any circumstances", "correct": false, "explanation": "PreToolUse hooks do receive tool_input, including arbitrary arguments like category, inside the callback; the claim that arguments can never be inspected is incorrect."}, {"letter": "C", "text": "In a second, separate matcher field called argument_matcher that the HookMatcher accepts alongside the tool-name matcher", "correct": false, "explanation": "HookMatcher does not expose a separate argument_matcher field; the only filtering field for tool hooks is the tool-name matcher."}, {"letter": "D", "text": "In the matcher string, by writing a regex such as refund_customer.*category=restricted so the SDK filters on the argument before invoking the callback", "correct": false, "explanation": "Matchers are compared against the tool name only; they have no visibility into tool_input fields, so embedding an argument condition in the matcher string would not work as intended."}], "correct": "A", "select": 1, "group": "E"}, {"id": "f1-034", "domain": 1, "task_id": "1.3", "objective": "Configure subagent invocation, context passing, and spawning", "situation": "An architect wants a \"lead-investigator\" subagent that, once running, can itself spawn narrower sub-investigator subagents to parallelize a large audit, while a separate \"final-summarizer\" subagent must never be allowed to spawn any subagents of its own.", "question": "How should the two AgentDefinitions differ to enforce this?", "options": [{"letter": "A", "text": "Set lead-investigator's model to a larger model and final-summarizer's model to a smaller one, since only larger models are capable of nested delegation", "correct": false, "explanation": "The ability to spawn subagents is governed by tool access, not model size; a smaller model with the invocation tool available could still attempt to spawn subagents."}, {"letter": "B", "text": "Set background to true on lead-investigator only, since background execution is what grants an agent the ability to spawn further subagents", "correct": false, "explanation": "background controls whether an agent runs as a non-blocking task, not whether it has permission to invoke further subagents."}, {"letter": "C", "text": "Include the subagent-invocation tool in lead-investigator's tools array, and omit it from (or add it to disallowedTools on) final-summarizer's definition", "correct": true, "explanation": "Whether a subagent can spawn its own subagents depends on whether the invocation tool is available to it; including it in lead-investigator's tools (or leaving tools unset) allows nesting, while omitting it or listing it in disallowedTools on final-summarizer blocks that capability entirely."}, {"letter": "D", "text": "Give both agents identical tools arrays, but set a lower maxTurns on final-summarizer so it runs out of turns before it could attempt to spawn anything", "correct": false, "explanation": "A turn limit constrains how much work an agent can do before stopping, not whether the invocation tool is available to it; it is not a reliable way to prevent spawning."}], "correct": "C", "select": 1, "group": "B"}, {"id": "f1-035", "domain": 1, "task_id": "1.2", "objective": "Orchestrate multi-agent systems with coordinator-subagent patterns", "situation": "After a coordinator aggregates results from search and analysis subagents into a synthesis report on 'emerging risks in supply chain cybersecurity,' the report thoroughly covers ransomware but omits any discussion of third-party vendor risk, which was part of the original scope.", "question": "What should the coordinator do next?", "options": [{"letter": "A", "text": "Publish the report as final, since the coordinator already invoked every available subagent once already", "correct": false, "explanation": "Publishing without addressing the identified gap leaves the report incomplete relative to its original scope, which is exactly the failure mode iterative refinement is meant to catch."}, {"letter": "B", "text": "Ask the synthesis subagent to rewrite the existing report in a different tone without gathering new source material", "correct": false, "explanation": "Rewriting the tone doesn't gather the missing information about third-party vendor risk, so the underlying coverage gap remains unresolved."}, {"letter": "C", "text": "Restart the entire pipeline from scratch with a completely new set of subagents to avoid compounding earlier synthesis errors", "correct": false, "explanation": "Restarting from scratch discards the useful ransomware coverage already produced and is unnecessary when a targeted, incremental re-delegation would fill the specific gap."}, {"letter": "D", "text": "Evaluate the synthesis output for gaps, then re-delegate targeted queries on vendor risk before re-invoking synthesis", "correct": true, "explanation": "The coordinator should evaluate synthesis output for gaps and re-delegate targeted queries to search and analysis subagents, then re-invoke synthesis until coverage is sufficient."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-036", "domain": 1, "task_id": "1.2", "objective": "Orchestrate multi-agent systems with coordinator-subagent patterns", "situation": "An architect designs a workflow where a coordinator's subagent spawns its own subagent, which spawns another, several levels deep.", "question": "According to Claude Code's documentation, what is the default maximum nesting depth for subagents below the main conversation, and what happens once a subagent reaches that depth?", "options": [{"letter": "A", "text": "Nesting is unlimited by default as long as every subagent keeps the Agent tool in its allowed tools list, no matter how many layers deep the delegation chain eventually runs.", "correct": false, "explanation": "Keeping the Agent tool in a subagent's allowed tools list does not lift the nesting cap. Claude Code enforces a default limit on how many layers of subagents can spawn further subagents below the main conversation, and once a subagent reaches that limit the Agent tool is withheld from it regardless of its configured tool list."}, {"letter": "B", "text": "By default, a subagent may spawn its own subagents up to three layers below the main conversation, and that default depth can be changed with an environment variable.", "correct": true, "explanation": "By default a subagent can spawn subagents of its own up to three layers below the main conversation. At that depth limit, Claude Code withholds the Agent tool from every regular subagent, so it must finish its delegated work itself and return a summary; a forked conversation keeps Agent in its tool list but the tool errors instead of spawning. The default depth can be raised or lowered by setting the CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH environment variable."}, {"letter": "C", "text": "Only the main coordinator agent may spawn subagents on its own; no subagent may ever spawn another one, regardless of how the session or its settings are configured.", "correct": false, "explanation": "Subagents are not restricted to only the main coordinator spawning them. A subagent can spawn subagents of its own, and that next layer can spawn again, up to a documented depth limit measured below the main conversation."}, {"letter": "D", "text": "Nesting stops automatically once a subagent is two layers below the main conversation unless its model is upgraded to a more capable tier for that specific delegation.", "correct": false, "explanation": "The nesting depth limit has nothing to do with which model a subagent uses. The default cap on how many layers of subagents can spawn further subagents is controlled by an environment variable, not by upgrading a subagent to a more capable model tier."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f1-037", "domain": 1, "task_id": "1.3", "objective": "Configure subagent invocation, context passing, and spawning", "situation": "A coordinator runs a \"web-researcher\" subagent that gathers ten sources on a topic, then spawns a \"synthesis\" subagent to write the final report. The synthesis subagent's output ignores nearly all of the researcher's findings and instead re-derives generic conclusions from scratch.", "question": "The coordinator's prompt to synthesis only said \"Write a report based on the research that was just completed.\" What change fixes this?", "options": [{"letter": "A", "text": "Switch the synthesis subagent's model to a larger model so it infers the missing research context from the phrase \"just completed\" more reliably", "correct": false, "explanation": "A larger model cannot infer specific research findings it was never shown; model choice does not substitute for missing context in the prompt."}, {"letter": "B", "text": "Grant synthesis the same search tools as web-researcher so it can re-run the original queries itself before drafting the final written report", "correct": false, "explanation": "Re-running the same searches is wasteful and non-deterministic; it does not guarantee the same findings and defeats the purpose of reusing completed work."}, {"letter": "C", "text": "Include the complete findings from web-researcher directly in synthesis's prompt, since subagents never automatically inherit parent or sibling context", "correct": true, "explanation": "Each subagent starts with a fresh context window; the only channel from parent to subagent is the prompt string, so any prior findings the subagent needs must be written into that prompt explicitly."}, {"letter": "D", "text": "Increase the synthesis subagent's maxTurns so it has enough turns to independently rediscover the same ten sources the researcher already found", "correct": false, "explanation": "More turns does not give the subagent access to data it was never given; it would just spend more turns without the source material."}], "correct": "C", "select": 1, "group": "B"}, {"id": "f1-038", "domain": 1, "task_id": "1.5", "objective": "Apply Agent SDK hooks for tool call interception and data normalization", "situation": "An architect is deciding between two enforcement strategies for a compliance rule that blocks database deletions outside business hours: a PreToolUse hook that checks the current time and denies the call, versus a system prompt clause instructing the agent to avoid deletions outside business hours.", "question": "Which statement correctly characterizes the tradeoff?", "options": [{"letter": "A", "text": "Both approaches are functionally equivalent because the Agent SDK's runtime translates system prompt rules into tool-use permission checks, so the deletion is always prevented regardless of enforcement method.", "correct": false, "explanation": "The Agent SDK does not automatically translate system prompt instructions into tool-use permission checks. Hooks are explicit developer-implemented code; prompt clauses rely on model generation, not runtime enforcement, so the approaches are not functionally equivalent."}, {"letter": "B", "text": "The prompt clause enforces the rule deterministically because Claude always follows explicit system prompt instructions verbatim once they are stated clearly enough, ensuring any outside-hours deletion is blocked.", "correct": false, "explanation": "Claude does not always follow system prompt instructions verbatim, even if clearly stated; it may misapply, ignore, or misinterpret rules in natural language. Deterministic hooks exist precisely to provide guaranteed enforcement that prompt instructions cannot ensure."}, {"letter": "C", "text": "The hook only affects what the user sees in the transcript, so the underlying tool call still executes even when the hook returns a deny decision, meaning the database record is deleted despite the compliance rule.", "correct": false, "explanation": "When a PreToolUse hook returns a deny decision, the SDK blocks the tool call from executing entirely, not just its appearance in the transcript. The deletion would not occur, contradicting the claim that the record is deleted."}, {"letter": "D", "text": "The hook deterministically enforces the rule on every matching tool call regardless of output, while the prompt clause only influences the likelihood that a compliant call is generated.", "correct": true, "explanation": "A PreToolUse hook executes as deterministic code outside the model's generation, and returning a deny decision unconditionally blocks the call regardless of model output. In contrast, a prompt clause only makes it more probable that the model generates a compliant call, but does not guarantee enforcement."}], "correct": "D", "select": 1, "group": "E"}, {"id": "f1-039", "domain": 1, "task_id": "1.6", "objective": "Design task decomposition strategies for complex workflows", "situation": "A team is tasked with adding comprehensive tests to a 200,000-line legacy codebase with no existing test suite and undocumented module boundaries.", "question": "Which sequence of decomposition steps best reflects an adaptive strategy for this open-ended task?", "options": [{"letter": "A", "text": "Map the codebase structure to find modules and dependencies, identify high-impact gaps, then build a test plan that adjusts as dependencies surface", "correct": true, "explanation": "For an open-ended task with unknown structure, first mapping the codebase, then identifying priorities, then adapting the plan as dependencies surface is the recommended way to decompose the work incrementally based on discoveries."}, {"letter": "B", "text": "Generate test files for every single source file in strict alphabetical order without first assessing how the modules actually relate to one another", "correct": false, "explanation": "Alphabetical ordering ignores which modules are highest-impact or most interdependent, so effort is not prioritized and discovered dependencies cannot reshape the plan."}, {"letter": "C", "text": "Write one exhaustive test file covering every function in the codebase in a single continuous pass before any tests are ever run", "correct": false, "explanation": "Attempting one exhaustive pass across 200,000 lines without first understanding structure risks missing high-impact areas and cannot adapt when unexpected dependencies are found mid-task."}, {"letter": "D", "text": "Have one subagent write every single test for the entire codebase in one uninterrupted session with no intermediate checkpoints along the way", "correct": false, "explanation": "A single uninterrupted session with no checkpoints prevents the plan from adapting to what is discovered about module boundaries partway through, which is essential for an undocumented legacy codebase."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f1-040", "domain": 1, "task_id": "1.4", "objective": "Implement multi-step workflows with enforcement and handoff patterns", "situation": "A PreToolUse hook gating process_refund throws an unhandled exception whenever the internal verification service it calls times out. During an outage of that service, every refund attempt in the session crashes the agent entirely instead of being cleanly denied.", "question": "What change to the hook fixes this without weakening the enforcement it provides?", "options": [{"letter": "A", "text": "Catch the timeout inside the hook and return permissionDecision \"deny\" with a reason explaining the outage, instead of letting the exception propagate and crash the session", "correct": true, "explanation": "Catching the error and returning a deny with a clear reason keeps the crash from propagating while preserving the enforcement guarantee: refunds still cannot proceed when verification cannot be confirmed."}, {"letter": "B", "text": "Catch the timeout and return permissionDecision \"allow\", since a service outage is not the customer's fault and refunds should never be penalized for infrastructure issues", "correct": false, "explanation": "Allowing the refund when verification could not be confirmed defeats the purpose of the gate; an inability to verify is not the same as a successful verification."}, {"letter": "C", "text": "Remove the hook entirely for the duration of any outage of the verification service, so refunds can proceed unblocked until that service comes back online", "correct": false, "explanation": "Removing the gate during an outage is exactly the scenario deterministic enforcement exists to prevent; it allows unverified refunds through when the check is most needed."}, {"letter": "D", "text": "Increase the hook's timeout value to several hours, so the hook keeps waiting on the verification service instead of failing quickly during a temporary outage", "correct": false, "explanation": "A multi-hour timeout does not fix the underlying failure mode; it just delays the same unhandled-exception crash instead of preventing it, and stalls every refund in the meantime."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f1-041", "domain": 1, "task_id": "1.3", "objective": "Configure subagent invocation, context passing, and spawning", "situation": "An architect defines a performance-optimizer subagent with the intention that the main Claude coordinator will delegate to it automatically whenever a user asks about slow database queries. In practice, the coordinator almost always answers query-tuning questions directly instead of delegating. The subagent definition's prompt field is detailed and technically accurate.", "question": "What is the most likely reason automatic delegation is failing?", "options": [{"letter": "A", "text": "The tools field lists Bash and Read, which are too permissive for a subagent Claude expects to invoke only for read-only analysis tasks.", "correct": false, "explanation": "While least-privilege tool access is a recommended subagent design consideration, a permissive tool list does not by itself prevent Claude from delegating to the subagent. Automatic delegation is triggered primarily by a clear description, not by whether the listed tools match a read-only analysis profile."}, {"letter": "B", "text": "The subagent's prompt is written in the second person, which the delegation matcher treats as intended for the end user rather than the subagent.", "correct": false, "explanation": "Anthropic documentation does not describe second-person prompt style as a delegation matcher issue. Subagent prompts commonly address Claude directly, and the routing signal is the subagent's description and declared purpose rather than the grammatical person used in the prompt."}, {"letter": "C", "text": "The coordinator's own system prompt was not updated to reference performance-optimizer, so the runtime cannot resolve the subagent name at dispatch time.", "correct": false, "explanation": "Anthropic's subagent pattern does not require the coordinator's system prompt to name every available subagent. Delegation is controlled by the subagent definition and its description; the runtime resolves the subagent through the configured tool/subagent interface, so omitting a reference in the coordinator prompt is not the documented cause."}, {"letter": "D", "text": "The subagent's description field is vague or missing, so Claude has no trigger signal for when this subagent applies and defaults to handling the task itself.", "correct": true, "explanation": "Anthropic's tool-use and subagent guidance treats the description field as the vital trigger signal for delegation. A vague or missing description leaves Claude without a clear mapping between user requests and the subagent's purpose, so the model may answer directly instead of invoking the tool. Since the prompt is detailed and technically accurate, the missing or weak description is the most likely cause."}], "correct": "D", "select": 1, "group": "B"}, {"id": "f1-042", "domain": 1, "task_id": "1.7", "objective": "Manage session state, resumption, and forking", "situation": "An architect is designing a serverless pipeline where each stage runs in a fresh, short-lived container and cannot guarantee access to any previous container's local disk. The pipeline still needs later stages to act on the conclusions of an earlier stage's investigation.", "question": "Which strategy best fits this constraint?", "options": [{"letter": "A", "text": "Capture the earlier stage's key results as application state and pass them into a new session's opening prompt", "correct": true, "explanation": "This is the recommended pattern for stateful serverless pipelines. Serverless environments are stateless by design; each invocation starts fresh without prior context. The official approach is to externalize state to a durable store (e.g., database, cache) and retrieve it at the beginning of each new session or step. Anthropic's own architecture for AI agents uses a 'session as durable event log' that persists past prompts and outputs, explicitly passing them into new sessions to overcome short-lived constraints."}, {"letter": "B", "text": "Pass the earlier stage's session ID to resume in the next container and expect the transcript to be found automatically", "correct": false, "explanation": "A session ID alone is insufficient because serverless functions do not retain local storage or automatically reconnect to previous sessions. The pipeline would need to store the transcript externally and retrieve it using the ID. Without explicit state capture and passing, a new container has no built-in mechanism to access the earlier stage's data."}, {"letter": "C", "text": "Set fork_session=True in the next container so it branches from the earlier stage's session ID directly", "correct": false, "explanation": "There is no standard serverless API or flag like fork_session=True that allows a new container to automatically inherit or branch from a previous container's session. Serverless functions are isolated and stateless; any sharing of context must be managed through external state stores."}, {"letter": "D", "text": "Rely on claude --continue in the next container to pick up the most recent local session automatically", "correct": false, "explanation": "The claude --continue command would attempt to resume a local session, but in a serverless environment, each container is ephemeral and does not have access to the previous container's local filesystem. Local sessions are not shared across invocations, so this approach would fail to find the earlier stage's data."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f1-043", "domain": 1, "task_id": "1.2", "objective": "Orchestrate multi-agent systems with coordinator-subagent patterns", "situation": "A coordinator needs to orchestrate a large-scale codebase migration involving on the order of two hundred independent file-level subtasks, far more than the handful of subagents a coordinator typically delegates to per turn.", "question": "Which approach is best suited to this scale?", "options": [{"letter": "A", "text": "Split the migration across multiple coordinators that each independently maintain their own separate pipeline", "correct": false, "explanation": "Multiple independent coordinators with no shared state would duplicate effort and lose the centralized tracking needed to know which of two hundred subtasks are complete."}, {"letter": "B", "text": "Invoke a single general-purpose subagent and have it sequentially handle all subtasks within one context window", "correct": false, "explanation": "A single subagent handling two hundred subtasks sequentially in one context window loses the parallelism and context-isolation benefits, and risks exhausting that context window."}, {"letter": "C", "text": "Keep using turn-by-turn subagent delegation but increase the coordinator's maxTurns to accommodate more invocations", "correct": false, "explanation": "Raising maxTurns lets a single agent take more turns, but turn-by-turn subagent delegation is still designed for a handful of subagents per turn, not hundreds of coordinated subtasks."}, {"letter": "D", "text": "Use the Workflow tool to move orchestration into a script the runtime executes outside the conversation itself", "correct": true, "explanation": "For runs coordinating dozens to hundreds of agents, the Workflow tool moves orchestration into a script the runtime executes outside the conversation, which fits this scale better than turn-by-turn delegation."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-044", "domain": 1, "task_id": "1.1", "objective": "Design and implement agentic loops for autonomous task execution", "situation": "A team is building an agent that queries a ticketing system, then a knowledge base, then drafts a reply, using the Messages API directly rather than a higher-level SDK. They debate whether their own code must execute the tools Claude requests.", "question": "Which statement accurately describes the responsibility split in this direct integration?", "options": [{"letter": "A", "text": "Claude executes the ticketing system and knowledge base calls internally on Anthropic's servers, so the application only renders the final answer", "correct": false, "explanation": "Anthropic's servers do not execute custom client-side tools such as a ticketing system lookup; those are implemented and run by the developer's own code."}, {"letter": "B", "text": "Tool execution happens only if the application registers a webhook URL that Anthropic calls back synchronously during response generation", "correct": false, "explanation": "There is no synchronous webhook-callback mechanism for client-side tools; the caller executes after receiving the response, not via a registered callback during generation."}, {"letter": "C", "text": "Their application must execute each requested tool and submit results as tool_result blocks, since the API never runs client-side tools itself", "correct": true, "explanation": "Client-side tools are defined and executed by the caller; Claude only returns a tool_use request, and the application must run it and return a tool_result to continue the loop."}, {"letter": "D", "text": "The Messages API automatically executes any tool whose name matches a function already defined within the calling application's runtime", "correct": false, "explanation": "The API cannot introspect or auto-invoke functions in the caller's runtime; execution is entirely the caller's responsibility after reading the tool_use block."}], "correct": "C", "select": 1, "group": "A"}, {"id": "f1-045", "domain": 1, "task_id": "1.2", "objective": "Orchestrate multi-agent systems with coordinator-subagent patterns", "situation": "A coordinator for a market-research assistant splits the broad query 'analyze the competitive landscape for electric vehicle charging networks' into narrow subtasks like 'find charger connector types' and 'list charging speeds,' each assigned to a separate subagent. The synthesized report ends up missing pricing models, regulatory incentives, and major competitors entirely.", "question": "What went wrong?", "options": [{"letter": "A", "text": "The synthesis step ran before all subagents had finished, so the coordinator aggregated partial results instead of complete ones", "correct": false, "explanation": "Nothing in the scenario indicates a timing or synchronization issue; the subagents' assigned subtasks themselves never covered pricing, incentives, or competitors."}, {"letter": "B", "text": "The subagents lacked sufficient tool access to search external sources, so they returned incomplete results for their subtasks", "correct": false, "explanation": "The scenario describes subagents accurately answering their narrow assigned subtasks; the problem is which subtasks were assigned, not their ability to search for information."}, {"letter": "C", "text": "The subagents duplicated each other's work on connector types, leaving no time to research the remaining subtopics", "correct": false, "explanation": "The two subtasks described are distinct, not duplicated, so overlapping effort isn't what caused the missing coverage."}, {"letter": "D", "text": "The coordinator decomposed the query too narrowly, so the subtasks covered only a few facets and left broad areas uncovered", "correct": true, "explanation": "Overly narrow task decomposition by the coordinator leaves broad research topics incompletely covered, since the chosen subtasks addressed only a slice of the competitive landscape."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-046", "domain": 1, "task_id": "1.6", "objective": "Design task decomposition strategies for complex workflows", "situation": "A customer onboarding workflow always performs three steps for every new account in the same order: create the account record, provision default permissions, and send a welcome email. None of these steps ever branch based on account data. Separately, a fraud investigation workflow examines a flagged account by pulling transaction history, and the specific records to inspect next depend entirely on what suspicious patterns are found in the transactions already reviewed.", "question": "Which pairing of decomposition strategies correctly matches each workflow?", "options": [{"letter": "A", "text": "Onboarding should use a fixed prompt chain since its steps never branch, while fraud investigation should use adaptive decomposition based on findings", "correct": true, "explanation": "Onboarding's three steps are fixed and never branch, matching the case for prompt chaining, while fraud investigation's next steps depend on discovered transaction patterns, matching the case for adaptive, dynamically generated subtasks."}, {"letter": "B", "text": "Both workflows should use a dynamic orchestrator, since dynamic decomposition is always the safer default no matter how predictable a workflow's steps are", "correct": false, "explanation": "Defaulting to dynamic orchestration for both workflows adds unnecessary runtime decision-making to onboarding, whose three steps are already fully predictable and never change."}, {"letter": "C", "text": "Both workflows should use a fixed prompt chain, since any workflow involving financial data must follow one strictly predetermined sequence of steps", "correct": false, "explanation": "Handling financial data does not by itself dictate a fixed sequence; the deciding factor is whether the steps are predictable in advance, and the fraud investigation explicitly is not."}, {"letter": "D", "text": "Onboarding should use adaptive decomposition since account creation always carries some risk of failure, while fraud investigation should use a fixed three-step chain", "correct": false, "explanation": "The possibility of account creation failing does not mean the onboarding steps themselves are unpredictable, and capping fraud investigation at exactly three steps ignores that the number of records to inspect depends on findings."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f1-047", "domain": 1, "task_id": "1.4", "objective": "Implement multi-step workflows with enforcement and handoff patterns", "situation": "An architect writes a PreToolUse hook with the regex matcher /refund/ intending it to gate a single custom tool named refund. During a review, a colleague notices the matcher would also fire on a tool named issue_refund_note that only writes an internal comment and should never be gated.", "question": "What is the cause and the correct fix?", "options": [{"letter": "A", "text": "The matcher /refund/ uses a regular expression, but regular expression matchers in hooks are automatically anchored to match exactly, so it would only fire on the tool named exactly refund. The colleague's concern is unfounded; no fix is needed.", "correct": false, "explanation": "Regular expression matchers in hooks are not automatically anchored. An unanchored regex like /refund/ will match any tool name containing the substring 'refund', such as issue_refund_note. Anchors (^ and $) must be explicitly added to force an exact match."}, {"letter": "B", "text": "The hook fires on every single tool call by default unless a timeout value is explicitly configured. Adding a timeout would stop it from matching issue_refund_note.", "correct": false, "explanation": "A PreToolUse hook fires only for tool calls that match its configured matcher pattern; it does not fire on every call. The timeout setting controls how long the hook's script is allowed to run, not which tool names the hook targets. It has no effect on matching or filtering tool names."}, {"letter": "C", "text": "Matchers can only target tool names that begin with the mcp__ server prefix. A custom in-process tool like refund must be renamed to include this prefix before any matcher can reliably target it.", "correct": false, "explanation": "Tool name matchers in PreToolUse hooks are not restricted to the mcp__ prefix. They can match any tool name, whether it is a built-in tool (e.g., Bash, Write), an MCP tool, or a custom in‑process tool. Renaming is unnecessary."}, {"letter": "D", "text": "The regex matcher /refund/ is unanchored, so it matches any tool name containing the substring 'refund'. The fix is to anchor the regex with ^ and $: /^refund$/, ensuring it only matches the exact tool name refund.", "correct": true, "explanation": "In Claude Code hooks, a matcher enclosed in forward slashes is treated as a JavaScript regular expression. Without anchors, the pattern matches any tool name that contains 'refund', such as issue_refund_note. Adding ^ (start of string) and $ (end of string) restricts the match to exactly 'refund'. Alternatively, the architect could use the exact string matcher 'refund' (without slashes), which performs an exact string match by default according to the official documentation, but given the current use of a regex, anchoring is the direct fix."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-048", "domain": 1, "task_id": "1.2", "objective": "Orchestrate multi-agent systems with coordinator-subagent patterns", "situation": "During a long-running research task, a foreground subagent produces several paragraphs of analysis before a server overload error cuts it off mid-response.", "question": "According to Anthropic’s official documentation (Claude Code v2.1.199+), what does the coordinator receive as the Agent tool result?", "options": [{"letter": "A", "text": "A completely empty result with no indication that an error occurred, requiring the coordinator to poll for status separately", "correct": false, "explanation": "Anthropic’s documentation states that when a foreground subagent is cut off by an API error after producing text output, the Agent tool returns the partial output with a note, not an empty result. No polling is required."}, {"letter": "B", "text": "The full expected output, reconstructed from cached intermediate tool calls made before the overload occurred", "correct": false, "explanation": "There is no documented mechanism to reconstruct the full output from cached tool calls. Only the partial text already generated is returned. If a subagent produces no text output or only tool calls before an error, it fails with a different message (‘Agent terminated early due to an API error’)."}, {"letter": "C", "text": "An automatic retry that silently reruns the subagent from the start until it completes without error", "correct": false, "explanation": "The default error handling does not include automatic retries. The documented behavior is to return the partial output and a note that the subagent didn’t finish, leaving it to the coordinator to decide how to proceed (e.g., using iterative retrieval)."}, {"letter": "D", "text": "The partial text output the subagent already produced, along with a note that the subagent didn’t finish", "correct": true, "explanation": "Anthropic’s official documentation explicitly states: ‘If a rate limit, overload, or server error cuts off a foreground subagent that already produced text output, the Agent tool returns that partial output with a note that the subagent didn’t finish.’ This behavior requires Claude Code v2.1.199 or later and gives the coordinator the incomplete analysis and an explicit status to inform follow-up actions."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-049", "domain": 1, "task_id": "1.7", "objective": "Manage session state, resumption, and forking", "situation": "An architect resumed a session this morning expecting yesterday's investigation, but the agent behaved as if starting completely fresh with no memory of prior findings. The team confirms the correct session ID was passed.", "question": "What should the architect check first?", "options": [{"letter": "A", "text": "Whether the teammate who ran the original session had sufficient tool permissions granted", "correct": false, "explanation": "tool permissions affect what actions the agent may take within a session, not whether the resume call can locate and load the correct transcript file."}, {"letter": "B", "text": "Whether the resume call ran from the same working directory the original session was started in", "correct": true, "explanation": "Sessions are stored under a path keyed by the encoded working directory; resuming from a different cwd looks in the wrong location and silently returns a fresh session, which matches the documented most common cause of this symptom."}, {"letter": "C", "text": "Whether the prompt text used to resume matched the original prompt text exactly", "correct": false, "explanation": "resume looks up a session by its ID or name, not by matching the new prompt text against the original prompt, so prompt wording differences do not explain a fresh session appearing."}, {"letter": "D", "text": "Whether the original session had exceeded its maximum turn or budget limit before ending", "correct": false, "explanation": "hitting a max-turns or max-budget limit ends a session with an error result but does not prevent it from being found and resumed later; it doesn't cause a fresh session to appear instead."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f1-050", "domain": 1, "task_id": "1.3", "objective": "Configure subagent invocation, context passing, and spawning", "situation": "A session has already analyzed an authentication module in depth and produced a JWT-based redesign plan. The architect now wants to explore what an OAuth2-based redesign would look like using the same accumulated analysis as a starting point, but only if the JWT thread can still be resumed unchanged afterward.", "question": "Which approach satisfies both requirements?", "options": [{"letter": "A", "text": "Continue the most recent session in the working directory, trusting that it preserves the JWT thread as a separate branch once a new topic is introduced", "correct": false, "explanation": "Continue simply resumes the most recent session in place; it does not create a separate branch, so the OAuth2 detour would still merge into the single existing thread."}, {"letter": "B", "text": "Start a brand-new session with no prior context and paste a short written summary of the JWT plan into the first prompt before asking about OAuth2", "correct": false, "explanation": "A fresh session with a pasted summary loses the full accumulated analysis (files read, intermediate reasoning) that forking preserves automatically."}, {"letter": "C", "text": "Resume the session with fork enabled to copy the current history into a new session id, then explore OAuth2 there while the original stays untouched", "correct": true, "explanation": "Forking creates a new session that starts as a copy of the source session's history and diverges from there; the original session and its id are left untouched, so it can still be resumed for the JWT thread."}, {"letter": "D", "text": "Resume the existing session directly and ask about OAuth2 in the same conversation, then resume it again afterward and ask it to disregard the OAuth2 detour", "correct": false, "explanation": "Resuming directly appends to the same history, so the OAuth2 exploration becomes permanently part of the one thread instead of leaving the original JWT-focused history untouched."}], "correct": "C", "select": 1, "group": "B"}, {"id": "f1-051", "domain": 1, "task_id": "1.4", "objective": "Implement multi-step workflows with enforcement and handoff patterns", "situation": "A coordinator spawns two subagents in parallel to investigate the billing and address concerns from a multi-concern ticket. Before writing the final customer response, the coordinator needs to know that both subagents have actually finished and to collect what each one found.", "question": "Which hook is designed for this?", "options": [{"letter": "A", "text": "Notification, which fires for status events such as permission prompts or idle states, not for subagent completion or its results", "correct": false, "explanation": "Notification fires for status events such as permission prompts or idle states; it is not the mechanism for tracking when a subagent finishes or what it found."}, {"letter": "B", "text": "PreToolUse, which fires immediately before a tool call executes and cannot preview a result that has not been produced yet", "correct": false, "explanation": "PreToolUse fires immediately before a tool call executes and only inspects or blocks that call; it has no visibility into a result that has not been produced yet."}, {"letter": "C", "text": "UserPromptSubmit, which fires when the customer's original ticket text is submitted, before either subagent has begun investigating", "correct": false, "explanation": "UserPromptSubmit fires when the user's prompt is submitted, before either subagent has investigated anything, so it cannot carry findings that do not exist yet."}, {"letter": "D", "text": "SubagentStop, which fires when a subagent finishes and lets the coordinator log or aggregate that subagent's results", "correct": true, "explanation": "SubagentStop fires when a subagent finishes its work, making it the point where the coordinator can log completion and aggregate results from parallel investigations before synthesis."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-052", "domain": 1, "task_id": "1.5", "objective": "Apply Agent SDK hooks for tool call interception and data normalization", "situation": "A support-automation agent has a refund_customer MCP tool. Company policy requires that refunds over $500 never be issued autonomously and must instead be routed to a human reviewer. The architect currently relies on a system prompt telling Claude never to approve refunds above that amount. During testing, the agent occasionally approves a $600 refund anyway.", "question": "What is the most effective fix?", "options": [{"letter": "A", "text": "You can register a PreToolUse hook on refund_customer that inspects tool_input.amount and returns permissionDecision \"deny\" with a reason whenever the amount exceeds $500 redirecting the workflow to escalation.", "correct": true, "explanation": "A PreToolUse hook intercepts the tool call before execution and can deterministically block it based on the amount in tool_input, preventing the refund entirely. This approach converts a prompt-level suggestion into an enforced business rule, ensuring the $500 threshold is never exceeded autonomously."}, {"letter": "B", "text": "Lower the model's temperature to 0.1 for any session that calls refund_customer, forcing deterministic tool-call decisions that respect the $500 limit more consistently and prevent unauthorized high-value approvals.", "correct": false, "explanation": "Reducing temperature to 0.1 makes generation less random but does not guarantee that every tool-call decision respects the $500 limit; the model may still make mistakes. This mitigates the probability of violations but cannot prevent them absolutely."}, {"letter": "C", "text": "Rewrite the system prompt to state the $500 threshold rule three times using varied phrasings and include a dedicated section that precedes every example conversation, making the constraint more salient during generation.", "correct": false, "explanation": "Prompt engineering remains probabilistic; no amount of repetition or salience guarantees that the model will not sometimes approve a $600 refund. Testing already showed that the agent violated the rule despite the system prompt, so reinforcing the prompt does not eliminate the risk."}, {"letter": "D", "text": "Register a PostToolUse hook matched to refund_customer that examines tool_input.amount and, if it exceeds $500, writes the transaction details to an audit log and triggers a post-hoc review process to flag the violation.", "correct": false, "explanation": "A PostToolUse hook executes after the refund tool has run, meaning it only logs the violation after the money has been transferred. It cannot prevent the autonomous approval because the tool call has already completed."}], "correct": "A", "select": 1, "group": "E"}, {"id": "f1-053", "domain": 1, "task_id": "1.1", "objective": "Design and implement agentic loops for autonomous task execution", "situation": "A build-automation loop's exception handler catches a tool execution failure, appends a tool_result block with \"is_error\": true describing the failure, and sends the updated conversation back to Claude. A reviewer argues the loop should terminate immediately on any tool failure instead.", "question": "Why is appending the error and continuing generally the better design here?", "options": [{"letter": "A", "text": "It lets Claude see the failure in context and decide the next step, such as retrying with different arguments, consistent with model-driven reasoning", "correct": true, "explanation": "Feeding the error back as a tool_result keeps it visible in context, letting Claude reason over it and decide autonomously whether to retry, adjust arguments, or take another path."}, {"letter": "B", "text": "It guarantees the same tool will succeed on the next attempt, because error tool_results reset that tool's internal rate limits", "correct": false, "explanation": "Marking a tool_result as an error has no effect on rate limits or future success; it only informs Claude that the prior attempt failed."}, {"letter": "C", "text": "It is required, since the Messages API automatically terminates any conversation that lacks a correctly formatted is_error field after any single failure", "correct": false, "explanation": "The Messages API does not require an is_error field or auto-terminate conversations lacking one; it is simply a convention for marking a failed tool_result."}, {"letter": "D", "text": "It removes the need for any stop_reason checks, since is_error becomes the sole termination signal for the rest of the loop", "correct": false, "explanation": "is_error does not replace stop_reason for termination; the loop still relies on stop_reason to know whether Claude is requesting another tool or has finished."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f1-054", "domain": 1, "task_id": "1.7", "objective": "Manage session state, resumption, and forking", "situation": "An application serves many concurrent users, each with their own long-running investigation conversation against the same repository directory. A backend service needs to resume a specific user's conversation on demand, potentially hours after the user's last message, while other users' sessions remain active in the same directory.", "question": "Which session option should the service use?", "options": [{"letter": "A", "text": "Track each user's captured session ID and pass it to resume when that user sends a follow-up", "correct": true, "explanation": "With multiple concurrent sessions sharing a directory, resume with a tracked session ID is required to return to a specific user's conversation rather than whichever one was most recently active."}, {"letter": "B", "text": "Set fork_session=True on every request so each user gets an isolated copy of the directory's latest session", "correct": false, "explanation": "fork_session branches from an already-selected session; without first resuming the correct session ID, it would just branch off the wrong user's latest activity."}, {"letter": "C", "text": "Pass continue_conversation=True (or continue: true) on every incoming request regardless of user", "correct": false, "explanation": "continue always targets the most recently active session in the directory, which would route different users into whichever conversation last had activity, not their own."}, {"letter": "D", "text": "Rely on the session picker to let the backend service select the correct user's conversation", "correct": false, "explanation": "the interactive session picker is a human-facing terminal feature and is not a programmatic mechanism a backend service can use to route requests."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f1-055", "domain": 1, "task_id": "1.4", "objective": "Implement multi-step workflows with enforcement and handoff patterns", "situation": "A customer's message covers three distinct concerns: a late shipment, a coupon that failed to apply, and a request to close their account. The agent has separate tools for shipment tracking, coupon validation, and account closure. The architect wants the fastest accurate resolution while keeping the final reply coherent.", "question": "What ordering of steps best achieves this?", "options": [{"letter": "A", "text": "Investigate only the account-closure request first, since closing the account would make the shipment and coupon concerns moot and therefore not worth investigating at all", "correct": false, "explanation": "Account closure does not resolve whether the shipment already occurred or whether the coupon issue needs a refund; skipping investigation of the other two concerns can leave real issues unresolved."}, {"letter": "B", "text": "Split the message into the three concerns, investigate shipment tracking, coupon validation, and account status concurrently, then combine the findings into one synthesized reply", "correct": true, "explanation": "Decomposing the multi-concern message into distinct items and investigating them in parallel with shared context, then synthesizing one unified resolution, resolves everything accurately without unnecessary round trips."}, {"letter": "C", "text": "Ask the customer to restate their single message as three separate, single-concern tickets before any investigation ever begins on any of the three items", "correct": false, "explanation": "Pushing the decomposition work back onto the customer is unnecessary friction when the agent is capable of splitting and investigating all three concerns itself."}, {"letter": "D", "text": "Investigate the shipment concern to full resolution and reply to the customer about it alone, then wait for a follow-up message before ever considering the coupon or account-closure concerns", "correct": false, "explanation": "Resolving only one concern and waiting for a follow-up message ignores the other two concerns the customer already raised and adds unnecessary delay."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f1-056", "domain": 1, "task_id": "1.1", "objective": "Design and implement agentic loops for autonomous task execution", "situation": "A team is implementing termination logic for an autonomous refactoring agent. Loop A exits when response.stop_reason == \"end_turn\". Loop B exits when the assistant's final text block is non-empty, on the assumption that any explanatory text means Claude is finished.", "question": "Which loop correctly implements the standard termination pattern?", "options": [{"letter": "A", "text": "Loop A, because stop_reason directly reports whether Claude finished its turn without requesting a tool, which is the documented signal", "correct": true, "explanation": "Claude can emit explanatory text alongside a tool_use block in the same turn, so checking for text presence stops the loop too early; stop_reason is the structured, authoritative field."}, {"letter": "B", "text": "Loop B, because a populated final text block is treated as the only reliable indicator that Claude has stopped requesting further tool calls", "correct": false, "explanation": "A response can include commentary text while still requesting another tool, so non-empty text does not reliably mean the turn is finished."}, {"letter": "C", "text": "Neither loop works, because the API only signals completion through the total count of content blocks returned in the response", "correct": false, "explanation": "Content block counts vary with output style and are not a documented completion signal in the Messages API."}, {"letter": "D", "text": "Both loops behave identically in practice, since stop_reason and trailing text emptiness always change together on every response", "correct": false, "explanation": "Text presence and stop_reason are not guaranteed to move together; a turn can carry both a tool_use block and accompanying text."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f1-057", "domain": 1, "task_id": "1.3", "objective": "Configure subagent invocation, context passing, and spawning", "situation": "A coordinator agent is configured with allowedTools set to [\"Read\", \"Grep\", \"Glob\"] and defines a \"research-assistant\" subagent in its agents map. When the coordinator tries to delegate a task to research-assistant, every invocation halts on a permission prompt instead of running automatically.", "question": "What is the most likely cause and fix?", "options": [{"letter": "A", "text": "The subagent's tools array lists more permissions than the coordinator holds, so the runtime blocks the escalation until an operator confirms it manually", "correct": false, "explanation": "A subagent's tools are a restriction subset, not an escalation; there is no privilege-escalation check that blocks delegation for this reason."}, {"letter": "B", "text": "The research-assistant definition is missing a model field, so the runtime cannot select a default model and pauses for confirmation before every call", "correct": false, "explanation": "model is optional on an agent definition and defaults to the main model when omitted; its absence does not trigger permission prompts."}, {"letter": "C", "text": "The coordinator's system prompt does not mention the subagent by name, so the permission layer treats each delegation as an unrecognized action requiring review", "correct": false, "explanation": "Naming a subagent in the prompt affects whether Claude chooses to invoke it, not whether the invocation is auto-approved once chosen."}, {"letter": "D", "text": "The subagent invocation tool, Task, is not listed in allowedTools, so each spawn attempt requires manual approval; add Task to allowedTools to auto-approve subagent calls", "correct": true, "explanation": "Spawning a subagent goes through the dedicated invocation tool, and that tool name must be present in allowedTools for the coordinator to auto-approve delegation; otherwise every call falls through to a permission prompt."}], "correct": "D", "select": 1, "group": "B"}, {"id": "f1-058", "domain": 1, "task_id": "1.1", "objective": "Design and implement agentic loops for autonomous task execution", "situation": "An architect is choosing between the Client SDK and the Agent SDK for an internal automation that must read files, run shell commands, and iterate until a task is done, while avoiding a hand-written tool-execution loop.", "question": "Which factor should most influence this decision?", "options": [{"letter": "A", "text": "The Agent SDK bundles built-in tools and runs the agentic loop internally, while the Client SDK requires manually inspecting stop_reason", "correct": true, "explanation": "The Agent SDK wraps the same loop and tool execution that power Claude Code, so Claude handles tools autonomously; the Client SDK leaves the developer to implement the loop and check stop_reason directly."}, {"letter": "B", "text": "Both SDKs need an identical amount of custom loop code to handle stop_reason, so the choice should rest solely on language preference", "correct": false, "explanation": "The two SDKs differ substantially in how much loop code the team must write, since the Agent SDK exists specifically to remove that burden."}, {"letter": "C", "text": "The Client SDK cannot return a stop_reason value at all, making any tool-use loop impossible to build without adopting the Agent SDK", "correct": false, "explanation": "The Client SDK does return stop_reason as part of the Messages API response; that field is exactly what a hand-built loop inspects."}, {"letter": "D", "text": "The Agent SDK requires the team to manually append tool_result blocks after every call, unlike the Client SDK, which appends them automatically", "correct": false, "explanation": "This reverses the real split: manual tool_result handling is a Client SDK responsibility, while the Agent SDK manages execution and history internally."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f1-059", "domain": 1, "task_id": "1.4", "objective": "Implement multi-step workflows with enforcement and handoff patterns", "situation": "An agent handles two kinds of unverified requests: viewing a masked order history (read-only, low risk) and issuing a refund (financial, irreversible). The architect wants to gate both with PreToolUse hooks but use different permission decisions based on risk.", "question": "Which pairing of returned permissionDecision values best fits the two cases before identity is verified?", "options": [{"letter": "A", "text": "\"allow\" for both the masked order history lookup and the refund, since neither action can realistically be reversed once the workflow reaches the hook stage.", "correct": false, "explanation": "PreToolUse hooks support permissionDecision values such as allow, ask, and deny. Using allow for both would bypass the permission prompt entirely, including for the irreversible refund. Anthropic's fail-closed, deliberately conservative guidance requires human oversight for high-risk actions like refunds, so allow is unsafe here."}, {"letter": "B", "text": "\"deny\" for both the masked order history lookup and the refund, since any unverified request should be treated identically regardless of the underlying tool.", "correct": false, "explanation": "deny blocks both actions, preventing even the low-risk read-only lookup. Anthropic's risk-based permission model differentiates actions: a masked read can be allowed or asked, while an irreversible refund should be asked to enable human confirmation. Treating all unverified requests identically ignores the documented distinction between reversible and irreversible operations."}, {"letter": "C", "text": "\"ask\" for the refund and \"allow\" for the masked order history lookup, because the irreversible refund needs human confirmation before proceeding, while the masked, read-only lookup is low risk and can be permitted automatically.", "correct": true, "explanation": "Anthropic's PreToolUse hooks support permissionDecision values including allow, ask, and deny. For an irreversible financial action like a refund, ask is appropriate: it requires a human to confirm the request before execution, providing a path to approval without automatic execution. The masked order history lookup is read-only and low risk; allow lets it proceed directly, which is acceptable because the data is masked and the operation has no irreversible side effects."}, {"letter": "D", "text": "\"ask\" for the masked order history lookup and \"deny\" for the refund, since the reversible read can tolerate a manual check while the irreversible refund should not proceed at all.", "correct": false, "explanation": "This incorrectly applies deny to the refund. A deny decision blocks the tool call entirely and provides no human confirmation path. For a refund that may be legitimate but simply lacks identity verification, Anthropic's ask decision is designed specifically for irreversible actions that require a human-in-the-loop to approve before execution. The masked read could be allowed or asked, but the refund should be asked, not denyed."}], "correct": "C", "select": 1, "group": "A"}, {"id": "f1-060", "domain": 1, "task_id": "1.2", "objective": "Orchestrate multi-agent systems with coordinator-subagent patterns", "situation": "A team wants several Claude Code agents to message each other directly and coordinate on a shared set of tasks over an extended session, rather than one agent invoking others as one-shot subtasks and receiving only a final summary back.", "question": "Which pattern fits this requirement better than standard coordinator subagents?", "options": [{"letter": "A", "text": "Filesystem-based subagents stored in .claude/agents, since only these support persistent shared state fully", "correct": false, "explanation": "Filesystem-based subagents are still standard subagents invoked as one-shot subtasks by a parent; the storage location doesn't grant them direct messaging or shared-task coordination with each other."}, {"letter": "B", "text": "Standard subagents with the Agent tool disabled, so they cannot spawn any further nested subagents at all", "correct": false, "explanation": "Disabling the Agent tool prevents a subagent from spawning nested subagents, which is the opposite of what's needed here and doesn't enable direct messaging between agents."}, {"letter": "C", "text": "Agent teams, which support inter-agent messaging and centralized management of multiple coordinating sessions", "correct": true, "explanation": "Agent teams coordinate multiple Claude Code sessions with shared tasks and inter-agent messaging, which fits ongoing back-and-forth coordination better than one-shot subagent delegation."}, {"letter": "D", "text": "A single subagent configured with a longer maxTurns value so it can simulate multiple separate participants at once", "correct": false, "explanation": "A longer maxTurns value only extends how many turns one agent can take; it doesn't give that single agent the ability to have multiple agents message each other directly."}], "correct": "C", "select": 1, "group": "A"}, {"id": "f1-061", "domain": 1, "task_id": "1.7", "objective": "Manage session state, resumption, and forking", "situation": "An architect is scripting a build pipeline where each single prompt is a one-shot task with no follow-up question expected afterward.", "question": "According to session-management guidance, what is the appropriate amount of session handling to add for this task?", "options": [{"letter": "A", "text": "Set continue_conversation=True so a follow-up could be added later without refactoring", "correct": false, "explanation": "The Anthropic Messages API does not have a continue_conversation parameter. Each request is self-contained; to maintain context across calls, you would need to include the full conversation history in subsequent requests rather than relying on a session flag."}, {"letter": "B", "text": "Capture the session ID and pass it to resume immediately after, to be safe", "correct": false, "explanation": "Capturing and immediately resuming a session adds unnecessary overhead for a one-shot task. Session resumption is designed for long-running or multi-turn interactions where context must be preserved. In a build pipeline where each prompt stands alone, this complexity is not justified."}, {"letter": "C", "text": "None beyond a single query() call, since no additional prompts will share context afterward", "correct": true, "explanation": "For one-shot, independent tasks, session management is unnecessary. Anthropic recommends starting a new session for each distinct task to prevent context convolution and potential hallucinations. Because these pipeline prompts do not require shared state, using a single stateless API call without session tracking is the simplest and most appropriate approach."}, {"letter": "D", "text": "Enable fork_session so a parallel branch is available if a follow-up becomes necessary", "correct": false, "explanation": "Forking a session is intended for creating parallel conversation branches when multiple continuations are needed from a shared starting point. Since the requirement is explicitly one-shot with no follow-up, enabling fork_session is irrelevant and would add unneeded complexity."}], "correct": "C", "select": 1, "group": "A"}, {"id": "f1-062", "domain": 1, "task_id": "1.5", "objective": "Apply Agent SDK hooks for tool call interception and data normalization", "situation": "An architect wants a compliance rule enforced so reliably that it must hold even if the model's system prompt is later edited by another team member who is unaware of the refund policy.", "question": "Which design property makes a PreToolUse hook the right mechanism for this, compared to keeping the rule only in the system prompt?", "options": [{"letter": "A", "text": "The hook is registered in application code separate from the prompt, so a prompt edit cannot silently remove the enforcement as it could remove a sentence describing the same rule.", "correct": true, "explanation": "The hook is defined in application code separate from the system prompt, so any subsequent editing or removal of the prompt text does not affect the hook's execution. This structural decoupling ensures the enforcement cannot be silently removed by altering the natural-language instructions."}, {"letter": "B", "text": "The hook is stored in the same configuration file as the system prompt, so any prompt edit is automatically validated against the hook's refund logic before being saved, preventing removal of the enforcement.", "correct": false, "explanation": "There is no automatic validation of prompt edits against hook logic; hooks are independent artifacts that do not intervene in the prompt editing or saving process. The enforcement mechanism relies on intercepting tool calls at runtime, not on checking configuration file changes."}, {"letter": "C", "text": "The hook increases the model's confidence score for refund-related completions by adjusting sampling parameters, making it statistically less likely to produce a noncompliant call even if the prompt is altered.", "correct": false, "explanation": "Hooks have no ability to modify the model's internal confidence scores or sampling parameters. They enforce compliance externally by blocking or modifying the tool call after the model has generated it, not by influencing generation probabilities."}, {"letter": "D", "text": "The hook automatically regenerates the system prompt on every session start from a template that includes the refund policy, so any manual edits are replaced with the compliant version.", "correct": false, "explanation": "Hooks do not regenerate or manage the system prompt; they operate purely at the tool invocation level. They intercept tool calls regardless of the current prompt content, so they cannot replace manual prompt edits with a compliant template."}], "correct": "A", "select": 1, "group": "E"}, {"id": "f1-063", "domain": 1, "task_id": "1.1", "objective": "Design and implement agentic loops for autonomous task execution", "situation": "A customer-support agent's loop exits when the assistant's response text contains the substring \"I have completed the task.\" During testing, the loop sometimes stops one step early and sometimes never stops at all on equally valid completions.", "question": "What is the underlying problem with this natural-language-based termination check?", "options": [{"letter": "A", "text": "Text-based matching requires a separate paid call to an external classification model before every single iteration, which this team has not yet implemented", "correct": false, "explanation": "Substring matching is a local string operation performed by the client's own code and requires no additional API call."}, {"letter": "B", "text": "Checking for that specific substring is disallowed under Anthropic's usage policy, so the API blocks any request that contains it", "correct": false, "explanation": "There is no Anthropic usage policy restricting particular substrings in assistant text, and the API does not block requests on that basis."}, {"letter": "C", "text": "The substring check throws a runtime exception on every call, because assistant text blocks are always stored internally as binary data", "correct": false, "explanation": "Assistant text blocks are returned as strings in the API response; there is no binary storage causing a runtime exception."}, {"letter": "D", "text": "Claude can phrase completion differently across responses or use similar wording while still requesting another tool, making substring checks unreliable", "correct": true, "explanation": "Natural-language phrasing is not a stable signal: wording can vary across otherwise-identical completions, or similar phrases can appear while a tool call is still pending, unlike the structured stop_reason field."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-064", "domain": 1, "task_id": "1.7", "objective": "Manage session state, resumption, and forking", "situation": "An architect resumed a session and asked the agent to revert a file to a state from earlier in that same conversation, expecting the session itself to have preserved the file's old contents. The agent instead reports it can only see the file's current, edited contents on disk.", "question": "What explains this behavior?", "options": [{"letter": "A", "text": "Sessions persist the conversation history, not a snapshot of the filesystem, so file reverts require a separate checkpointing mechanism", "correct": true, "explanation": "Sessions save conversation history, tool calls, and results, but not a filesystem snapshot; reverting actual file contents to an earlier point requires the separate file-checkpointing feature, not session resume alone."}, {"letter": "B", "text": "The agent's read tool caches results only for the current turn, so earlier reads are always inaccessible after resuming", "correct": false, "explanation": "earlier tool results remain visible in the resumed conversation history; the issue is that history doesn't restore the file's actual on-disk bytes, not that earlier reads vanish from context."}, {"letter": "C", "text": "fork_session was required to preserve the file's earlier state, and it was omitted from this particular resume call", "correct": false, "explanation": "fork_session controls whether resuming creates a branch versus continuing in place; it has no effect on whether filesystem contents are snapshotted or restorable."}, {"letter": "D", "text": "The resume call must have used the wrong session ID, since a correct resume would restore the earlier file contents automatically", "correct": false, "explanation": "resuming the correct session ID restores the conversation's history and reasoning, but conversation history was never a mechanism for restoring file contents, regardless of which session ID was used."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f1-065", "domain": 1, "task_id": "1.7", "objective": "Manage session state, resumption, and forking", "situation": "An architect is building a custom session picker for an internal tool and needs a way to enumerate every session on disk for a repository along with reading a particular session's full message history, without shelling out to the interactive CLI picker.", "question": "Which SDK-exposed functions fit this need?", "options": [{"letter": "A", "text": "renameSession() to enumerate sessions and tagSession() to read a session's messages", "correct": false, "explanation": "renameSession() is used to rename a session's file, not to enumerate sessions. There is no tagSession() function in the SDK. The correct functions are listSessions() for enumeration and getSessionMessages() for reading messages."}, {"letter": "B", "text": "listSessions() to enumerate sessions and getSessionMessages() to read a session's messages", "correct": true, "explanation": "Anthropic's Agent SDK (both Python and TypeScript) provides listSessions() (or list_sessions() in Python) to return a list of all sessions on disk sorted by last-modified time, and getSessionMessages() (or get_session_messages()) to retrieve the full message history for a given session. These functions are the recommended approach for building custom session pickers or transcript viewers per official documentation."}, {"letter": "C", "text": "getSessionInfo() to enumerate all sessions and resume() to read a session's messages", "correct": false, "explanation": "getSessionInfo() is not documented as a function to enumerate all sessions; rather, information about a specific session can be obtained from the data returned by listSessions(). resume() is used to continue an existing session, not to read its message history. Use getSessionMessages() to access historical messages."}, {"letter": "D", "text": "fork_session() to enumerate sessions and continue() to read a session's messages", "correct": false, "explanation": "fork_session() is a function to create a new session based on an existing one, not to enumerate sessions. continue() is not the standard way to read message history; the SDK provides getSessionMessages() for that purpose. For enumeration, use listSessions()."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f1-066", "domain": 1, "task_id": "1.4", "objective": "Implement multi-step workflows with enforcement and handoff patterns", "situation": "A team registers three independent PreToolUse hooks for process_refund: one checks identity verification, one checks fraud score, and one logs the attempt for audit. On one ticket, the verification hook returns \"deny\" while the fraud-score hook returns \"allow\" and the logging hook returns an empty object.", "question": "What happens to the tool call?", "options": [{"letter": "A", "text": "The call proceeds, because a majority of the registered hooks either allowed the call or expressed no objection to it", "correct": false, "explanation": "Decisions are not resolved by majority vote; a single deny is sufficient to block the call regardless of how many other hooks allowed it."}, {"letter": "B", "text": "The call is paused and escalated to the user for manual approval, because the hooks produced a mixed set of decisions", "correct": false, "explanation": "Deny outranks ask in the priority order, so a mixed result does not fall back to a manual-approval prompt when one hook already denied it."}, {"letter": "C", "text": "The call proceeds using only the fraud-score hook's decision, because it returned an explicit \"allow\" rather than an empty object", "correct": false, "explanation": "All matching hooks' outputs are considered together; an empty object from one hook does not remove it from the resolution, and deny still overrides allow."}, {"letter": "D", "text": "The call is blocked, because when multiple hooks disagree, a \"deny\" from any hook overrides \"allow\" results from the others", "correct": true, "explanation": "When several hooks fire for the same event, the most restrictive decision wins: deny takes priority over defer, which takes priority over ask, which takes priority over allow."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-067", "domain": 1, "task_id": "1.5", "objective": "Apply Agent SDK hooks for tool call interception and data normalization", "situation": "For refunds between $500 and $1000, policy requires the agent to pause and let a human reviewer approve or reject before the refund proceeds, rather than blocking it outright or letting it run automatically.", "question": "Which PreToolUse hookSpecificOutput configuration matches this requirement?", "options": [{"letter": "A", "text": "permissionDecision set to \"ask\", so the operation is surfaced for approval instead of executing automatically or being silently rejected", "correct": true, "explanation": "According to Anthropic's documentation, setting hookSpecificOutput.permissionDecision to \"ask\" in a PreToolUse hook prompts the user for confirmation before the tool executes. This directly fulfills the requirement to pause and wait for a human reviewer to approve or reject the refund. It is the recommended mechanism for manual approval in Claude Code when an operation must not run automatically or be blocked silently."}, {"letter": "B", "text": "async set to true with asyncTimeout raised to 60000, so the hook has enough time to reach a human reviewer before the call proceeds", "correct": false, "explanation": "Asynchronous hooks (async: true) run in the background and do not pause tool execution for human review. They are designed for non-blocking tasks like logging or notifications, not interactive approval workflows. The asyncTimeout only governs how long the hook itself may run asynchronously, not a wait for human input."}, {"letter": "C", "text": "permissionDecision set to \"deny\", paired with a permissionDecisionReason that instructs the model to contact a human reviewer on its own", "correct": false, "explanation": "Setting permissionDecision to \"deny\" blocks the tool execution outright. While the reason can inform the model, the policy calls for the operation to be surfaced for approval, not rejected. This approach does not guarantee a human review; it only requires the model to potentially try a different action based on the feedback."}, {"letter": "D", "text": "permissionDecision set to \"allow\", combined with an additionalContext note asking the model to mention the amount to the user afterward", "correct": false, "explanation": "permissionDecision: \"allow\" bypasses the permission system entirely and lets the tool run automatically without any pause for human input. The additionalContext field is used to inject extra context into the model, not to trigger a human approval step. This configuration would ignore the required review for refunds in the $500–$1000 range."}], "correct": "A", "select": 1, "group": "E"}, {"id": "f1-068", "domain": 1, "task_id": "1.2", "objective": "Orchestrate multi-agent systems with coordinator-subagent patterns", "situation": "A coordinator has several custom subagents defined, each with a description field summarizing what it's for. When a new query arrives, the coordinator does not explicitly name any subagent.", "question": "How does it typically decide which, if any, subagent to invoke?", "options": [{"letter": "A", "text": "It requires the query to include the subagent's exact name, otherwise no subagent is ever considered for delegation", "correct": false, "explanation": "Explicit naming is one way to guarantee a specific subagent is used, but it is not required; automatic delegation based on description matching works without naming a subagent."}, {"letter": "B", "text": "It always invokes every defined subagent and discards the ones whose output isn't relevant to the query", "correct": false, "explanation": "Invoking every subagent regardless of relevance wastes effort and contradicts the coordinator's role in dynamically selecting which subagents a query actually needs."}, {"letter": "C", "text": "It invokes subagents in the fixed order they were defined in the agents configuration, regardless of the query", "correct": false, "explanation": "Definition order in the agents configuration has no bearing on invocation; matching is based on the description field's relevance to the query, not declaration sequence."}, {"letter": "D", "text": "It matches the query against each subagent's description field and delegates automatically when a match is found", "correct": true, "explanation": "Claude determines whether to invoke a subagent based on how well the query matches that subagent's description field, delegating automatically without requiring an explicit name."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-069", "domain": 1, "task_id": "1.7", "objective": "Manage session state, resumption, and forking", "situation": "A CI worker runs claude --resume <session-id> to continue a session that was created and last active on a different ephemeral build machine. The resume call returns a brand-new session with none of the prior history instead of the expected conversation.", "question": "What is the most likely root cause?", "options": [{"letter": "A", "text": "The session's transcript file only exists on the original machine and was never copied to the new worker", "correct": true, "explanation": "Session transcripts are written to local disk under the working directory's project path; without mirroring that JSONL file to the new machine (or matching the cwd), resume cannot find the prior history and starts fresh."}, {"letter": "B", "text": "The session name was too long for the resume lookup to match it against the stored transcript index", "correct": false, "explanation": "this scenario resumes by session ID, not by name, so session name length is irrelevant to the lookup."}, {"letter": "C", "text": "The session ID was generated with fork_session enabled, which prevents cross-machine resumption entirely", "correct": false, "explanation": "fork_session controls whether a resume produces a branch versus continuing in place; it does not by itself block resumption across machines."}, {"letter": "D", "text": "CI workers are restricted from resuming any session that was started interactively via the CLI", "correct": false, "explanation": "there is no such restriction tied to how a session was originally started; the failure mode described is about locating the transcript file, not the session's origin."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f1-070", "domain": 1, "task_id": "1.6", "objective": "Design task decomposition strategies for complex workflows", "situation": "An architect is asked to design an agent that investigates why a production incident occurred, given only a vague alert message and no prior knowledge of which service is at fault. The number of logs, services, and code paths to inspect cannot be known ahead of time.", "question": "Which decomposition approach is most appropriate?", "options": [{"letter": "A", "text": "A fixed sequential chain that always inspects the database, then the cache layer, then the load balancer, in that fixed order", "correct": false, "explanation": "A hardcoded order of components to check assumes the fault location in advance, which contradicts the premise that the responsible service is unknown and could waste steps checking irrelevant systems."}, {"letter": "B", "text": "A single prompt that asks the model to name the root cause immediately from only the wording of the alert message", "correct": false, "explanation": "A single-shot guess without any investigation into logs, services, or code paths is likely to be wrong and does not use the available tools to gather evidence before concluding."}, {"letter": "C", "text": "A dynamic orchestrator that generates and prioritizes new investigation subtasks based on what each prior step uncovers", "correct": true, "explanation": "Because the scope and required steps are unknown in advance and depend on intermediate findings, an adaptive, dynamically decomposed investigation plan is the right pattern rather than a fixed pipeline."}, {"letter": "D", "text": "A prompt chain with a hardcoded set of five investigation steps that always runs in full regardless of what is found", "correct": false, "explanation": "Fixing the number of steps in advance ignores that some incidents resolve in fewer steps and others need many more, defeating the purpose of adapting to discovered evidence."}], "correct": "C", "select": 1, "group": "A"}, {"id": "f1-071", "domain": 1, "task_id": "1.7", "objective": "Manage session state, resumption, and forking", "situation": "A team has finished a shared analysis of a monolith's test suite in one session and now wants to compare two independent refactoring strategies (extract-service vs. strangler-fig) starting from that same analyzed baseline, without letting either exploration corrupt the other or the original session.", "question": "What is the most appropriate mechanism?", "options": [{"letter": "A", "text": "Start two brand-new sessions and describe the test-suite findings from memory in each opening prompt.", "correct": false, "explanation": "This discards the preserved analysis session and relies on manual restatement, which is error-prone and may omit important details. It does not branch from the original context and creates inconsistent starting points for the two strategies."}, {"letter": "B", "text": "Resume the analysis session twice with --fork-session set, producing two independent branches from the shared baseline.", "correct": true, "explanation": "Resuming with --fork-session clones the selected conversation into a new session, so each exploration gets an isolated branch from the same baseline and the original session remains unchanged. Running this twice gives two independent paths that cannot contaminate one another. This is the intended way to compare alternative approaches from a shared checkpoint."}, {"letter": "C", "text": "Resume the analysis session once and alternate prompts between the two strategies within that single conversation.", "correct": false, "explanation": "In a single resumed conversation, both strategies share the same context and history, so prompts from one path can influence the other and alter the original session. It does not provide the isolation or independent branching required for a fair comparison."}, {"letter": "D", "text": "Resume the analysis session, complete the first strategy, then use /clear and resume again for the second.", "correct": false, "explanation": "Using /clear wipes the conversation history, so the second strategy loses the shared baseline and the first exploration may alter the original session. It does not create isolated branches from the same starting state and can corrupt the original analysis."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f1-072", "domain": 1, "task_id": "1.2", "objective": "Orchestrate multi-agent systems with coordinator-subagent patterns", "situation": "Two sub-agents independently research the release year of a major historical event: one reports 1969, the other reports 1970, citing different sources.", "question": "Under the orchestrator-subagent pattern, how should the orchestrator handle this discrepancy before presenting a final answer?", "options": [{"letter": "A", "text": "The conflicting results should be discarded to maintain output consistency, and the orchestrator should provide no answer.", "correct": false, "explanation": "Discarding conflicting results hides important uncertainty and removes source provenance. The orchestrator should not silently drop valid subagent output; it should preserve the discrepancy and present the relevant references."}, {"letter": "B", "text": "The orchestrator aggregates the findings and presents the conflicting reports to the end user with references to the different sources, rather than unilaterally resolving the discrepancy.", "correct": true, "explanation": "Anthropic describes the orchestrator as synthesizing subagent findings and presenting a final answer with citations. In cases of conflicting evidence that cannot be reconciled, the orchestrator should preserve both findings and present them to the user with references to the different sources, allowing human judgment rather than forcing an unsupported resolution."}, {"letter": "C", "text": "Whichever sub-agent returned its result first takes precedence, since first-to-respond is authoritative in a hub-and-spoke architecture.", "correct": false, "explanation": "The orchestrator-subagent pattern does not assign authority based on response order. Subagents often run in parallel, and the lead agent must aggregate all returned findings; arrival order does not determine correctness."}, {"letter": "D", "text": "The end user handles the discrepancy because sub-agents are not permitted to report ambiguous or conflicting findings to the orchestrator.", "correct": false, "explanation": "Subagents are expected to return their findings to the orchestrator, including conflicts. The orchestrator is responsible for aggregating and, when necessary, surfacing those conflicts to the user with references; human intervention is a fallback for complex or unresolvable discrepancies."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f1-073", "domain": 1, "task_id": "1.1", "objective": "Design and implement agentic loops for autonomous task execution", "situation": "An architect is deciding between two designs for an incident-response agent: Design X lets Claude choose which diagnostic tool to call next based on the evolving conversation, while Design Y hardcodes a fixed sequence (check logs, then check metrics, then restart service) regardless of what earlier tool results reveal.", "question": "Which statement correctly characterizes the tradeoff?", "options": [{"letter": "A", "text": "Design X lets the model adapt its next action to intermediate findings, while Design Y is a fixed tree that ignores what earlier results show", "correct": true, "explanation": "Model-driven decision-making lets Claude reason over current context to pick the next tool, while a fixed sequence executes the same steps regardless of what the logs or metrics actually showed."}, {"letter": "B", "text": "Design X requires disabling stop_reason inspection entirely, while Design Y depends on stop_reason to pick the next hardcoded step in its sequence", "correct": false, "explanation": "Both designs still rely on stop_reason to detect a tool call versus a final answer; Design Y's hardcoding concerns which tool runs next, not stop_reason usage."}, {"letter": "C", "text": "Design X necessarily issues more tool_use blocks per request, because every model-driven loop always batches all available tools in one turn", "correct": false, "explanation": "Model-driven loops do not necessarily request more tools per turn; Claude can request a single tool at a time just as easily as a scripted design."}, {"letter": "D", "text": "Design X and Design Y always produce identical conversation histories, differing only in which tool names appear in the system prompt text", "correct": false, "explanation": "The histories differ meaningfully: Design X reflects Claude's own tool choices, while Design Y reflects a scripted sequence the model did not choose."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f1-074", "domain": 1, "task_id": "1.6", "objective": "Design task decomposition strategies for complex workflows", "situation": "A team is building an automated pipeline that takes a raw customer support transcript, produces a structured summary, then translates that summary into three fixed target languages. The steps and their order never change between runs.", "question": "Which decomposition strategy best fits this workflow?", "options": [{"letter": "A", "text": "An adaptive investigation plan that generates new subtasks based on which language the transcript happens to mention", "correct": false, "explanation": "Adaptive subtask generation is suited to open-ended investigation, not a predictable two-step transform where the steps never change based on content."}, {"letter": "B", "text": "A fixed prompt-chaining pipeline with a programmatic check after the summarization step before translation begins", "correct": true, "explanation": "The task decomposes cleanly into fixed, predictable subtasks (summarize, then translate), which is exactly the case where prompt chaining with a checkpoint between steps is recommended over a dynamic agent."}, {"letter": "C", "text": "A single subagent that reads the transcript once and produces all three translations without an intermediate summary artifact", "correct": false, "explanation": "Skipping the intermediate summary artifact removes the checkpoint that lets the pipeline verify quality before translation, and merges two distinct concerns into one pass."}, {"letter": "D", "text": "A dynamic orchestrator that decides at runtime whether summarization or translation should happen first", "correct": false, "explanation": "The order of steps is fixed and known in advance, so introducing runtime decision-making about step order adds unnecessary complexity for no benefit."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f1-075", "domain": 1, "task_id": "1.4", "objective": "Implement multi-step workflows with enforcement and handoff patterns", "situation": "A support agent has two tools, get_customer and process_refund. The team's system prompt says \"Always call get_customer and confirm the identity before calling process_refund.\" During testing, the agent occasionally calls process_refund first when a ticket is phrased as an urgent complaint. The architect wants this ordering to hold every time, not just most of the time.", "question": "What should they implement?", "options": [{"letter": "A", "text": "A PreToolUse hook matched to process_refund that checks for a verified customer ID in session state and returns permissionDecision \"deny\" when none exists", "correct": true, "explanation": "A PreToolUse hook is programmatic enforcement: it inspects state before the tool executes and can deny the call outright, so the ordering holds regardless of how the prompt is phrased."}, {"letter": "B", "text": "A follow-up user-facing reminder message injected after every ticket that restates the required call order before the agent responds", "correct": false, "explanation": "An injected reminder is still a prompt-based nudge the model can override under unusual phrasing; it is not a gate on the tool call itself."}, {"letter": "C", "text": "A rewritten system prompt that repeats the ordering requirement twice and adds the word \"must\" in place of \"always\" to strengthen the instruction", "correct": false, "explanation": "Rewording the same prompt-based instruction still leaves compliance dependent on the model following it; it does not close the non-zero failure rate the architect is trying to eliminate."}, {"letter": "D", "text": "A larger context window for the agent so it retains the ordering instruction even when the ticket phrasing shifts partway through the conversation", "correct": false, "explanation": "A larger context window helps retention of information but does not create a hard block on calling process_refund out of order."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f1-076", "domain": 1, "task_id": "1.7", "objective": "Manage session state, resumption, and forking", "situation": "During an interactive session investigating a caching bug, the architect wants to explore an alternate hypothesis without losing the current line of reasoning, and wants to keep working in the terminal rather than scripting against the SDK.", "question": "Which in-session action creates a divergent copy of the conversation while leaving the current one intact?", "options": [{"letter": "A", "text": "Run /resume and select the same session again from the picker to duplicate it in place", "correct": false, "explanation": "/resume reopens a previously saved session, but it does not duplicate it. Selecting the same session simply continues that session; it does not create a parallel copy while leaving the current session open."}, {"letter": "B", "text": "Run /clear to empty context and begin reasoning about the alternate hypothesis immediately", "correct": false, "explanation": "/clear wipes all conversation history and context, causing you to lose the current line of reasoning entirely. It does not retain the original session for simultaneous or later reference."}, {"letter": "C", "text": "Run /compact to replace the history with a summary focused on the alternate hypothesis", "correct": false, "explanation": "/compact condenses the existing conversation into a summary to manage context length, but it alters the current session's history rather than creating an independent, divergent copy. It does not preserve the original session intact for later return."}, {"letter": "D", "text": "Run /fork with an optional name to switch into a copy of the conversation so far", "correct": true, "explanation": "In Claude Code, the /fork command creates a new session that inherits the entire conversation context up to the current point. This allows you to explore an alternate hypothesis in a separate session while preserving the original session intact, all within the terminal."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-077", "domain": 1, "task_id": "1.1", "objective": "Design and implement agentic loops for autonomous task execution", "situation": "A support-ticket agent receives a simple, fully answerable question. On the very first request, Claude's response has stop_reason \"end_turn\" and contains only a text answer, with no tool_use block at all. An engineer flags this as a bug, assuming every task must go through at least one tool call before the loop can end.", "question": "Is this assumption correct?", "options": [{"letter": "A", "text": "Yes, because the API rejects any first response carrying stop_reason \"end_turn\" and requires the client to resend the request", "correct": false, "explanation": "The API does not reject first responses carrying stop_reason \"end_turn\"; ending on the very first turn is an entirely normal outcome."}, {"letter": "B", "text": "Yes, the loop must always force at least one tool_use round trip before accepting stop_reason \"end_turn\" as genuine completion", "correct": false, "explanation": "Forcing an unnecessary tool call before accepting completion is itself an anti-pattern; not every task needs a tool, and the loop should trust stop_reason as given."}, {"letter": "C", "text": "No, but only because the Messages API silently inserts a placeholder tool_use block whenever no real tool call was necessary", "correct": false, "explanation": "The API does not insert placeholder tool_use blocks; a response with no tool_use block simply means no tool was requested."}, {"letter": "D", "text": "No, stop_reason \"end_turn\" on the first response is a valid immediate completion whenever Claude can answer without needing any tool", "correct": true, "explanation": "There is no requirement that a task pass through a tool call; if Claude can answer directly, stop_reason \"end_turn\" on the very first response is a legitimate, complete result."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-078", "domain": 1, "task_id": "1.5", "objective": "Apply Agent SDK hooks for tool call interception and data normalization", "situation": "An agent's inventory_lookup MCP tool returns stock levels as a numeric status code (0, 1, 2) meaning in-stock, low-stock, and out-of-stock respectively, while a separate warehouse_lookup tool returns the same concept as plain strings. The architect wants the model to reason over one consistent vocabulary for stock status regardless of which tool answered.", "question": "Which hook change achieves this with a deterministic guarantee?", "options": [{"letter": "A", "text": "A PreToolUse hook matched to both tools that rewrites tool_input so both tools receive identical request parameters before they execute", "correct": false, "explanation": "Rewriting tool_input changes the outgoing request, not the returned stock-status value; the inconsistency exists in the response payload, which a PreToolUse hook does not see."}, {"letter": "B", "text": "A PostToolUse hook matched to both tools that maps each tool's raw response onto the same set of string labels and returns it via updatedToolOutput", "correct": true, "explanation": "PostToolUse sees each tool's actual response and can deterministically rewrite it into a shared vocabulary via updatedToolOutput before the model ever reads the discrepancy, regardless of which tool answered."}, {"letter": "C", "text": "A SessionStart hook that documents the numeric-to-string mapping once in a system message shown to the user at the beginning of the session", "correct": false, "explanation": "A SessionStart message is shown to the user, not injected as a guaranteed transformation of tool output, and does not run per tool call, so later responses in the session remain unnormalized."}, {"letter": "D", "text": "A UserPromptSubmit hook that reminds the model at the start of every turn to translate numeric status codes into the equivalent string labels itself", "correct": false, "explanation": "Asking the model to translate the codes itself each turn is a prompt-based, probabilistic approach; the model can still misinterpret or forget the mapping, which is exactly the reliability problem normalization hooks are meant to solve."}], "correct": "B", "select": 1, "group": "E"}, {"id": "f1-079", "domain": 1, "task_id": "1.1", "objective": "Design and implement agentic loops for autonomous task execution", "situation": "An architect is building a client-side agent loop against the Messages API for a data-cleanup task. The first response has stop_reason set to \"tool_use\" and contains a tool_use block requesting a file-listing tool.", "question": "What should the loop do next to correctly continue the agentic execution?", "options": [{"letter": "A", "text": "Execute the requested tool, append a tool_result block tagged with the tool_use_id, and send the full updated conversation back to Claude", "correct": true, "explanation": "The client executes the tool and returns a tool_result block tied to the tool_use_id in the next request; this is how the agentic loop lifecycle continues."}, {"letter": "B", "text": "Run the tool locally, write its output only to an application log file, and resend the prior conversation exactly as it stood before", "correct": false, "explanation": "If the tool output never reaches the conversation as a tool_result, Claude has no way to see it and cannot reason about the next action."}, {"letter": "C", "text": "Hold off on running the requested tool until Claude produces assistant text describing exactly what output it expects the tool call to return", "correct": false, "explanation": "stop_reason \"tool_use\" is itself the signal to execute immediately; Claude does not need to narrate expected output before the tool runs."}, {"letter": "D", "text": "Append only the tool_use block to history and send a new request with just the original prompt, dropping the tool result needed to interpret it", "correct": false, "explanation": "Dropping later context and resending only the original prompt discards the pending tool call and any prior reasoning, breaking the model's view of what it already asked for."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f1-080", "domain": 1, "task_id": "1.3", "objective": "Configure subagent invocation, context passing, and spawning", "situation": "A coordinator spawns a \"codebase-explorer\" subagent that reads dozens of files while investigating a bug, then reports back a three-paragraph summary. The team is happy that the main conversation's context stayed small despite the exploration.", "question": "What underlying behavior explains why the parent conversation did not grow with every file the subagent read?", "options": [{"letter": "A", "text": "The subagent's intermediate tool calls and results stay inside its own isolated context; only its final message is returned to the parent as the Task tool's result", "correct": true, "explanation": "Context isolation means a subagent's intermediate tool calls and results remain inside its own run; the parent only receives the subagent's final message as the result of the invocation, keeping the main conversation compact."}, {"letter": "B", "text": "The coordinator automatically deletes the subagent's transcript from disk as soon as the Task call completes, so nothing persists to bloat future context", "correct": false, "explanation": "Subagent transcripts are persisted to disk (subject to cleanup settings) rather than deleted immediately; the parent's context stays small because of isolation, not deletion."}, {"letter": "C", "text": "The parent's context window is expanded dynamically to absorb subagent activity, so the growth is real but simply not visible in the rendered conversation", "correct": false, "explanation": "The parent's context window does not silently expand; the reduction in visible growth reflects genuine isolation of the subagent's intermediate work, not hidden growth."}, {"letter": "D", "text": "The subagent compresses every file it reads into a one-line hash before responding, and the parent only ever sees these hashes rather than real content", "correct": false, "explanation": "There is no hashing or compression mechanism for file contents; the subagent's tool results are simply not forwarded to the parent's context at all."}], "correct": "A", "select": 1, "group": "B"}, {"id": "f1-081", "domain": 1, "task_id": "1.6", "objective": "Design task decomposition strategies for complex workflows", "situation": "A reviewer is designing an agentic code review for a pull request that touches 40 files across a monorepo. Reviewers have noticed that when a single pass tries to hold all 40 files in context at once, subtle cross-file issues are missed and comments become generic.", "question": "What restructuring addresses this attention dilution problem?", "options": [{"letter": "A", "text": "Raise the review prompt's sampling temperature so the single combined pass considers a wider range of possible issues across all the files", "correct": false, "explanation": "Adjusting sampling temperature does not address the underlying problem of too much context competing for attention in a single pass; it changes output variability, not scope management."}, {"letter": "B", "text": "Run a per-file local analysis pass on each file independently, then run a separate cross-file integration pass over the per-file findings", "correct": true, "explanation": "Splitting the review into focused per-file passes avoids attention dilution, and a dedicated cross-file pass afterward catches integration issues that per-file review alone would miss, matching the recommended pattern for large reviews."}, {"letter": "C", "text": "Keep the single combined pass but instruct the model to prioritize whichever files appear first in the diff ordering", "correct": false, "explanation": "Prioritizing by diff order in a single combined pass still holds all 40 files in context at once and does not solve the dilution problem, it only biases which files get more attention."}, {"letter": "D", "text": "Randomize the file order within the single combined pass before each run so different files receive more attention each time", "correct": false, "explanation": "Randomizing file order across runs does not reduce the total context held in any single pass and produces inconsistent coverage rather than systematically avoiding dilution."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f1-082", "domain": 1, "task_id": "1.3", "objective": "Configure subagent invocation, context passing, and spawning", "situation": "A coordinator runs a \"web-search\" subagent and a \"document-analysis\" subagent, both of which return findings that must later be cited with their original sources when a \"synthesis\" subagent writes the final report. Reviewers keep finding that the final report attributes a claim to the wrong source URL or page number.", "question": "What is the best way to pass the prior agents' findings into the synthesis subagent's prompt to prevent this?", "options": [{"letter": "A", "text": "Pass only a short summary of each prior agent's findings and let synthesis re-fetch every original source itself to confirm the exact URL and page number for each claim", "correct": false, "explanation": "Re-fetching sources is redundant, slower, and does not guarantee the re-fetched content maps back to the same claims already gathered, and it discards the point of reusing completed research."}, {"letter": "B", "text": "Pass the findings exactly as free-form conversational text copied from the subagents' chat-style replies, since exact wording matters more than structure for citation accuracy", "correct": false, "explanation": "Free-form conversational text without explicit metadata fields still forces synthesis to guess at attribution rather than reading it directly from structured fields."}, {"letter": "C", "text": "Pass the findings as one continuous block of prose combining every source's text together, trusting synthesis to keep the origin of each sentence straight from context", "correct": false, "explanation": "Merging everything into undifferentiated prose is exactly what causes mis-attribution, since there is no reliable boundary between where one source's content ends and another's begins."}, {"letter": "D", "text": "Pass the findings as structured entries that separate content from metadata, such as source URL, document name, and page number, so synthesis preserves correct attribution", "correct": true, "explanation": "Using structured data that keeps each finding's content distinct from its source metadata lets the synthesis subagent map claims back to the correct URL, document, or page number reliably, rather than relying on the model to infer attribution from unstructured prose."}], "correct": "D", "select": 1, "group": "B"}, {"id": "f1-083", "domain": 1, "task_id": "1.7", "objective": "Manage session state, resumption, and forking", "situation": "A prior session spent many turns reading and cross-referencing a large data-pipeline codebase, and most of its tool results are now stale because the pipeline was heavily refactored afterward. The architect has already captured the session's final conclusions in a structured summary (such as progress notes or a summary document) and only needs those conclusions—not the intermediate tool calls—to continue the next phase of work.", "question": "Which continuation strategy is more reliable here?", "options": [{"letter": "A", "text": "Fork the prior session and continue exploration straight from its unmodified stale history", "correct": false, "explanation": "Forking creates a continuation from the same stale history, so the agent would still rely on outdated tool outputs and intermediate context. A fresh context initialized with the already-prepared structured summary is more reliable and avoids the stale intermediate calls."}, {"letter": "B", "text": "Resume the prior session and issue /clear right after resuming to reset its context window", "correct": false, "explanation": "Issuing /clear immediately after resuming would erase the session context, including the final conclusions, unless those conclusions were first saved externally. The stated approach does not inject the prepared summary, so it does not reliably preserve the required continuation context."}, {"letter": "C", "text": "Resume the prior session so the agent inherits every cached tool result automatically as-is", "correct": false, "explanation": "Resuming the prior session loads stale cached tool results from before the refactor, which no longer reflect the current pipeline. The architect only needs the final conclusions, not the intermediate tool calls, so inheriting them as-is is unreliable and may mislead the next phase."}, {"letter": "D", "text": "Start a fresh session and inject the already-prepared structured summary of the prior conclusions as the opening prompt", "correct": true, "explanation": "Anthropic recommends compaction: take a conversation nearing the context window limit, summarize its contents, and reinitiate a new context window with that summary to improve long-term coherence. For multi-session work, leaving clear artifacts (like a structured summary or progress notes) lets the next session read only the needed context. Because the stem states the summary of final conclusions already exists, starting fresh with that summary avoids inheriting stale tool results while preserving the needed conclusions."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-084", "domain": 1, "task_id": "1.5", "objective": "Apply Agent SDK hooks for tool call interception and data normalization", "situation": "Two different PostToolUse hooks are registered for the same query_database tool: one truncates overly long result sets to a fixed row limit, and one converts embedded timestamps into ISO 8601. Both need to apply to the same tool response in sequence for the final output the model sees to be both trimmed and normalized.", "question": "What should the architect verify about how these hooks combine?", "options": [{"letter": "A", "text": "Whether the hooks run in a way that lets the second hook operate on the first hook's updatedToolOutput, since if both run independently against the original response only one transformation may be applied.", "correct": true, "explanation": "Multiple PostToolUse hooks for the same tool may run independently, each receiving the original tool response by default, without automatically chaining their updates. The architect must verify whether the SDK ensures that the second hook's input includes the first hook's updatedToolOutput, otherwise only one transformation might be applied."}, {"letter": "B", "text": "Whether the SDK executes PostToolUse hooks in alphabetical order by hook name, because the hook registration system sorts callbacks by name to ensure deterministic processing, and reversing names could swap the truncation and normalization steps.", "correct": false, "explanation": "The Anthropic SDK does not guarantee alphabetical execution order for PostToolUse hooks; they may run concurrently or in registration order, not sorted by name. Therefore, alphabetical ordering cannot be relied upon to ensure the truncation step runs before normalization."}, {"letter": "C", "text": "Whether the hooks are declared using the same HookMatcher timeout value, because the SDK uses timeouts to resolve conflicts when multiple hooks register for the same event, and a mismatch causes one hook's output to be ignored if it finishes later.", "correct": false, "explanation": "The timeout value controls how long the SDK waits for an individual hook to return; it is not used to resolve conflicts or order multiple hooks. A mismatch in timeouts may cause a hook to time out if it takes longer than allowed, but does not suppress its output based on finishing order relative to another hook."}, {"letter": "D", "text": "Whether both hooks share the same tool_use_id, because the SDK requires hooks to have identical matchers in order to compose their transformations on the same response, and mismatched IDs would cause the second hook to ignore the first's output.", "correct": false, "explanation": "The tool_use_id correlates a specific tool call across PreToolUse and PostToolUse events but does not determine how multiple hooks compose their outputs. Hooks do not require matching matchers based on tool_use_id to combine transformations, and mismatched IDs would not cause the second hook to ignore the first's output."}], "correct": "A", "select": 1, "group": "E"}, {"id": "f1-085", "domain": 1, "task_id": "1.7", "objective": "Manage session state, resumption, and forking", "situation": "An architect wants to send a single follow-up question, 'summarize what we changed,' to an already-completed non-interactive session and capture the answer as structured data for a script to parse, without opening the interactive terminal UI.", "question": "Which invocation fits this need?", "options": [{"letter": "A", "text": "claude --resume <session-id> and then type the follow-up question at the interactive prompt", "correct": false, "explanation": "this opens the interactive terminal UI rather than a scriptable, non-interactive invocation, and produces no structured JSON output for a script to consume."}, {"letter": "B", "text": "claude --continue --fork-session \"summarize what we changed\" piped into a JSON parser", "correct": false, "explanation": "--continue targets the most recent session rather than a specific captured one, and forking here creates an unnecessary branch instead of simply asking the existing session a follow-up."}, {"letter": "C", "text": "claude --from-pr <number> \"summarize what we changed\" with no output format specified", "correct": false, "explanation": "--from-pr resumes a session tied to a pull request rather than a specific captured session ID, and omitting an output format leaves the response unstructured for scripting."}, {"letter": "D", "text": "claude -p --resume <session-id> --output-format json \"summarize what we changed\"", "correct": true, "explanation": "claude -p --resume <session-id> sends a follow-up prompt to an existing session non-interactively, and --output-format json produces structured output a script can parse directly."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-086", "domain": 1, "task_id": "1.4", "objective": "Implement multi-step workflows with enforcement and handoff patterns", "situation": "A developer implements a PreToolUse hook that gates the process_refund tool by checking a boolean flag is_verified. The flag is expected to be set to true by a separate mark_verified tool after a human reviewer approves a photo ID. In an incident, the mark_verified tool executed and the human reviewer explicitly rejected the ID, but due to a software bug the is_verified flag was incorrectly set to true. The PreToolUse hook consequently allowed process_refund, resulting in an unauthorized refund.", "question": "What change to the PreToolUse hook would best prevent this category of failure?", "options": [{"letter": "A", "text": "Add a PostToolUse hook on process_refund that verifies the flag again after the refund has been initiated, so it can reverse the transaction if the flag is invalid.", "correct": false, "explanation": "PostToolUse hooks fire after tool execution, by which time the model has already received and processed the tool’s response. This approach reverses the unauthorized action rather than preventing it, adding complexity without true prevention."}, {"letter": "B", "text": "Replace the model with a larger, more capable language model that can independently re-read the entire conversation and determine whether the human review actually succeeded, overriding the flag when necessary.", "correct": false, "explanation": "Agent safety relies on deterministic architectural controls rather than probabilistic model reasoning. A larger model may still misinterpret context and cannot provide the same level of guarantee as a structured parameter check in a hook."}, {"letter": "C", "text": "Modify the PreToolUse hook to inspect the explicit verification result included in the process_refund tool call parameters, confirming that the human review expressly passed, instead of relying on a separate boolean flag that can be set incorrectly.", "correct": true, "explanation": "Embedding the verification outcome as a parameter to process_refund allows the PreToolUse hook (which has access to tool inputs) to deterministically confirm the actual review result. This follows Anthropic’s recommended practice of using explicit, self-contained tool interactions to avoid reliance on decoupled, mutable global state that can be corrupted."}, {"letter": "D", "text": "Increase the timeout on the PreToolUse hook to re-check the flag periodically; the hook will eventually notice the review was incorrect and block the refund.", "correct": false, "explanation": "PreToolUse hooks execute once before tool invocation and are not designed for periodic re-checks; the flag value would remain unchanged, so this does not prevent the failure."}], "correct": "C", "select": 1, "group": "A"}, {"id": "f1-087", "domain": 1, "task_id": "1.1", "objective": "Design and implement agentic loops for autonomous task execution", "situation": "A monitoring agent's loop executes a requested metrics-fetch tool and prepares the next request. The engineer building the payload includes the new tool_result but replaces the entire prior messages array with just that single tool_result, instead of appending it to the existing history.", "question": "What will most likely go wrong when this request is sent?", "options": [{"letter": "A", "text": "The Messages API rejects the request with an authentication error, since tool_result blocks require a session token from the first call", "correct": false, "explanation": "Authentication is handled through API keys, not session tokens tied to message history; omitting prior messages does not trigger an auth failure."}, {"letter": "B", "text": "Claude loses the original prompt and the reasoning that led to the tool call, so it cannot correctly interpret what the result is answering", "correct": true, "explanation": "Each Messages API call must include the full relevant history; replacing it with only the tool_result strips the original prompt and prior turns, leaving Claude without context for the result."}, {"letter": "C", "text": "Claude automatically re-fetches the full prior conversation from server-side storage, so the replaced array has no practical effect", "correct": false, "explanation": "The Messages API does not persist and auto-restore conversation history by default; the client must supply the full message history on every request."}, {"letter": "D", "text": "The tool_result is silently converted into a system prompt, permanently altering Claude's instructions for every later turn", "correct": false, "explanation": "tool_result blocks are user-turn content and are never automatically promoted into system-prompt instructions."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f1-088", "domain": 1, "task_id": "1.1", "objective": "Design and implement agentic loops for autonomous task execution", "situation": "A platform team is building a shared agentic-loop library for several internal agents.", "question": "One engineer proposes checking response.stop_reason in (\"tool_use\", \"end_turn\") and treating both values as \"continue the loop.\" What is wrong with treating \"end_turn\" as a continue condition in this shared loop?", "options": [{"letter": "A", "text": "\"end_turn\" marks a billing boundary in the API, so continuing past it causes duplicate charges for tokens that were already generated", "correct": false, "explanation": "stop_reason values are not tied to billing boundaries; billing follows token usage reported per request, independent of which stop_reason came back."}, {"letter": "B", "text": "\"end_turn\" is only ever returned on the very first request in a conversation, so treating it as continue causes an infinite loop from turn two onward", "correct": false, "explanation": "\"end_turn\" can occur on any request where Claude produces a final answer, not just the first one in a conversation."}, {"letter": "C", "text": "\"end_turn\" means Claude produced a final response with no further tool request, so continuing on it sends requests after the task is already done", "correct": true, "explanation": "\"end_turn\" signals Claude has finished and is not requesting a tool; looping on it keeps sending requests after the model already considers the task complete."}, {"letter": "D", "text": "\"end_turn\" and \"tool_use\" cannot both be valid stop_reason values for the same account, so the proposed condition would always evaluate to false", "correct": false, "explanation": "Both are simply possible values of the same field across different responses; the proposed check would still run, just incorrectly continuing the loop rather than failing outright."}], "correct": "C", "select": 1, "group": "A"}, {"id": "f1-089", "domain": 1, "task_id": "1.7", "objective": "Manage session state, resumption, and forking", "situation": "A Claude Code session from three days ago read and cached the contents of a configuration file (e.g., a CLAUDE.md or agent definition). Since then, another engineer has substantially rewritten that file in a separate branch that was just merged. The architect needs to continue work while accounting for the file rewrite and preserving the accumulated reasoning about the surrounding system.", "question": "What is the best approach?", "options": [{"letter": "A", "text": "Resume the session and run /compact immediately before asking any follow-up question about the file.", "correct": false, "explanation": "Running /compact does reload the project's configuration file from disk, so it would pick up the rewrite. But compaction achieves that by replacing the conversation with a structured summary that keeps only the request, key technical concepts, and decisions — the underlying tool outputs and intermediate reasoning behind them do not survive. Resuming and asking the agent to re-read the specific file that changed gets the same fresh copy without trading away any of that detail."}, {"letter": "B", "text": "Resume the session and explicitly tell the agent the configuration file changed, prompting it to re-read that file.", "correct": true, "explanation": "Resuming restores the full prior conversation exactly as it was — every file already read, every analysis already performed, every decision already made — so telling the agent the configuration file changed and prompting it to re-read that file fixes the one thing that has gone stale without discarding anything else. The agent's own Read tool call for that file pulls in the merged, current version, and none of the accumulated reasoning is lost in the process."}, {"letter": "C", "text": "Start a new session and provide it with a concise summary of the previous session’s findings and reasoning.", "correct": false, "explanation": "Starting a new session does load the current configuration file fresh from disk, so the rewrite would be picked up. But a new session otherwise starts from nothing, so the only reasoning that survives is whatever the architect manages to compress into a written summary — every other detail of the prior analysis is gone. Resuming and simply asking the agent to re-read the changed file achieves the same fresh read without that loss."}, {"letter": "D", "text": "Resume the session and trust its cached understanding of the configuration file since resumption restores full context.", "correct": false, "explanation": "Resuming restores the conversation as it was recorded, including whichever version of the configuration file the session read into it three days ago; it does not automatically detect or re-check that an external file changed on disk afterward. Trusting the cached understanding means the architect keeps reasoning from the pre-rewrite version of the file, which is exactly the failure this scenario describes."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f1-090", "domain": 1, "task_id": "1.7", "objective": "Manage session state, resumption, and forking", "situation": "A team runs multiple ad-hoc investigation sessions per week and wants every session to be easy to locate and resume using a human-readable handle, especially when several tasks run in parallel on the same day.", "question": "What is the most direct way to make a session resumable by a memorable handle?", "options": [{"letter": "A", "text": "Give the session a descriptive name at startup or via /rename so it can later be resumed with claude --resume <name>", "correct": true, "explanation": "According to Anthropic's official Claude Code CLI documentation, sessions can be named at startup using the --name flag or renamed during a session with the /rename slash command. These named sessions can then be resumed by name with claude --resume <name>. This provides a human-readable, memorable handle. If the exact name is forgotten, claude --resume (without arguments) launches an interactive session picker that lists recent sessions with summaries, allowing easy selection. Using descriptive, consistent naming remains the most direct way to create a handle that is both memorable and resumable."}, {"letter": "B", "text": "Note the raw session ID in a spreadsheet and resume with claude --resume <session-id> each time it's needed", "correct": false, "explanation": "Although resuming by session ID is technically supported, this method relies on manually tracking machine‑generated IDs that are not human‑readable or memorable. It increases the risk of error and does not align with the goal of a memorable handle."}, {"letter": "C", "text": "Keep every session running continuously so that it never needs to be resumed by name at a later point", "correct": false, "explanation": "This approach is impractical and wastes resources. Session persistence is designed to let you save, close, and later resume work, making continuous execution unnecessary."}, {"letter": "D", "text": "Rely on the default auto-generated display name that combines the directory name with a random suffix", "correct": false, "explanation": "Claude Code may auto‑generate a display title from the first prompt, but this title is for display only and cannot be used as a resume handle. Sessions must be explicitly named to be resumed by a specific name."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f1-091", "domain": 1, "task_id": "1.4", "objective": "Implement multi-step workflows with enforcement and handoff patterns", "situation": "An architect is drafting the handoff protocol for cases where an agent must escalate mid-process to a human supervisor who cannot see the conversation. The draft template currently has one field: a free-text \"notes\" box the agent fills in however it sees fit.", "question": "What is the strongest improvement to make the handoffs reliably useful?", "options": [{"letter": "A", "text": "Replace the free-text field with required structured fields for customer details, root cause analysis, and a recommended action, so every handoff contains the same essentials", "correct": true, "explanation": "A structured handoff with defined fields for customer details, root cause, and recommended action guarantees the human agent gets consistent, complete information every time, rather than depending on what an unstructured note happens to include."}, {"letter": "B", "text": "Keep the single free-text field but instruct the agent, via the system prompt, to always remember to mention the customer's name somewhere within its written notes", "correct": false, "explanation": "This is still a prompt-based nudge on an unstructured field; it does not guarantee root cause analysis or a recommended action are included, only that a name might appear somewhere."}, {"letter": "C", "text": "Remove the notes field entirely and have the human supervisor call the customer back so they can re-explain the entire situation again from the beginning", "correct": false, "explanation": "Requiring the customer to re-explain everything defeats the purpose of an agent-compiled handoff and creates a worse customer experience than a good summary would."}, {"letter": "D", "text": "Keep the single free-text field but increase its maximum character limit so the agent has more room available to describe everything it happened to observe", "correct": false, "explanation": "A larger character limit does not add structure; the agent could still omit the root cause or the recommended action entirely within a longer unstructured note."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f1-092", "domain": 1, "task_id": "1.3", "objective": "Configure subagent invocation, context passing, and spawning", "situation": "A coordinator calls the same \"endpoint-finder\" subagent twice in a row, once for the billing service and once for the notifications service, using two separate Task invocations without resuming a specific agent id. On the second call, the subagent has no awareness that a billing-service scan happened earlier and re-explains basic conventions it already covered.", "question": "Why does this happen, and is it expected behavior?", "options": [{"letter": "A", "text": "Yes, expected: each invocation starts a fresh context unless a specific prior agent is explicitly resumed, so separate calls to the same agent type share no memory", "correct": true, "explanation": "Each subagent invocation is a fresh context by default; nothing is retained between separate calls to the same agent type unless the coordinator explicitly resumes that specific subagent's prior session by its agent id."}, {"letter": "B", "text": "Yes, this is expected, but only because the two calls targeted different services; invoking endpoint-finder twice for the same service would have shared memory automatically", "correct": false, "explanation": "The target service is irrelevant; two separate invocations of the same agent type never share memory automatically regardless of what they are analyzing."}, {"letter": "C", "text": "No, this is a bug: subagent definitions cache their reasoning across calls automatically, so two invocations of endpoint-finder should share memory by default", "correct": false, "explanation": "There is no automatic caching of reasoning across separate subagent invocations; retaining prior context requires explicitly resuming that subagent."}, {"letter": "D", "text": "No, this is a bug: all subagents invoked within the same coordinator session automatically share one combined context window across every Task call", "correct": false, "explanation": "Subagents do not share a combined context window with each other or with the coordinator; each spawn is isolated unless deliberately resumed."}], "correct": "A", "select": 1, "group": "B"}, {"id": "f1-093", "domain": 1, "task_id": "1.6", "objective": "Design task decomposition strategies for complex workflows", "situation": "An agent is tasked with refactoring a shared authentication module that is imported by an unknown number of other modules across a codebase of uncertain size. Before any refactor can proceed safely, the impact radius must be understood, but that radius cannot be known without first searching the codebase.", "question": "How should this task be decomposed?", "options": [{"letter": "A", "text": "Start with a discovery subtask that searches for all importers of the module, then generate follow-up subtasks for each caller as it is found", "correct": true, "explanation": "Since the impact radius is unknown until the codebase is searched, an adaptive plan that starts with discovery and generates follow-up subtasks for each caller found reflects the recommended approach of adjusting decomposition based on what is discovered."}, {"letter": "B", "text": "Begin refactoring the module immediately using a fixed three-step chain of edit, test, and deploy, since refactors always follow those same stages", "correct": false, "explanation": "Jumping straight to a fixed edit-test-deploy chain assumes the affected callers and edit scope are already known, when the premise is that the impact radius must first be discovered."}, {"letter": "C", "text": "Have the model estimate the number of affected callers from the module's name alone, before searching the codebase or making any edits", "correct": false, "explanation": "Estimating impact from the module's name alone skips the actual investigation needed to find real importers and risks missing callers that a search would reveal."}, {"letter": "D", "text": "Split the codebase into per-file passes with no cross-file integration step, treating each file's imports as unrelated to any other file", "correct": false, "explanation": "A shared authentication module's importers are exactly the kind of cross-file relationship that requires an integration pass; treating each file as unrelated ignores the dependency the task is centered on."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f1-094", "domain": 1, "task_id": "1.4", "objective": "Implement multi-step workflows with enforcement and handoff patterns", "situation": "A single support ticket reads: \"I was double-charged for my subscription and I also need my shipping address updated before the next shipment.\" The agent has tools for billing lookups and for address changes. The architect wants both concerns resolved efficiently in one response rather than sequentially re-reading the ticket twice.", "question": "What is the recommended approach?", "options": [{"letter": "A", "text": "Ask the customer to submit two separate tickets, since a single support workflow is only designed to track and resolve one concern per conversation", "correct": false, "explanation": "Pushing the decomposition work onto the customer defeats the purpose of building a workflow that can handle multi-concern requests directly."}, {"letter": "B", "text": "Resolve the billing concern fully first, close that thread, then open an entirely new session with no memory of the ticket to investigate the address change", "correct": false, "explanation": "Starting a fresh session with no memory discards the shared context from the first investigation and forces the ticket to be re-read, which is the inefficiency the architect wants to avoid."}, {"letter": "C", "text": "Decompose the ticket into the billing concern and the address concern, investigate both in parallel using shared context, then synthesize the findings into one unified resolution", "correct": true, "explanation": "Multi-concern requests are handled by splitting them into distinct items and investigating each concurrently within the same shared context, then merging the results into a single coherent reply."}, {"letter": "D", "text": "Skip the address change silently and resolve only the billing concern, since financial issues take unconditional precedence over account-detail updates", "correct": false, "explanation": "Dropping a stated concern without addressing it is not a resolution strategy; both concerns were explicitly raised and both need a response."}], "correct": "C", "select": 1, "group": "A"}, {"id": "f1-095", "domain": 1, "task_id": "1.3", "objective": "Configure subagent invocation, context passing, and spawning", "situation": "A team is building a \"doc-reviewer\" subagent that should be able to read and comment on documentation but must never be able to modify files, even accidentally. The coordinator itself has Read, Edit, Write, Grep, and Glob available.", "question": "How should the doc-reviewer's AgentDefinition be configured to guarantee this constraint?", "options": [{"letter": "A", "text": "Set the coordinator's own allowedTools to just [\"Read\", \"Grep\"] for the duration of the review so neither agent can access Edit or Write", "correct": false, "explanation": "Restricting the coordinator's own tools would also block the coordinator's other legitimate work and is unrelated to constraining what a specific subagent can do."}, {"letter": "B", "text": "Set tools to [\"Read\", \"Grep\"] on the doc-reviewer definition so it inherits only the listed read tools regardless of what the coordinator has available", "correct": true, "explanation": "The tools field on an AgentDefinition is an explicit allowlist; specifying only Read and Grep guarantees the subagent has no access to Edit or Write, independent of what the coordinator or parent session can do."}, {"letter": "C", "text": "Give doc-reviewer the same tools as the coordinator, then rely on the description field to signal that the subagent is read-only in practice", "correct": false, "explanation": "The description field only affects when Claude chooses to invoke the subagent; it does not restrict which tools the subagent can actually call once running."}, {"letter": "D", "text": "Leave tools unset on the doc-reviewer definition and instead instruct it in the prompt field never to call Edit or Write during its review", "correct": false, "explanation": "Leaving tools unset means the subagent inherits all available tools; a prompt instruction is only a suggestion the model could still deviate from, not an enforced restriction."}], "correct": "B", "select": 1, "group": "B"}, {"id": "f1-096", "domain": 1, "task_id": "1.3", "objective": "Configure subagent invocation, context passing, and spawning", "situation": "A coordinator's prompt says only: \"Use the code-reviewer agent to check the authentication module.\" The team wants to guarantee code-reviewer is invoked rather than risk Claude answering the review directly, since automatic delegation based on the description field has been unreliable for this task in the past.", "question": "Does this prompt achieve that guarantee, and why?", "options": [{"letter": "A", "text": "Yes, naming the subagent by name in the prompt is explicit invocation, which bypasses automatic description-based matching and directly invokes that subagent.", "correct": false, "explanation": "This is incorrect. Simply mentioning the subagent's name in a natural language prompt is not a guaranteed forced invocation. The model interprets the prompt and may still decide to handle the task itself. As per Anthropic's documentation, a guarantee requires a forced mechanism like setting the tool_choice parameter to the subagent's name. Explicit naming in the prompt is only a prompt engineering technique, not a hard guarantee."}, {"letter": "B", "text": "No, explicit invocation by name only works for built-in subagents like the general-purpose agent, not for custom AgentDefinition entries like code-reviewer.", "correct": false, "explanation": "There is no such limitation in Anthropic's documentation. Explicit invocation mechanisms (when they are forced via tool_choice or dedicated syntax) work for all subagents, whether built‑in or custom. This option incorrectly claims a restriction that does not exist."}, {"letter": "C", "text": "Yes, but only if the description field is also removed from the AgentDefinition, since a populated description field always overrides an explicit name mention.", "correct": false, "explanation": "This is incorrect. The description field does not override an explicit forced invocation. When using the guaranteed invocation mechanism (e.g., tool_choice or @subagent-name), the description field is irrelevant. Moreover, a simple name mention in the prompt is not the dedicated mechanism required for a guarantee."}, {"letter": "D", "text": "No, the prompt alone is not a guaranteed invocation. While explicitly asking for the code-reviewer agent influences the model, it does not force its use; the system may still respond directly. To reliably enforce delegation, you must use a programmatic constraint such as setting the tool_choice parameter to that subagent or using the dedicated subagent invocation syntax (e.g., @code-reviewer in Claude Code).", "correct": true, "explanation": "This is correct. Anthropic's official documentation indicates that to force a specific tool or subagent, you should use the tool_choice API parameter or the supported subagent invocation syntax (like @subagent-name). A natural language request like \"Use the code-reviewer agent\" is a prompt suggestion, not a guaranteed enforcement. Therefore, the prompt alone does not achieve the team's required guarantee."}], "correct": "D", "select": 1, "group": "B"}, {"id": "f1-097", "domain": 1, "task_id": "1.5", "objective": "Apply Agent SDK hooks for tool call interception and data normalization", "situation": "An agent connects to two MCP servers named \"billing\" and \"inventory\". An architect wants a single PreToolUse hook to run for every tool exposed by the \"billing\" server, without matching any tool from \"inventory\" or any built-in tool like Bash or Read.", "question": "Which matcher achieves this?", "options": [{"letter": "A", "text": "^mcp__billing__$", "correct": false, "explanation": "This adds explicit anchors, so it is evaluated as a regular expression rather than an exact string, but the anchors force it to require an exact match against the literal string mcp__billing__ with nothing following. Real MCP tool names always continue past the server's double underscore with an action name, for example mcp__billing__create_invoice, so no tool name equals this pattern exactly and it matches nothing."}, {"letter": "B", "text": "billing", "correct": false, "explanation": "This matches only a tool literally named \"billing\", which does not exist; MCP tool names always include the mcp__ prefix and server name as part of a longer string."}, {"letter": "C", "text": "^mcp__", "correct": false, "explanation": "This regex matches every tool name starting with mcp__, which includes both billing and inventory tools, so it fails to isolate billing-only calls as required."}, {"letter": "D", "text": "mcp__billing__.*", "correct": true, "explanation": "MCP tools are named mcp__<server>__<action>. Because this pattern contains a regex metacharacter (.*), it's evaluated as an unanchored regex and matches every tool name beginning with mcp__billing__, covering all billing tools without matching inventory or built-in tools."}], "correct": "D", "select": 1, "group": "E"}, {"id": "f1-098", "domain": 1, "task_id": "1.4", "objective": "Implement multi-step workflows with enforcement and handoff patterns", "situation": "Midway through investigating a billing dispute, an agent determines the case requires a policy exception only a human supervisor can approve. The human agent who picks up the case will not have access to the conversation transcript.", "question": "Which handoff summary is most useful to that human agent?", "options": [{"letter": "A", "text": "Customer ID 88213; root cause: a duplicate authorization hold from a retried gateway call was never released; recommended action: void the hold.", "correct": true, "explanation": "The supervisor picking this case up cannot see the conversation, so the summary is the only channel that carries the investigation forward, and this one is self-contained: it names the account the case belongs to, states what the investigation concluded, and proposes one concrete action the supervisor can approve or override. The Claude Agent SDK documents the same constraint when an agent hands work to a fresh context window - the delegating prompt is the only content passed across, so the identifiers, errors and decisions the recipient needs must be written into it rather than left in a record the recipient has to go and find. That is also the answer to the objection that this material could be tracked elsewhere: wherever it is stored, the handoff message itself has to carry it, because nothing else reaches the reader."}, {"letter": "B", "text": "See the attached conversation transcript for full details; the customer's most recent message explains the situation better than a summary could", "correct": false, "explanation": "The human agent explicitly has no access to the transcript, so a summary that points at it transmits nothing at all - it is the one shape of handoff guaranteed to fail under the stated constraint. Handing a case to a reader without the prior conversation means distilling the trace into the details that still matter, which is precisely what summarisation for a new context window is for; forwarding the raw conversation instead pushes the whole investigation back onto the person receiving it."}, {"letter": "C", "text": "The customer seems upset about a charge and asked several questions before the case was escalated for further human review of the account history", "correct": false, "explanation": "Describing the customer's mood and the fact that several questions were asked narrates the conversation instead of distilling it: there is no account reference, no statement of what the disputed charge was, and no finding from the investigation. That is the redundant narrative a summary written for a fresh reader is meant to discard, not the substance it is meant to preserve, so the supervisor still has to reconstruct everything that matters."}, {"letter": "D", "text": "Billing issue escalated; customer wants money back; please review and use your judgment on what discount or credit, if any, is appropriate here", "correct": false, "explanation": "Announcing an escalation with no account reference, no finding from the investigation that was just performed, and no proposed resolution leaves the supervisor to work the case again from the beginning. Delegation guidance asks for an objective, an output format and clear task boundaries, and inviting the reader to 'use your judgment on what discount or credit, if any' supplies none of those - it transfers the decision without transferring the basis for it."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f1-099", "domain": 1, "task_id": "1.6", "objective": "Design task decomposition strategies for complex workflows", "situation": "An architect is reviewing a design where a large code review agent analyzes every file in a pull request in one combined prompt, then a second prompt asks it to summarize cross-file issues. Stakeholders report that the cross-file summary frequently misses real integration bugs that are visible when files are compared directly.", "question": "What is the most likely cause, and what change would fix it?", "options": [{"letter": "A", "text": "The two passes should run in reverse order, producing the cross-file summary before any individual file has actually been analyzed", "correct": false, "explanation": "Running the cross-file summary before any individual file has been analyzed leaves it with no per-file findings to compare, making it even less likely to catch integration bugs."}, {"letter": "B", "text": "The first pass already dilutes attention across all files at once, so splitting it into per-file analyses gives the cross-file pass richer findings", "correct": true, "explanation": "Combining all files into one first pass causes attention dilution, producing shallow per-file analysis; separating that into individual per-file passes before the cross-file pass gives the integration step higher-quality inputs to work from, which is the documented fix for this exact pattern."}, {"letter": "C", "text": "The cross-file pass should be dropped and replaced with one combined pass that covers every file and every possible relationship in a single prompt", "correct": false, "explanation": "Removing the cross-file pass eliminates the only step designed to catch integration issues, which is the opposite of what the stakeholders need."}, {"letter": "D", "text": "The two-prompt structure is fine as is, but the model needs a longer system prompt that more precisely defines what an integration bug is", "correct": false, "explanation": "A longer definition of integration bugs does not solve the underlying problem that the first pass never produced precise per-file findings to begin with, since all files competed for attention in one prompt."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f1-100", "domain": 1, "task_id": "1.1", "objective": "Design and implement agentic loops for autonomous task execution", "situation": "A developer debugging an agent loop notices that on iteration three, Claude requests the same file-read tool with the same arguments used on iteration one, as if it never saw the earlier result. The loop's history-building code only appends the assistant's text blocks to conversation history, never the tool_use or tool_result blocks.", "question": "What is the most likely cause of the repeated call?", "options": [{"letter": "A", "text": "The Messages API automatically clears tool_use blocks from history every two iterations to control overall token usage.", "correct": false, "explanation": "Anthropic offers a \"tool result clearing\" feature that can automatically clear old tool use results as token limits are approached, but there is no documented behavior of clearing tool_use blocks every two iterations. Clearing is based on token limits and context management, not a fixed iteration count, and it does not cause the model to forget a result if that result is properly included in the history."}, {"letter": "B", "text": "The tool_result from iteration one was never added to the context sent back to Claude, so the model has no record the file was already read.", "correct": true, "explanation": "According to Anthropic's tool use documentation, when using client-executed tools, you must send a new request containing the original messages, the assistant's response, and a user message with the tool_result blocks. If the history-building code only appends assistant text blocks and omits tool_use and tool_result blocks, the model never receives the result of its earlier tool call, so it has no record that the file was already read. This causes the model to repeat the same tool request on a later iteration."}, {"letter": "C", "text": "Claude's context window silently resets whenever stop_reason returns \"tool_use\" on two consecutive iterations in a row.", "correct": false, "explanation": "Anthropic documentation does not state that the context window resets when stop_reason returns \"tool_use\" on consecutive iterations. The context window retains conversation history, including tool results, when those results are correctly provided. The described issue arises because the history-building code omits the tool blocks entirely."}, {"letter": "D", "text": "The tool being called is inherently non-idempotent, so the Messages API requires Claude to call it again on every single subsequent iteration.", "correct": false, "explanation": "Non-idempotency refers to whether a tool has side effects when called multiple times; it is not a reason for the Messages API to require repeated calls on every iteration. The repeated call in this scenario is caused by missing tool result context in the conversation history, not by an inherent property of the tool."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f1-101", "domain": 1, "task_id": "1.2", "objective": "Orchestrate multi-agent systems with coordinator-subagent patterns", "situation": "An architect delegates a task to a subagent using the Agent tool, expecting it to reference a decision the user made three turns earlier in the main conversation about which authentication provider to use. The subagent's response ignores that decision entirely and proposes a different provider.", "question": "What is the most likely cause, and how should the architect fix it?", "options": [{"letter": "A", "text": "The coordinator's tool permissions blocked the authentication context from being read, so the architect must widen the subagent's tool access", "correct": false, "explanation": "Tool permissions govern which actions a subagent can take, not whether it receives conversation history, so widening tool access would not surface the missing decision."}, {"letter": "B", "text": "The subagent inherited a stale cached copy of the conversation and needs its session resumed to pick up the recent turns", "correct": false, "explanation": "Subagents don't cache or partially load prior turns; they simply never receive the parent conversation at all, so there is no stale copy to resume."}, {"letter": "C", "text": "The subagent's context starts fresh each call, so the decision must be included directly in the Agent tool's prompt text", "correct": true, "explanation": "A subagent's context window starts fresh and does not automatically inherit the parent's conversation history, so any needed decisions must be passed explicitly in the Agent tool's prompt string."}, {"letter": "D", "text": "The subagent's system prompt overrides user decisions by design, so the architect must disable its custom system prompt", "correct": false, "explanation": "A custom system prompt defines the subagent's role and expertise; it doesn't override user decisions, and disabling it would remove needed specialization without fixing the missing context."}], "correct": "C", "select": 1, "group": "A"}, {"id": "f1-102", "domain": 1, "task_id": "1.6", "objective": "Design task decomposition strategies for complex workflows", "situation": "A support-ticket triage system always performs the same three actions on every incoming ticket: classify category, extract customer sentiment, and generate a brief summary. None of these actions ever depend on the outcome of another action within the same ticket.", "question": "Based on Anthropic's workflow patterns, which decomposition or execution pattern is most appropriate, and what is the key justification?", "options": [{"letter": "A", "text": "A dynamic orchestrator, because sentiment extraction might reveal information that changes how many classification steps are needed.", "correct": false, "explanation": "Anthropic's orchestrator-workers workflow is intended for complex tasks where the required subtasks cannot be predicted in advance and may depend on intermediate findings. The scenario explicitly states that the same three actions are always performed and that none depend on another action within the ticket, so dynamic orchestration is not needed and the justification contradicts the premise."}, {"letter": "B", "text": "A fixed prompt chain, because the subtasks are the same for every ticket and can be cleanly predetermined regardless of ticket content.", "correct": false, "explanation": "A fixed prompt chain is sequential: each LLM call processes the output of the previous one. Anthropic describes prompt chaining as useful when subtasks are fixed and sequential, prioritizing accuracy over latency. Here the subtasks are fixed but independent, so chaining them would add unnecessary latency and does not match the recommended parallelization pattern for independent subtask execution."}, {"letter": "C", "text": "Execute the three actions in parallel using separate, concurrent LLM calls, then aggregate the results. Justification: The subtasks are independent, so parallelization improves efficiency and follows Anthropic's recommended pattern for independent subtask execution.", "correct": true, "explanation": "Anthropic recommends a Parallelization workflow when a task can be broken into independent subtasks, including the Sectioning variation where independent subtasks are run concurrently. Because the three actions (classification, sentiment extraction, summary generation) do not depend on each other, separate concurrent LLM calls can handle each consideration more efficiently, reducing latency while preserving accuracy. This is the appropriate execution pattern for fixed, independent subtasks."}, {"letter": "D", "text": "An adaptive investigation plan, because triage requires generating new subtasks based on what earlier steps in the ticket discover.", "correct": false, "explanation": "An adaptive or open-ended investigation process is suited to problems where the number and type of steps are not predictable in advance. This scenario is the opposite: every ticket receives the same three predetermined, independent actions, so no adaptive generation of new subtasks is required."}], "correct": "C", "select": 1, "group": "A"}, {"id": "f1-103", "domain": 1, "task_id": "1.6", "objective": "Design task decomposition strategies for complex workflows", "situation": "An architect is choosing between prompt chaining and orchestrator-workers for a document-processing pipeline that converts a PDF invoice into a structured JSON record, validates required fields are present, and stores the record. Every invoice follows the same three-field schema and the steps never branch.", "question": "What is the strongest argument for choosing prompt chaining here?", "options": [{"letter": "A", "text": "Orchestrator-workers setups cannot process PDF documents under any circumstances, making chaining the only technically feasible option here", "correct": false, "explanation": "Orchestrator-workers setups are capable of processing documents like PDFs; the reason to prefer chaining here is the task's predictability, not a technical limitation of the alternative pattern."}, {"letter": "B", "text": "Prompt chaining always produces more accurate JSON extraction results than any other decomposition pattern, no matter how predictable the task is", "correct": false, "explanation": "Prompt chaining's benefit is predictability and consistency for well-defined tasks, not a general claim of superior extraction accuracy over other architectures in all cases."}, {"letter": "C", "text": "The task decomposes cleanly into fixed, predictable subtasks always in the same order, giving chaining predictability without delegation overhead", "correct": true, "explanation": "The stated guidance is that fixed workflows should be chosen when the process steps can be reliably predicted and follow consistent patterns, which matches an invoice pipeline that always uses the same three-field schema and step order."}, {"letter": "D", "text": "Chaining is required for any workflow with more than a single step, no matter how those steps depend on or relate to one another", "correct": false, "explanation": "Having multiple steps alone does not mandate chaining; the deciding factor is whether those steps are predictable and fixed versus needing to emerge dynamically from intermediate results."}], "correct": "C", "select": 1, "group": "A"}, {"id": "f1-104", "domain": 1, "task_id": "1.7", "objective": "Manage session state, resumption, and forking", "situation": "A session investigating a rate-limiter bug hit error_max_turns before reaching a conclusion. The architect wants to keep the analysis already performed and simply let the agent keep working with a higher turn ceiling, rather than repeating the investigation.", "question": "What should the architect do?", "options": [{"letter": "A", "text": "Fork the session and set a higher max_turns only on the newly forked branch of it", "correct": false, "explanation": "The research does not support forking a session as the recommended fix for error_max_turns. The documented recovery method is to resume the original session_id with a higher max_turns on the follow-up query, so the agent continues with the full context from the prior run."}, {"letter": "B", "text": "Resume the session's ID with a higher max_turns value configured on the new follow-up query", "correct": true, "explanation": "error_max_turns indicates the Claude agent session exceeded its configured conversational or tool-use turn limit; it is not the same as an HTTP 429 rate_limit_error. The documented recovery path is to resume the same session with a higher limit, preserving accumulated context. Capture the session_id from the ResultMessage, then pass that session_id and the new, higher max_turns value in the ClaudeAgentOptions or Options when calling query()."}, {"letter": "C", "text": "Start a fresh session and manually restate every file the agent had already read before", "correct": false, "explanation": "Starting a fresh session forces the agent to re-read and reprocess work that the previous session already completed, wasting time and cost. The documented approach for recovering from error_max_turns is to resume the existing session with a higher max_turns, not to rebuild context manually."}, {"letter": "D", "text": "Rerun the original prompt from scratch since a turn limit indicates the approach was flawed", "correct": false, "explanation": "error_max_turns is a configurable limit on how many agentic turns may occur in a session; hitting it does not mean the approach itself was flawed. The correct action is to resume the session with a higher turn ceiling, allowing the agent to continue its existing analysis rather than repeating it from the beginning."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f1-105", "domain": 1, "task_id": "1.3", "objective": "Configure subagent invocation, context passing, and spawning", "situation": "A coordinator prompt for a \"research-assistant\" subagent currently reads: \"Step 1: search for the top 5 articles. Step 2: open each one. Step 3: extract the publish date. Step 4: return a table.\" Reviewers find the subagent produces a table even when the most useful sources have no clear publish date, or when six sources would answer the question better than five.", "question": "What change to the coordinator prompt would improve outcomes here?", "options": [{"letter": "A", "text": "Keep the same four numbered steps but increase the required article count from five to six so the subagent always gathers slightly more source material overall", "correct": false, "explanation": "Hard-coding a different fixed count still leaves the subagent locked into procedural steps and does not address sources with no clear publish date or cases needing a different number of sources."}, {"letter": "B", "text": "Rewrite the prompt around the research goal and quality bar, such as finding credible, relevant sources and reporting what is verifiable, rather than a fixed step sequence", "correct": true, "explanation": "Framing the coordinator's prompt around goals and quality criteria rather than a rigid procedure lets the subagent adapt, such as gathering more or fewer sources or handling missing fields sensibly, instead of following steps that don't fit every case."}, {"letter": "C", "text": "Replace the step list with an even more granular ten-step sequence covering search syntax, click order, and field extraction to remove ambiguity entirely", "correct": false, "explanation": "Making the procedure even more granular increases rigidity rather than adaptability, making it harder, not easier, for the subagent to handle sources that don't fit the assumed shape."}, {"letter": "D", "text": "Keep the same four numbered steps but add a fifth step instructing the subagent to double-check the publish date it already extracted back in step three", "correct": false, "explanation": "Adding a verification step keeps the same rigid structure and does not help when a source simply has no publish date to extract."}], "correct": "B", "select": 1, "group": "B"}, {"id": "f1-106", "domain": 1, "task_id": "1.6", "objective": "Design task decomposition strategies for complex workflows", "situation": "An engineering lead wants an agent workflow that first generates a marketing blog post outline, then expands the outline into full prose, then proofreads the prose for grammar. Each stage's output is always exactly what the next stage needs, and the number of stages never varies.", "question": "Which architectural choice best matches this scenario?", "options": [{"letter": "A", "text": "A single subagent with a general instruction to write and polish a blog post, leaving stage boundaries implicit within one response", "correct": false, "explanation": "Merging all stages into one implicit response removes the checkpoints that let a malformed outline be caught before wasted effort is spent expanding it into prose."}, {"letter": "B", "text": "An adaptive investigation plan that generates new subtasks whenever the proofreading stage finds an issue", "correct": false, "explanation": "Treating grammar issues as triggers for new adaptive subtasks overcomplicates a proofreading step that should simply be a bounded stage in the fixed chain, not an open-ended investigation."}, {"letter": "C", "text": "A dynamic orchestrator-workers setup that decides how many drafting stages to run based on the outline's content", "correct": false, "explanation": "The number of stages is stated to never vary, so a dynamic orchestrator that decides stage count at runtime adds decision-making overhead the fixed workflow does not need."}, {"letter": "D", "text": "A prompt chain with a checkpoint between each stage so a malformed outline or draft can be caught before the next stage proceeds", "correct": true, "explanation": "The stages are fixed, predictable, and always follow the same order, which is the textbook case for prompt chaining with programmatic checkpoints between steps to verify progress."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f1-107", "domain": 1, "task_id": "1.3", "objective": "Configure subagent invocation, context passing, and spawning", "situation": "A coordinator dynamically builds a \"security-reviewer\" AgentDefinition and needs the strict variant of the review to use a noticeably more capable model than routine reviews, without changing the model used for the rest of the coordinator's own work.", "question": "Which configuration achieves this?", "options": [{"letter": "A", "text": "Resume the coordinator's session with a different model argument each time a strict review is requested, then dispatch the subagent from that resumed session", "correct": false, "explanation": "Resuming a session continues prior conversation history; it is not how per-subagent model selection is configured, and it would unnecessarily entangle session state with model choice."}, {"letter": "B", "text": "Fork the coordinator's session before every strict review so the fork inherits a separate, upgraded default model for all subsequent subagent calls", "correct": false, "explanation": "Forking copies conversation history into a new session; it does not change which model any agent uses and is unrelated to per-subagent model overrides."}, {"letter": "C", "text": "Add the more capable model's name as an entry in the security-reviewer's tools array so the runtime treats it as an available capability for that call", "correct": false, "explanation": "The tools array lists tool names the subagent may call, not model identifiers; models are configured through the dedicated model field."}, {"letter": "D", "text": "Set the model field on the strict security-reviewer's AgentDefinition to the more capable model, leaving the coordinator's own model configuration untouched", "correct": true, "explanation": "The model field on an AgentDefinition overrides the model for that specific subagent only; it can be set per definition (e.g., built dynamically by a factory function) without touching the coordinator's own model."}], "correct": "D", "select": 1, "group": "B"}, {"id": "f1-108", "domain": 1, "task_id": "1.5", "objective": "Apply Agent SDK hooks for tool call interception and data normalization", "situation": "During a security review, an architect is asked to justify why a refund-blocking rule was implemented as a PreToolUse hook instead of as a stricter, more detailed system prompt paragraph that explicitly lists dollar thresholds and escalation steps.", "question": "Which justification is most accurate?", "options": [{"letter": "A", "text": "A hook is required because the Agent SDK does not allow system prompts to reference numeric thresholds like dollar amounts, and the only way to enforce refund-blocking based on amounts is through a hook that intercepts tool use.", "correct": false, "explanation": "The Agent SDK allows system prompts to include numeric thresholds; there is no technical restriction preventing prompts from referencing dollar amounts. The choice of a hook is based on enforcement reliability, not a limitation of prompt content."}, {"letter": "B", "text": "A hook is faster to write than a detailed prompt paragraph because it avoids specifying exact dollar thresholds and escalation steps, making it the preferred method when rapid deployment is prioritized during a security review.", "correct": false, "explanation": "The speed of writing a hook versus a detailed prompt paragraph is not the architectural justification for choosing a hook. The core reason is deterministic enforcement, not how quickly the rule can be deployed."}, {"letter": "C", "text": "A hook produces a more articulate explanation to the end user than a system prompt would, since hookSpecificOutput text is always shown verbatim in the chat transcript, providing clearer messages than natural language prompts.", "correct": false, "explanation": "While a hook can provide verbatim output to the user, this is a user experience detail, not the primary justification for using a hook over a system prompt for policy enforcement. The key advantage is deterministic, out-of-model enforcement, not message clarity."}, {"letter": "D", "text": "A hook enforces the rule entirely outside of model generation, so it cannot be bypassed by unusual phrasing, prompt injection in tool results, or the model simply misapplying a complex instruction.", "correct": true, "explanation": "Because the hook executes as deterministic code outside the model's generation process, it is immune to bypass via unusual phrasing, prompt injection in tool results, or model misinterpretation. This guarantees enforcement of hard compliance rules that a system prompt cannot reliably provide."}], "correct": "D", "select": 1, "group": "E"}, {"id": "f1-109", "domain": 1, "task_id": "1.1", "objective": "Design and implement agentic loops for autonomous task execution", "situation": "A junior engineer proposes capping an autonomous research agent at exactly 5 tool-call iterations and treating that cap as the loop's primary stopping mechanism, regardless of stop_reason.", "question": "What is the main architectural problem with relying on a fixed iteration cap this way?", "options": [{"letter": "A", "text": "It causes the API to immediately reject the request, since the Messages API enforces its own hard limit of five tool calls total", "correct": false, "explanation": "The Messages API has no built-in limit on tool calls per conversation; any cap is an application-level choice."}, {"letter": "B", "text": "It can cut the loop off while stop_reason is still \"tool_use\", forcing the agent to abandon work Claude has not finished reasoning through", "correct": true, "explanation": "An iteration cap is arbitrary and unrelated to task completion, so using it as the primary stop condition can truncate the loop mid-task; termination should instead follow stop_reason == \"end_turn\"."}, {"letter": "C", "text": "It forces every one of the later requests to omit the tools parameter entirely, disabling any further tool use for the rest of that session", "correct": false, "explanation": "Reaching an iteration cap only affects whether the loop sends another request; it does not alter the tools parameter of any request."}, {"letter": "D", "text": "It stops tool_result blocks from being appended to conversation history entirely once the third iteration has completed", "correct": false, "explanation": "Appending tool_result blocks is a client-side responsibility that has nothing to do with any iteration counter the loop maintains."}], "correct": "B", "select": 1, "group": "A"}, {"id": "g7", "domain": 2, "scenario": "Multi-agent Research System", "situation": "Production logs show a persistent pattern: requests like “analyze the uploaded quarterly report” are routed to the web-search agent 45% of the time instead of the document analysis agent. Reviewing tool definitions, you find that the web-search agent has a tool `analyze_content` described as “analyzes content and extracts key information,” while the document analysis agent has a tool `analyze_document` described as “analyzes documents and extracts key information.” How should you fix the misrouting problem?", "question": "How should you fix the misrouting problem?", "options": [{"letter": "A", "text": "Add a pre-routing classifier that detects whether the user refers to uploaded files or web content before the coordinator decides on delegation.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Rename the web-search tool to `extract_web_results` and update its description to “processes and returns information retrieved from web search and URLs.”", "correct": true, "explanation": "Renaming the web-search tool to `extract_web_results` and updating its description to explicitly reference web search and URLs directly removes the root cause by eliminating semantic overlap between the two tool names and descriptions. This makes each tool’s purpose unambiguous, enabling the coordinator to reliably distinguish document analysis from web search."}, {"letter": "C", "text": "Add few-shot examples to the coordinator prompt showing correct routing: “User uploads a quarterly report → document analysis agent” and “User asks about a web page → web-search agent.”", "correct": false, "explanation": ""}, {"letter": "D", "text": "Expand the document analysis tool description with usage examples like “Use for uploaded PDFs, Word docs, and spreadsheets,” leaving the web-search tool unchanged.", "correct": false, "explanation": ""}], "correct": "B", "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "group": "J"}, {"id": "g10", "domain": 2, "scenario": "Multi-agent Research System", "situation": "In your system design, you gave the document analysis agent access to a general-purpose tool `fetch_url` so it could download documents by URL. Production logs show this agent now frequently downloads search engine results pages to perform ad hoc web search—behavior that should be routed through the web-search agent—causing inconsistent results. Which fix is most effective?", "question": "Which fix is most effective?", "options": [{"letter": "A", "text": "Replace `fetch_url` with a `load_document` tool that validates that URLs point to document formats.", "correct": true, "explanation": "Replacing a general-purpose tool with a document-specific tool that validates URLs against document formats fixes the root cause by constraining capability at the interface level. This follows the principle of least privilege, making undesired search behavior impossible rather than merely discouraged."}, {"letter": "B", "text": "Remove `fetch_url` from the document analysis agent and route all URL fetching through the coordinator to the web-search agent.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Implement filtering that blocks `fetch_url` calls to known search engine domains while allowing other URLs.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Add instructions to the document analysis agent prompt that `fetch_url` should only be used to download document URLs, not to search.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "2.3", "objective": "Distribute tools appropriately across agents and configure tool choice", "group": "B"}, {"id": "g15", "domain": 2, "scenario": "Multi-agent Research System", "situation": "In testing, you observe that the synthesis agent often needs to verify specific claims while merging results. Currently, when verification is needed, the synthesis agent returns control to the coordinator, which calls the web-search agent and then re-invokes synthesis with the results. This adds 2–3 extra loops per task and increases latency by 40%. Your assessment shows 85% of these verifications are simple fact checks (dates, names, stats) and 15% require deeper research. Which approach most effectively reduces overhead while preserving system reliability?", "question": "Which approach is most effective?", "options": [{"letter": "A", "text": "Give the synthesis agent access to all web-search tools so it can handle any verification need directly without coordinator loops.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Have the synthesis agent accumulate all verification needs and return them as a batch to the coordinator at the end, which then sends them all to the web-search agent at once.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Have the web-search agent proactively cache extra context around each source during initial research in anticipation of synthesis needing verification.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Give the synthesis agent a limited-scope `verify_fact` tool for simple checks, while routing complex verifications through the coordinator to the web-search agent.", "correct": true, "explanation": "A limited-scope fact-verification tool lets the synthesis agent handle 85% of simple checks directly, eliminating most loops, while preserving the coordinator delegation path for the 15% of complex verifications. This applies least privilege while significantly reducing latency."}], "correct": "D", "task_id": "2.3", "objective": "Distribute tools appropriately across agents and configure tool choice", "group": "B"}, {"id": "g18", "domain": 2, "scenario": "Claude Code for Continuous Integration", "situation": "Your code review component is iterative: Claude analyzes the changed file, then may request related files (imports, base classes, tests) via tool calls to understand context before providing final feedback. Your application defines a tool that lets Claude request file contents; Claude calls the tool, gets results, and continues analysis. You’re evaluating batch processing to reduce API cost. What is the primary technical limitation when considering batch processing for this workflow?", "question": "What is the primary technical limitation?", "options": [{"letter": "A", "text": "Batch processing does not include correlation IDs to map outputs back to input requests.", "correct": false, "explanation": ""}, {"letter": "B", "text": "The asynchronous model cannot execute tools mid-request and return results for Claude to continue analysis.", "correct": true, "explanation": "A “fire-and-forget” asynchronous Batch API model has no mechanism to intercept a tool call during a request, execute the tool, and return results for Claude to continue analysis. This is fundamentally incompatible with iterative tool-calling workflows that require multiple tool request/response rounds within a single logical interaction."}, {"letter": "C", "text": "The Batch API does not support tool definitions in request parameters.", "correct": false, "explanation": ""}, {"letter": "D", "text": "The batch processing latency of up to 24 hours is too slow for pull request feedback, although the workflow would otherwise function.", "correct": false, "explanation": ""}], "correct": "B", "task_id": "4.5", "objective": "Design efficient batch processing strategies", "group": "H"}, {"id": "g46", "domain": 2, "scenario": "Customer Support Agent", "situation": "While testing, you notice the agent often calls `get_customer` when users ask about order status, even though `lookup_order` would be more appropriate. What should you check first to address this problem?", "question": "What should you check first?", "options": [{"letter": "A", "text": "Implement a preprocessing classifier to detect order-related requests and route them directly to `lookup_order`.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Reduce the number of tools available to the agent to simplify choice.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Add few-shot examples to the system prompt covering all possible order request patterns to improve tool selection.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Check the tool descriptions to ensure they clearly differentiate each tool’s purpose.", "correct": true, "explanation": "Tool descriptions are the primary input the model uses to decide which tool to call. When an agent consistently picks the wrong tool, the first diagnostic step is to verify that tool descriptions clearly separate each tool’s purpose and usage boundaries."}], "correct": "D", "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "group": "J"}, {"id": "g51", "domain": 2, "scenario": "Customer Support Agent", "situation": "Production logs show that in 12% of cases your agent skips `get_customer` and calls `lookup_order` directly using only the customer-provided name, sometimes leading to misidentified accounts and incorrect refunds. What change most effectively fixes this reliability problem?", "question": "What change is most effective?", "options": [{"letter": "A", "text": "Add few-shot examples showing that the agent always calls `get_customer` first, even when customers voluntarily provide order details.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Implement a routing classifier that analyzes each request and enables only a subset of tools appropriate for that request type.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Add a programmatic precondition that blocks `lookup_order` and `process_refund` until `get_customer` returns a verified customer identifier.", "correct": true, "explanation": "A programmatic precondition provides a deterministic guarantee that required sequencing is followed. It’s the most effective approach because it eliminates the possibility of skipping verification, regardless of LLM behavior."}, {"letter": "D", "text": "Strengthen the system prompt stating that customer verification via `get_customer` is mandatory before any order operations.", "correct": false, "explanation": ""}], "correct": "C", "task_id": "1.5", "objective": "Apply Agent SDK hooks for tool call interception and data normalization", "group": "E"}, {"id": "g55", "domain": 2, "scenario": "Customer Support Agent", "situation": "Your `get_customer` tool returns all matches when searching by name. Currently, when there are multiple results, Claude picks the customer with the most recent order, but production data shows this selects the wrong account 15% of the time for ambiguous matches. How should you address this?", "question": "How should you address this?", "options": [{"letter": "A", "text": "Implement a confidence scoring system that acts autonomously above 85% confidence and requests clarification below the threshold.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Instruct Claude to request an additional identifier (email, phone, or order number) when `get_customer` returns multiple matches before taking any customer-specific action.", "correct": true, "explanation": "Asking the user for an additional identifier is the most reliable way to resolve ambiguity because the user has definitive knowledge of their identity. One extra conversational turn is a small price to pay to eliminate a 15% error rate caused by choosing the wrong account."}, {"letter": "C", "text": "Modify `get_customer` to return only a single most-likely match based on a ranking algorithm, eliminating ambiguity.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Add few-shot examples to the prompt demonstrating correct reasoning and tool sequencing for ambiguous matches.", "correct": false, "explanation": ""}], "correct": "B", "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "group": "J"}, {"id": "g57", "domain": 2, "scenario": "Customer Support Agent", "situation": "Production logs show the agent often calls `get_customer` when users ask about orders (e.g., “check my order #12345”) instead of calling `lookup_order`. Both tools have minimal descriptions (“Gets customer information” / “Gets order details”) and accept similar-looking identifier formats. What is the most effective first step to improve tool selection reliability?", "question": "What is the most effective first step?", "options": [{"letter": "A", "text": "Implement a routing layer that analyzes user input before each turn and preselects the correct tool based on detected keywords and ID patterns.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Combine both tools into a single `lookup_entity` that accepts any identifier and internally decides which backend to query.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Add few-shot examples to the system prompt demonstrating correct tool selection patterns, with 5–8 examples routing order-related queries to `lookup_order`.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Expand each tool’s description to include input formats, example queries, edge cases, and boundaries explaining when to use it versus similar tools.", "correct": true, "explanation": "Expanding tool descriptions with input formats, example queries, edge cases, and clear boundaries directly fixes the root cause—minimal descriptions that don’t give the LLM enough information to distinguish similar tools. It’s a low-effort, high-impact first step that improves the primary mechanism the LLM uses for tool selection."}], "correct": "D", "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "group": "J"}, {"id": "g59", "domain": 2, "scenario": "Customer Support Agent", "situation": "Production logs show the agent misinterprets outputs from your MCP tools: Unix timestamps from `get_customer`, ISO 8601 dates from `lookup_order`, and numeric status codes (1=pending, 2=shipped). Some tools are third-party MCP servers you cannot modify. Which approach to data format normalization is most maintainable?", "question": "Which approach is most maintainable?", "options": [{"letter": "A", "text": "Use a PostToolUse hook to intercept tool outputs and apply formatting transformations before the agent processes them.", "correct": true, "explanation": "A PostToolUse hook provides a centralized, deterministic point to intercept and normalize all tool outputs—including third-party MCP server data—before the agent processes them. It’s more maintainable because transformations live in code and apply uniformly, rather than relying on LLM interpretation."}, {"letter": "B", "text": "Modify tools you control to return human-readable formats and create wrappers for third-party tools.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Create a `normalize_data` tool that the agent calls after every data retrieval to transform values.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Add detailed format documentation to the system prompt explaining each tool’s data conventions.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "1.5", "objective": "Apply Agent SDK hooks for tool call interception and data normalization", "group": "E"}, {"id": "g61", "domain": 2, "scenario": "Conversational AI Architecture Patterns", "situation": "Your `remove_team_member` tool uses a `dry_run: boolean` parameter for previewing impacts before execution. Production monitoring shows the agent bypasses the preview step by calling with `dry_run=false` directly. You need to ensure every removal is preceded by a preview that the user explicitly confirms.", "question": "What is the most reliable approach?", "options": [{"letter": "A", "text": "Add server-side validation that permits `dry_run=false` only when a `dry_run=true` call with identical parameters occurred within the past 60 seconds.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Annotate the tool as requiring confirmation and configure the orchestration layer to prompt the user for approval before forwarding any calls to annotated tools.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Add detailed instructions and few-shot examples to the tool description requiring the agent to always call with `dry_run=true` first and wait for user confirmation before calling again.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Replace with two tools: `preview_remove_member` returns impact details and a single-use confirmation token; `execute_remove_member` requires that token, binding execution to the preview.", "correct": true, "explanation": "The two-tool token-binding approach makes it architecturally impossible to execute without a prior preview—the execute tool literally requires a token that only the preview tool can generate. This is the only approach that enforces the constraint at the code level rather than relying on LLM compliance with instructions (C), timing heuristics (A), or orchestration infrastructure (B)."}], "correct": "D", "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "group": "J"}, {"id": "g62", "domain": 2, "scenario": "Conversational AI Architecture Patterns", "situation": "Production monitoring shows your `search_catalog` tool fails 12% of the time: 8% are network timeouts that succeed when retried, and 4% are query syntax errors that never succeed regardless of retries. Currently both error types are returned identically, causing wasted retries.", "question": "How should you modify the tool's error handling?", "options": [{"letter": "A", "text": "Add few-shot examples to your system prompt demonstrating how to distinguish network errors from syntax errors.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Apply exponential backoff retry logic to all errors uniformly.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Implement automatic retry with backoff for network timeouts inside the tool; return syntax errors immediately with parameter validation details.", "correct": true, "explanation": "Handling retries at the tool level for transient errors is the correct abstraction boundary—the tool has definitive knowledge of the error type and can implement deterministic retry logic without relying on the agent to interpret a flag (D) or follow prompt-level instructions (A). Uniform backoff (B) wastes time on syntax errors that will never succeed."}, {"letter": "D", "text": "Return all errors with a `retryable` boolean flag and error type details.", "correct": false, "explanation": ""}], "correct": "C", "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "group": "J"}, {"id": "m3", "domain": 2, "scenario": null, "situation": "The document analysis agent has a single `analyze_document` tool that takes a document and a free-text instruction parameter. During evaluation, requests like \"extract the key financial metrics\" often return narrative summaries, while \"summarize the methodology\" sometimes returns raw data tables. The synthesis agent reports that 35% of analysis results require re-requests with clarified instructions.", "question": "What's the most effective way to improve reliability?", "options": [{"letter": "A", "text": "Split the generic tool into purpose-specific tools—`extract_data_points`, `summarize_content`, `verify_claim_against_source`—each with defined input/output contracts.", "correct": true, "explanation": "Correct. Free-text instructions put the semantics in prose, which the model interprets inconsistently. Purpose-specific tools give the model an explicit, well-typed contract to pick between."}, {"letter": "B", "text": "Keep the single tool but add an `analysis_type` enum parameter requiring explicit selection between extraction, summarization, and verification modes.", "correct": false, "explanation": "Better than free text, but still a single tool with a single output shape. Output still has to be a generic string, and the model can mode-mismatch. Separate tools enforce distinct I/O contracts."}, {"letter": "C", "text": "Have the coordinator pre-classify each analysis request before passing instructions to the document analysis agent.", "correct": false, "explanation": "Moves the ambiguity upstream without fixing it. The tool contract is still fuzzy; the coordinator's classification is one more error surface."}, {"letter": "D", "text": "Enhance the tool description with detailed examples showing how different instruction phrasings should map to different output formats.", "correct": false, "explanation": "Examples help, but tool-use reliability comes first from clear tool boundaries and output schemas, not from heroic prompt engineering on one generic tool."}], "correct": "A", "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "group": "J"}, {"id": "m16", "domain": 2, "scenario": null, "situation": "After integrating a local MCP server providing code analysis tools (`analyze_dependencies`, `find_dead_code`, `calculate_complexity`), you verify the server is healthy and tools appear in the tools/list response. However, you observe that the agent consistently uses Grep to search for import statements instead of calling `analyze_dependencies`—even when users explicitly ask about \"code dependencies.\" Examining tool definitions reveals: MCP: `analyze_dependencies` - \"Analyzes dependency graph\" Built-in: Grep - \"Search file contents for a pattern using regular expressions. Returns matching lines with line numbers and surrounding context.\"", "question": "What's the most effective approach to improve the agent's selection of MCP tools?", "options": [{"letter": "A", "text": "Remove Grep from available tools when the MCP server is connected to eliminate functional overlap.", "correct": false, "explanation": "Crippling a general tool to force adoption of a specialized one punishes other legitimate Grep use cases."}, {"letter": "B", "text": "Add routing instructions to the system prompt specifying that dependency-related questions should use MCP tools rather than Grep.", "correct": false, "explanation": "Prompt-level routing can help, but the root cause is that the MCP tool's description is thinner than Grep's. Fix the tool descriptions first."}, {"letter": "C", "text": "Split `analyze_dependencies` into granular tools (`list_imports`, `resolve_transitive_deps`, `detect_circular_deps`) so each has a focused purpose less likely to overlap with Grep.", "correct": false, "explanation": "Splitting is sometimes right, but here the sibling tools would still have weak descriptions. The same selection bug would reappear."}, {"letter": "D", "text": "Expand MCP tool descriptions to detail capabilities and outputs—e.g., \"Builds dependency graph showing direct imports, transitive dependencies, and cycles.\"", "correct": true, "explanation": "Correct. Tool selection is driven by the descriptions the model sees. A one-line description like 'Analyzes dependency graph' loses against Grep's rich description. Beef up the MCP tool's description."}], "correct": "D", "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "group": "J"}, {"id": "m17", "domain": 2, "scenario": null, "situation": "An engineer asks the agent to find all callers of a function before removing it. The function is defined in a core library but is also exposed through wrapper modules that rename the function for domain-specific use (e.g., calculateTax in the library becomes computeOrderTax in the orders module).", "question": "What exploration strategy will most reliably identify all callers?", "options": [{"letter": "A", "text": "Read the library and wrapper modules to identify all exposed names for the function, then Grep for each name across the codebase.", "correct": true, "explanation": "Correct. You have to enumerate every name the function is exposed under — otherwise renamed wrappers hide callers. Read the relevant modules, gather all aliases, then grep for each."}, {"letter": "B", "text": "Use Grep to find all files that import from the library or wrapper modules, then read each file to check whether it uses the function.", "correct": false, "explanation": "Misses dynamic imports, re-exports, and indirect call chains. Also scales poorly."}, {"letter": "C", "text": "Use Grep to search for the function's original name across the codebase.", "correct": false, "explanation": "The original name won't appear anywhere a wrapper has renamed it — you'd miss a significant class of callers."}, {"letter": "D", "text": "Search for the function name in project documentation to understand intended usage patterns and navigate to documented integration points.", "correct": false, "explanation": "Docs are incomplete and often stale. You can't safely remove a function based on documentation-level evidence."}], "correct": "A", "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "group": "G"}, {"id": "m27", "domain": 2, "scenario": null, "situation": "After adding an MCP server with specialized code refactoring tools (`extract_function`, `rename_variable`, `inline_function`), you notice the agent still uses basic text manipulation via Write and Bash sed commands for refactoring tasks. The MCP server is connected and healthy. Examining the configuration, you find each MCP tool has a minimal description like \"`extract_function`: extracts a function from code.\"", "question": "What's the most effective way to improve adoption of the MCP refactoring tools?", "options": [{"letter": "A", "text": "Implement a request classifier that detects refactoring intent and automatically routes those requests to the MCP server before the agent processes them.", "correct": false, "explanation": "A pre-classifier is a separate system to maintain and can misroute. The cheaper fix is to make the tool descriptions strong enough that the agent picks them on its own."}, {"letter": "B", "text": "Remove the Write tool from the agent's configuration for refactoring sessions so it must use the MCP tools for code modifications.", "correct": false, "explanation": "Stripping general tools forces adoption by subtraction, and breaks legitimate Write use cases."}, {"letter": "C", "text": "Accept this as expected behavior since simpler tools like sed are more predictable than specialized refactoring tools.", "correct": false, "explanation": "Capitulating defeats the purpose of integrating refactoring tools. The integration is fine — the descriptions are the problem."}, {"letter": "D", "text": "Enhance the MCP tool descriptions to explain when each tool is preferable to text manipulation and clarify expected inputs and outputs.", "correct": true, "explanation": "Correct. Tool selection is driven by the descriptions Claude sees. When the MCP tools say 'extracts a function from code' and Write/sed come with rich documentation, Claude picks Write/sed. Beef up the descriptions."}], "correct": "D", "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "group": "J"}, {"id": "m29", "domain": 2, "scenario": null, "situation": "Your agent needs to insert a new helper function into the middle of a 150-line utility module, between two existing functions. The Edit tool fails because its `old_string` parameter cannot find unique text to match — the file has repetitive docstrings, variable names, and structural patterns.", "question": "What's the most reliable way to complete this insertion?", "options": [{"letter": "A", "text": "Use Edit with an extremely long `old_string` capturing 30+ lines of context to guarantee uniqueness", "correct": false, "explanation": "Long, brittle match strings frequently miss due to whitespace or minor edits and produce confusing failures."}, {"letter": "B", "text": "Use Edit's `replace_all` parameter to target a common pattern and embed the new function in the replacement text", "correct": false, "explanation": "`replace_all` would mutate every occurrence of the pattern — corrupting the whole file."}, {"letter": "C", "text": "Use Bash to append the function definition to the end of the file using heredoc syntax", "correct": false, "explanation": "Appending puts the function at the wrong location. The requirement is to insert between two existing functions."}, {"letter": "D", "text": "Use Read to load the file, add the function at the appropriate location, then Write the updated file", "correct": true, "explanation": "Correct. When Edit's unique-match contract can't be satisfied in a repetitive file, fall back to Read → modify in memory at the intended line → Write the full file back."}], "correct": "D", "task_id": "2.5", "objective": "Select and apply built-in tools (Read, Write, Edit, Bash, Grep, Glob) effectively", "group": "F"}, {"id": "m38", "domain": 2, "scenario": null, "situation": "Production logs reveal inconsistent error handling: when `lookup_order` fails, the agent sometimes retries 5+ times (wasteful when the order ID doesn't exist), sometimes escalates immediately (premature for temporary network issues), and sometimes asks users for clarification (inappropriate when the issue is a backend permission error). Investigation shows your MCP tool returns uniform error responses: {\"isError\": true, \"content\": [{\"type\": \"text\", \"text\": \"Operation failed\"}]}. The agent cannot distinguish between error types.", "question": "What's the most effective improvement?", "options": [{"letter": "A", "text": "Enhance error responses with structured metadata: include errorCategory (transient/validation/permission), isRetryable boolean, and a description of what caused the failure.", "correct": true, "explanation": "Correct. Give the agent the information it needs to make the right decision: category, retryability, and a human-readable cause. That replaces guessing with deterministic policy."}, {"letter": "B", "text": "Create an `analyze_error` MCP tool the agent calls after any failure to determine the error category and recommended action.", "correct": false, "explanation": "Adds an extra round-trip for something the original tool already knows. Put the metadata in the original response."}, {"letter": "C", "text": "Implement retry logic with exponential backoff in your MCP server for all errors, returning to the agent only after retries are exhausted.", "correct": false, "explanation": "Blanket retries hurt on permanent errors (order not found) and hide useful distinctions from the agent."}, {"letter": "D", "text": "Add few-shot examples to the system prompt demonstrating how to interpret error message patterns and select appropriate responses for each.", "correct": false, "explanation": "If the tool returns 'Operation failed' for every failure, no amount of few-shots can extract category info that isn't there."}], "correct": "A", "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "group": "J"}, {"id": "m43", "domain": 2, "scenario": null, "situation": "When implementing your `lookup_order` MCP tool, the backend sometimes returns errors (e.g., \"Order not found\" or temporary database failures).", "question": "What is the correct pattern for communicating these errors back to the agent?", "options": [{"letter": "A", "text": "Log the error server-side and return an empty result to avoid confusing the model", "correct": false, "explanation": "Returning empty successes makes the agent think no data exists, rather than that something went wrong — a different and worse confusion."}, {"letter": "B", "text": "Return the error message in the tool result content with the isError flag set to true", "correct": true, "explanation": "Correct. MCP's designed pattern: put the error text in the content field and mark isError=true. Claude sees both the failure flag and a readable message to reason about."}, {"letter": "C", "text": "Throw an exception from the tool handler so the agent framework can catch and log it", "correct": false, "explanation": "Uncaught exceptions break the tool protocol and don't give the model anything to reason with."}, {"letter": "D", "text": "Return a success response with a \"status\" field indicating the error type", "correct": false, "explanation": "Ad-hoc 'status' fields vary across tools and the model has no standard way to interpret them. isError is the standard."}], "correct": "B", "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "group": "J"}, {"id": "m44", "domain": 2, "scenario": null, "situation": "Your `process_refund` tool returns two types of errors: technical errors (\"503 Service Unavailable\", \"Connection timeout\") that are transient (5% of calls), and business errors (\"Order exceeds 30-day return window\", \"Item already refunded\") that are permanent (12% of calls). Monitoring shows the agent wastes 3-4 turns retrying business errors that can never succeed. Currently, both error types return only a plain text message to Claude.", "question": "What's the most effective way to reduce wasted retries while improving customer-facing response quality?", "options": [{"letter": "A", "text": "Return structured error responses with retryable: false for business errors and a customer-friendly explanation for Claude to use.", "correct": true, "explanation": "Correct. A retryable flag tells Claude deterministically 'don't retry,' and a ready-made customer-friendly message improves the outgoing reply. Fixes both problems at once."}, {"letter": "B", "text": "Add few-shot examples showing how to distinguish retryable from non-retryable errors by parsing error message text.", "correct": false, "explanation": "Relying on the model to text-parse error strings is fragile and exactly the instability you see today."}, {"letter": "C", "text": "Add a `check_refund_eligibility` tool that must be called before `process_refund` to prevent business rule violations.", "correct": false, "explanation": "Useful in principle but adds a round-trip to every refund to guard against 12% of cases, and doesn't help when `process_refund` still fails for other business reasons."}, {"letter": "D", "text": "Implement automatic retry logic at the tool level for technical errors only, passing business errors to Claude without retries.", "correct": false, "explanation": "Hiding transient retries inside the tool can mask latency and takes the model out of the loop on recovery decisions. Business errors also still arrive as a plain string, so the customer-facing response doesn't improve."}], "correct": "A", "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "group": "J"}, {"id": "m55", "domain": 2, "scenario": null, "situation": "Your pipeline uses a tool called `extract_metadata` with a JSON schema for paper details. You've also defined `lookup_citations` and `verify_doi` tools for enrichment. During testing, you notice that when users include requests like \"extract the metadata and tell me how cited it is,\" Claude sometimes calls `lookup_citations` first, which fails because it needs the DOI that `extract_metadata` would provide.", "question": "What's the most effective way to ensure structured metadata extraction happens first?", "options": [{"letter": "A", "text": "Set `tool_choice` to \"any\" so Claude must use a tool, combined with system prompt instructions prioritizing `extract_metadata`.", "correct": false, "explanation": "'any' forces a tool call but doesn't force the right one. You'd still see `lookup_citations` picked first."}, {"letter": "B", "text": "Set `tool_choice` to \"auto\" and reorder the tool definitions so `extract_metadata` appears first in the tools array, since Claude prioritizes earlier-listed tools.", "correct": false, "explanation": "There's no documented ordering preference to rely on; this assumes behavior that isn't contractual."}, {"letter": "C", "text": "Set `tool_choice` to {\"type\": \"tool\", \"name\": \"`extract_metadata`\"} and process the enrichment requests in subsequent turns after receiving the extracted metadata.", "correct": true, "explanation": "Correct. `tool_choice`=specific-tool deterministically forces `extract_metadata` on the first turn. Then you hand control back to 'auto' to let the model use citations/DOI enrichment with the metadata in context."}, {"letter": "D", "text": "Set `tool_choice` to {\"type\": \"tool\", \"name\": \"`extract_metadata`\"} for every API call in the pipeline, ensuring Claude always extracts metadata before any enrichment can occur.", "correct": false, "explanation": "Pinning `extract_metadata` on every call prevents the enrichment tools from ever being used."}], "correct": "C", "task_id": "2.3", "objective": "Distribute tools appropriately across agents and configure tool choice", "group": "B"}, {"id": "f2-001", "domain": 2, "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "situation": "A cancel_subscription MCP tool rejects a cancellation because the account is locked in a legal hold, a policy condition that will not change no matter how the request is retried or reformatted. The engineer must choose between labeling this a validation error or a business error.", "question": "Which choice is correct, and why?", "options": [{"letter": "A", "text": "It is a business error because the request itself is well-formed and the rejection stems from a policy rule about the account's state rather than malformed input.", "correct": true, "explanation": "The request is well-formed and contains a valid account ID; the rejection is caused by a business policy (legal hold) on the account's state, not by malformed or invalid input. This is the hallmark of a business error."}, {"letter": "B", "text": "Business error, because the legal hold check occurs in a separate service after request validation, so the rejection is a business rule violation, not a schema issue.", "correct": false, "explanation": "The fact that the legal hold check runs in a separate service after request validation is an implementation detail and does not change the error category. Classification should be based on the nature of the failure (policy violation vs. malformed input), not on internal service boundaries."}, {"letter": "C", "text": "Validation fails because the account ID in the request is the specific field that, when evaluated against the account's legal hold status, causes the rejection.", "correct": false, "explanation": "Although the account ID field triggers the evaluation against the legal hold status, the ID itself is valid and the rejection is due to the account's policy state, not any structural or format issue with the input. This makes it a business error, not a validation error."}, {"letter": "D", "text": "Validation error, because any rejection after initial schema checks indicates the input, when checked against account state, does not pass full system validation.", "correct": false, "explanation": "Passing initial schema checks and then being rejected due to a policy rule is a business rejection, not a validation error. Not all post-schema rejections are validation failures; business rules are enforced after schema validation is successful."}], "correct": "A", "select": 1, "group": "J"}, {"id": "f2-002", "domain": 2, "task_id": "2.4", "objective": "Integrate MCP servers into Claude Code and agent workflows", "situation": "An architect is rolling out a GitHub MCP server for the whole engineering team. Every teammate has their own GitHub personal access token, and the config must be checked into the repo without ever committing a real secret.", "question": "How should the architect configure this?", "options": [{"letter": "A", "text": "Add the server with project scope, then add .mcp.json to .gitignore so the checked-in repository never actually contains the shared server configuration file", "correct": false, "explanation": "Ignoring .mcp.json removes the shared configuration entirely, meaning no teammate gets the server automatically and the stated goal of a checked-in, team-wide config is not met."}, {"letter": "B", "text": "Add the server with project scope in .mcp.json, and set the header to Authorization: Bearer ${GITHUB_TOKEN} so each teammate's environment supplies the value at connection", "correct": true, "explanation": "Project scope stores the server in .mcp.json for team-wide sharing via version control, and ${GITHUB_TOKEN} expansion pulls the value from each user's own environment at connection time, so no secret is ever committed."}, {"letter": "C", "text": "Add the server with user scope in ~/.claude.json, and paste each teammate's literal token value into the shared header field before committing that file to the repo", "correct": false, "explanation": "User scope stores the entry in ~/.claude.json, which is private to one machine and not shared with the team, and pasting literal tokens defeats the goal of never committing secrets."}, {"letter": "D", "text": "Add the server with local scope, then have every teammate individually edit their own copy of .mcp.json to insert their personal token in place of a placeholder", "correct": false, "explanation": "Local scope is private to the current project on one machine and is not checked into version control at all, so it cannot serve as the shared team configuration the scenario requires."}], "correct": "B", "select": 1, "group": "J"}, {"id": "f2-003", "domain": 2, "task_id": "2.3", "objective": "Distribute tools appropriately across agents and configure tool choice", "situation": "A synthesis agent's job is to combine findings that a separate research agent has already gathered into a final answer. Because it shares a tool registry with the research agent, the synthesis agent also has access to a web_search tool. During testing, the synthesis agent repeatedly calls web_search mid-synthesis instead of using the findings already provided to it, producing inconsistent citations.", "question": "What is the best explanation and fix for this behavior?", "options": [{"letter": "A", "text": "The research agent is passing incomplete findings, so the synthesis agent searches for missing information; updating the research agent to include complete source lists would prevent the extra calls.", "correct": false, "explanation": "There is no evidence in the scenario that the research agent's findings are incomplete; the synthesis agent is simply misusing a tool it should not have access to. Adding more data to the research output would not discourage the synthesis agent from invoking web_search when the tool remains available."}, {"letter": "B", "text": "Agents tend to misuse tools outside their specialization when given access to them; web_search should be removed from the synthesis agent's tool set and left with the research agent.", "correct": true, "explanation": "Agents are prone to misusing tools that fall outside their designated role because the availability of a tool signals it is an appropriate action, even when it conflicts with the agent's purpose. Removing web_search from the synthesis agent's tool set enforces role‑based access and eliminates the extraneous calls."}, {"letter": "C", "text": "The synthesis agent's temperature is likely too high, causing it to explore tool calls rather than follow its findings; lowering it to a more focused value like 0.2 would reduce unnecessary searches.", "correct": false, "explanation": "Temperature controls the randomness of token selection, not the agent's decision to call a tool outside its scope. The behavior stems from over‑broad tool access rather than sampling variability, so lowering temperature would not resolve the underlying scoping problem."}, {"letter": "D", "text": "The web_search tool's description is too vague for the synthesis agent to interpret correctly; rewriting it with guidance that it is a research tool and should not be used during synthesis would prevent the extra calls.", "correct": false, "explanation": "The root issue is not the quality of the tool description but rather the inclusion of an out‑of‑scope tool in the agent's registry. Even a perfectly written description would not prevent misuse because the tool itself is inappropriate for the synthesis agent's responsibilities."}], "correct": "B", "select": 1, "group": "B"}, {"id": "f2-004", "domain": 2, "task_id": "2.4", "objective": "Integrate MCP servers into Claude Code and agent workflows", "situation": "A contributor clones a team repository that includes a checked-in .mcp.json defining a deploy-tools server. They open the project in Claude Code for the first time and run claude mcp list.", "question": "What should they expect to see for deploy-tools before they take any further action?", "options": [{"letter": "A", "text": "It connects immediately and silently, because checking a server into .mcp.json is itself treated as implicit approval from every contributor who clones the repo", "correct": false, "explanation": "Being checked into version control does not grant automatic trust; each contributor must approve project-scoped servers themselves before they connect."}, {"letter": "B", "text": "It fails to load at all, because project-scoped servers only activate for the teammate who originally added them via claude mcp add --scope project", "correct": false, "explanation": "Project scope is specifically designed so the configuration is shared with the whole team via .mcp.json, not restricted to only the original author's machine."}, {"letter": "C", "text": "It is renamed automatically with a numeric suffix, because Claude Code assumes any newly cloned .mcp.json server name conflicts with an existing one", "correct": false, "explanation": "Automatic renaming with a suffix happens for name collisions during imports from Claude Desktop, not as a default behavior for freshly cloned project-scoped servers."}, {"letter": "D", "text": "It appears as pending approval, because project-scoped servers from .mcp.json require the contributor to explicitly approve them before Claude Code connects", "correct": true, "explanation": "For security reasons, Claude Code prompts for approval before using project-scoped servers from .mcp.json, so a freshly cloned repo's server shows as pending approval until the contributor reviews and approves it."}], "correct": "D", "select": 1, "group": "J"}, {"id": "f2-005", "domain": 2, "task_id": "2.5", "objective": "Select and apply built-in tools (Read, Write, Edit, Bash, Grep, Glob) effectively", "situation": "Claude Code is fixing a bug and wants to reproduce it first by running the project's test suite and capturing the failing stack trace before making any code changes.", "question": "Which tool should Claude use to run the suite and view its output?", "options": [{"letter": "A", "text": "Grep, to search the codebase for the word test and treat matching file names as evidence that the suite has already passed", "correct": false, "explanation": "Grep only searches file contents for a pattern; finding files whose names or contents mention the word test says nothing about whether the suite currently passes or fails, and cannot execute anything."}, {"letter": "B", "text": "Glob, to list all files matching **/*.test.* and treat the presence of test files as confirmation that the suite runs cleanly", "correct": false, "explanation": "Glob only lists file paths matching a naming pattern; the mere existence of test files provides no information about whether they currently pass when run, and Glob cannot execute code."}, {"letter": "C", "text": "Bash, to invoke the project's test runner command and capture its stdout and stderr, including the stack trace, in the result", "correct": true, "explanation": "Bash executes terminal commands such as the project's test runner and returns its output, including any stack trace printed to stdout or stderr, which is exactly what is needed to reproduce and observe the failure."}, {"letter": "D", "text": "Read, to open the test runner's configuration file and infer the current pass or fail status of the suite from its settings", "correct": false, "explanation": "A test runner's configuration file describes how tests are set up to run, not the current pass or fail outcome; reading it cannot substitute for actually executing the suite."}], "correct": "C", "select": 1, "group": "F"}, {"id": "f2-006", "domain": 2, "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "situation": "A single tool, analyze_document, extracts data points, produces summaries, and verifies claims against a source. Users report it inconsistently performs only one of these behaviors for ambiguous requests.", "question": "Which redesign best addresses this?", "options": [{"letter": "A", "text": "Split the tool into extract_data_points, summarize_content, and verify_claim_against_source, each with a narrow contract.", "correct": true, "explanation": "Splitting a generic multi-purpose tool into purpose-specific tools with defined input/output contracts removes the ambiguity about which behavior a call should trigger."}, {"letter": "B", "text": "Merge the tool with unrelated tools into one larger tool so the model faces fewer total choices overall.", "correct": false, "explanation": "Merging with unrelated tools increases the surface area of a single tool's responsibilities, worsening overlapping-purpose confusion rather than resolving it."}, {"letter": "C", "text": "Keep the single tool but instruct the model, in the system prompt, to always call it three separate times per request.", "correct": false, "explanation": "Forcing three calls per request wastes invocations on tasks that need only one behavior and does not resolve the ambiguity about which behavior was intended."}, {"letter": "D", "text": "Add a required mode enum parameter to the existing tool, but leave its single overarching description unchanged.", "correct": false, "explanation": "Adding a mode parameter without updating the description still leaves the model guessing at the tool's overall purpose and when each mode applies."}], "correct": "A", "select": 1, "group": "J"}, {"id": "f2-007", "domain": 2, "task_id": "2.4", "objective": "Integrate MCP servers into Claude Code and agent workflows", "situation": "A shared .mcp.json points a stdio server's args at ${API_REGION:-us-east-1}.", "question": "On a machine where the API_REGION environment variable is unset, what value does Claude Code pass to the server?", "options": [{"letter": "A", "text": "The literal string us-east-1, because the ${VAR:-default} syntax falls back to the default when the variable is not set", "correct": true, "explanation": "${VAR:-default} expands to the variable's value when it is set and to the text after :- when it is not, so with API_REGION unset the server receives the literal us-east-1. Claude Code's MCP documentation lists args among the fields where this expansion happens, alongside command, env, url, and headers."}, {"letter": "B", "text": "An empty string, because Claude Code always expands an unset variable to blank rather than substituting the trailing default text", "correct": false, "explanation": "Claude Code does not blank out a referenced variable that carries a default; the whole purpose of the :- form is to substitute the text after :- instead of an empty value. Nothing in the documented expansion rules yields an empty string here."}, {"letter": "C", "text": "The literal text ${API_REGION:-us-east-1}, because default-value expansion only applies inside the env block and not in args", "correct": false, "explanation": "Expansion is not confined to the env block. The documented expansion locations are command, args, env, url, and headers, so a value inside args is substituted like any other."}, {"letter": "D", "text": "A parse failure, because Claude Code requires every referenced environment variable to be set even when a default is supplied", "correct": false, "explanation": "No failure occurs: a default was supplied, which is exactly what :- is for. Even with no default at all the file still parses — Claude Code loads the config, reports a missing-variable warning for that server in claude mcp list, and passes the unexpanded ${VAR} text through."}], "correct": "A", "select": 1, "group": "J"}, {"id": "f2-008", "domain": 2, "task_id": "2.3", "objective": "Distribute tools appropriately across agents and configure tool choice", "situation": "A team builds a single agent that handles research, drafting, fact-checking, and formatting for a report-generation pipeline. The agent is given all 18 tools used across these functions. Reviewers notice it frequently picks a plausible-but-wrong tool, or stalls comparing similar options, even though each individual tool works correctly in isolation.", "question": "Which change is most likely to fix the selection accuracy problem?", "options": [{"letter": "A", "text": "Keep the single agent but raise its max_tokens limit so it has more room to reason through the tool list", "correct": false, "explanation": "More output budget does not address the underlying problem of an oversized, overlapping tool surface; the model still has to discriminate among 18 similar options on every turn."}, {"letter": "B", "text": "Rewrite the 18 tool descriptions to be shorter so the model spends less time reading each one before deciding", "correct": false, "explanation": "Shortening descriptions reduces prompt tokens but does not reduce the number of competing tools the model must choose among, which is the actual driver of misselection at 18 tools."}, {"letter": "C", "text": "Split the work across specialized subagents so each one is exposed to only the 4-5 tools relevant to its own role", "correct": true, "explanation": "Decision complexity scales with the number of candidate tools an agent must discriminate between; scoping each subagent to the 4-5 tools its role actually needs restores selection reliability, which is the core motivation for distributing tools across specialized agents rather than one generalist."}, {"letter": "D", "text": "Sort the 18 tools alphabetically in the tools array so the model scans them in a consistent order", "correct": false, "explanation": "Ordering the array changes presentation only; it does not shrink the set of tools competing for selection on each turn, so accuracy would not meaningfully improve."}], "correct": "C", "select": 1, "group": "B"}, {"id": "f2-009", "domain": 2, "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "situation": "A support agent has both search_web and fetch_webpage_results as separate tools. Testing shows the model almost always calls search_web, even for tasks better suited to fetch_webpage_results.", "question": "The system prompt contains \"Always prefer searching for the most up to date information.\" What most likely explains the bias?", "options": [{"letter": "A", "text": "The search tool carries a lower internal temperature value in its schema, making it statistically favored during sampling.", "correct": false, "explanation": "Tool schemas do not carry a temperature value; temperature is a model-level sampling parameter unrelated to individual tool definitions."}, {"letter": "B", "text": "Keyword-sensitive system prompt wording creates an unintended association that overrides the more accurate tool description.", "correct": true, "explanation": "The phrase 'always prefer searching' repeatedly emphasizes the word 'search,' creating a keyword-driven bias that overrides the more suitable tool's description."}, {"letter": "C", "text": "The fetch tool cannot be selected during parallel tool use, so it is filtered out before the model can consider it.", "correct": false, "explanation": "Parallel tool use restrictions are not described here, and there is no indication the fetch tool is structurally excluded from the candidate set."}, {"letter": "D", "text": "The fetch tool's description exceeds a fixed token threshold, causing the model to systematically avoid tools of that length.", "correct": false, "explanation": "There is no fixed token-length threshold that categorically excludes longer tool descriptions from being selected by the model."}], "correct": "B", "select": 1, "group": "J"}, {"id": "f2-010", "domain": 2, "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "situation": "An MCP server exposes update_inventory. When the request body is missing the required sku field, the server currently returns isError: true with errorCategory: \"transient\" and isRetryable: true. An agent retries the exact same malformed request repeatedly and never succeeds.", "question": "What is wrong with this error classification?", "options": [{"letter": "A", "text": "isRetryable should remain true because validation errors are inherently transient once the missing field is eventually supplied by a later, unrelated request", "correct": false, "explanation": "The field being supplied in some unrelated future request doesn't make this particular failed call retryable; the current request itself needs correction before it can succeed."}, {"letter": "B", "text": "The tool should have used a JSON-RPC protocol-level error instead of a tool result, since any error involving a missing field must always be surfaced at the protocol layer", "correct": false, "explanation": "Whether an error is a missing-field validation problem doesn't dictate protocol-level vs. tool-result reporting; a well-formed request with invalid field values is exactly the kind of business-logic-adjacent case reported via isError in the tool result."}, {"letter": "C", "text": "The categorization is correct as written, since the agent's repeated retries are the expected and desired behavior for any error carrying isRetryable: true", "correct": false, "explanation": "The scenario explicitly shows the agent wasting retries on a request that can never succeed unmodified, which demonstrates the classification is causing the exact problem structured metadata is meant to prevent."}, {"letter": "D", "text": "A missing required field is a validation error, not a transient one, so marking it retryable causes the agent to resend an identically malformed request instead of correcting the input", "correct": true, "explanation": "A missing required field is a validation problem the caller must fix by changing the request, not a condition that resolves with time; labeling it transient and retryable misleads the agent into resending the same broken payload."}], "correct": "D", "select": 1, "group": "J"}, {"id": "f2-011", "domain": 2, "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "situation": "A team observes that adding 'If in doubt, use the search tool' to the system prompt caused the model to call search_web even when the lookup_internal_docs tool was more appropriate.", "question": "What does this scenario illustrate?", "options": [{"letter": "A", "text": "System prompts can significantly influence tool selection, and explicit instructions may override the model's assessment of which tool is most appropriate.", "correct": true, "explanation": "This scenario illustrates that the system prompt has a strong influence on tool selection. Even when a tool's description makes it more appropriate, an explicit directive like 'If in doubt, use the search tool' can cause the model to disregard its normal tool-evaluation logic and favor the instructed tool. Anthropic's documentation emphasizes that prompts can bias completions and recommends clear, unambiguous instructions."}, {"letter": "B", "text": "The lookup_internal_docs tool must have a malformed JSON schema, since that is the only way a tool can be excluded from selection.", "correct": false, "explanation": "Tool selection is not solely based on JSON schema validity; the model considers the full context including tool descriptions and the prompt. Here, the system prompt's explicit instruction overrode the descriptions, not a schema defect."}, {"letter": "C", "text": "System prompts have no measurable effect on tool selection; the behavior must be caused by a defect in the model.", "correct": false, "explanation": "System prompts do influence tool selection, as they act as a persistent directive layer that shapes model behavior. Anthropic's documentation notes that prompts can bias completions, so the observed behavior is expected when the prompt contains explicit instructions, not a model defect."}, {"letter": "D", "text": "The word 'search' appearing anywhere in a tool's name always takes absolute priority over any other tool regardless of prompt content.", "correct": false, "explanation": "The model considers the complete tool description and the entire prompt context, not just keywords in tool names. This scenario shows the effect of an explicit instruction in the prompt, not an absolute naming rule."}], "correct": "A", "select": 1, "group": "J"}, {"id": "f2-012", "domain": 2, "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "situation": "Two tools, create_ticket and escalate_ticket, both include the phrase \"handles customer issues\" in their descriptions. The model regularly calls create_ticket for cases that should instead be escalated.", "question": "Which revision best resolves the ambiguity?", "options": [{"letter": "A", "text": "Instruct the model, in the system prompt, to always call both tools together on every customer issue regardless of context.", "correct": false, "explanation": "Calling both tools on every issue wastes invocations on cases that don't need escalation and does not address why the model can't distinguish the two triggers."}, {"letter": "B", "text": "Add a short delay before escalate_ticket executes so the model has more time to reconsider which tool it selected.", "correct": false, "explanation": "Execution timing has no bearing on the reasoning the model uses to select a tool; the ambiguity exists before either tool is invoked."}, {"letter": "C", "text": "Combine both tools' functionality into create_ticket and pass an escalate flag, without changing any description text.", "correct": false, "explanation": "Merging functionality without updating any description text leaves the model with the same ambiguous signal it had before, now hidden behind an undocumented flag."}, {"letter": "D", "text": "Rewrite each description to state its specific trigger condition, and note explicitly when to use the other tool instead.", "correct": true, "explanation": "Stating the specific trigger condition for each tool and cross-referencing the alternative gives the model the differentiating signal it was missing."}], "correct": "D", "select": 1, "group": "J"}, {"id": "f2-013", "domain": 2, "task_id": "2.5", "objective": "Select and apply built-in tools (Read, Write, Edit, Bash, Grep, Glob) effectively", "situation": "An architect asks Claude Code to standardize indentation across a 900-line generated file where nearly every line's leading whitespace is inconsistent, affecting the entire file rather than a few isolated lines. Claude has already read the file once earlier and no one else has modified it since.", "question": "Which approach is most appropriate for applying this change?", "options": [{"letter": "A", "text": "Call Glob to match the file's own path, then rely on Glob's sort-by-modification-time behavior to normalize the file's whitespace", "correct": false, "explanation": "Glob locates files by pattern and can sort results by metadata such as modification time, but it has no capability to edit file contents or normalize whitespace. Finding the file path is not a formatting operation."}, {"letter": "B", "text": "Call Grep with output mode content to retrieve every line of the file, since retrieving lines through Grep also rewrites them in place", "correct": false, "explanation": "Grep is a read-only search tool that returns matching lines or content, but it does not rewrite or modify files in place. Even if it could retrieve every line, retrieving content is not the same as applying an indentation change."}, {"letter": "C", "text": "Issue one Edit call per line of the file, each with old_string set to that single line's current indentation and text", "correct": false, "explanation": "A line-by-line Edit strategy would require hundreds of calls and greatly increases the chance of mismatches, conflicts, and inconsistent formatting. When a change affects almost every line, rewriting the affected file content in one operation is more reliable and efficient."}, {"letter": "D", "text": "Read the file again to confirm current content, then call Write with the entire re-indented file content in a single call", "correct": true, "explanation": "This follows the standard read-modify-write pattern for whole-file reformatting: confirm the current file state with Read, perform the re-indentation, then use a single Write call to replace the file. A single whole-file write is more appropriate than hundreds of per-line edits when nearly every line is affected."}], "correct": "D", "select": 1, "group": "F"}, {"id": "f2-014", "domain": 2, "task_id": "2.4", "objective": "Integrate MCP servers into Claude Code and agent workflows", "situation": "A team built a custom MCP server exposing a find_symbol_usages tool that is far more accurate than text search, but the agent keeps reaching for the built-in Grep tool instead.", "question": "The tool's current description is just \"Finds symbol usages.\" What change is most likely to fix this?", "options": [{"letter": "A", "text": "Rewrite the description to explain what it returns and when it beats text search, such as resolving usages across renamed imports and generated code grep cannot match", "correct": true, "explanation": "A thin, generic description gives the agent no reason to prefer the MCP tool; explaining its concrete capabilities and outputs in detail is the documented way to prevent it from defaulting to a familiar built-in tool like Grep."}, {"letter": "B", "text": "Rename the tool from find_symbol_usages to grep so the agent recognizes it as a drop-in upgrade for its existing habit of reaching for text search", "correct": false, "explanation": "Renaming the tool to collide with a built-in tool's name does not change what information the agent has about its capabilities, and reserved/ambiguous naming causes confusion rather than better tool selection."}, {"letter": "C", "text": "Set the tool's permission mode to require manual approval on every single call so the agent is forced to weigh it before falling back to Grep", "correct": false, "explanation": "Requiring approval on every call adds friction to using the tool but does nothing to inform the agent's judgment about which tool is more capable, so Grep would still often be chosen first."}, {"letter": "D", "text": "Remove Grep from the list of available built-in tools entirely so the agent has no remaining alternative but to call the MCP tool for every search it runs", "correct": false, "explanation": "Disabling a core built-in tool for the whole session is a heavy-handed workaround that breaks unrelated tasks relying on Grep, rather than addressing the actual cause: an uninformative tool description."}], "correct": "A", "select": 1, "group": "J"}, {"id": "f2-015", "domain": 2, "task_id": "2.3", "objective": "Distribute tools appropriately across agents and configure tool choice", "situation": "A legal-document analysis agent has a single retrieve_clause tool that can pull arbitrary text ranges from any uploaded file by byte offset, which the model frequently misuses to grab unrelated or malformed spans. The team wants to replace it with a constrained alternative that only ever returns whole, well-defined clauses.", "question": "Which redesign best follows the pattern of replacing a generic tool with a constrained one?", "options": [{"letter": "A", "text": "Keep retrieve_clause unchanged and add a second agent whose only job is to double-check the byte ranges after retrieval", "correct": false, "explanation": "Adding a checking agent works around the problem after the fact rather than constraining the tool itself, and increases pipeline complexity without fixing the root cause."}, {"letter": "B", "text": "Replace retrieve_clause with a get_clause_by_id tool that only accepts a validated clause identifier from a pre-parsed clause index", "correct": true, "explanation": "Swapping the unconstrained byte-offset tool for get_clause_by_id, which only accepts validated identifiers from a pre-parsed index, mirrors the documented pattern of replacing a generic tool with a narrower one that structurally prevents the misuse."}, {"letter": "C", "text": "Keep retrieve_clause but double the number of example byte-offset calls in its description so the model learns better offsets", "correct": false, "explanation": "More examples may improve offset guesses somewhat, but the tool still accepts arbitrary byte ranges, so the structural cause of malformed spans remains unaddressed."}, {"letter": "D", "text": "Give the agent broader access by also adding a raw_file_read tool so it can cross-check offsets against the full document", "correct": false, "explanation": "Adding another broad, unconstrained tool increases the surface for misuse and decision complexity rather than replacing the problematic tool with a scoped alternative."}], "correct": "B", "select": 1, "group": "B"}, {"id": "f2-016", "domain": 2, "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "situation": "A platform team is designing error responses for a fleet of internal MCP tools. One engineer proposes that every tool failure, regardless of cause, return the same generic text \"Operation failed\" with isError: true, arguing this keeps the interface simple for tool authors.", "question": "What is the strongest architectural objection to this proposal?", "options": [{"letter": "A", "text": "A uniform generic message gives the agent no basis for choosing among retrying, adjusting input, or escalating, so it cannot make an appropriate recovery decision for each failure.", "correct": true, "explanation": "When all errors return the same generic message, the agent cannot distinguish between transient, input, or system failures, leaving it no basis to retry, adjust, or escalate. This prevents appropriate recovery decisions and undermines the agent's ability to handle errors gracefully."}, {"letter": "B", "text": "Returning a constant error string for every failure adds metadata overhead that pushes the total block size beyond the MCP protocol's maximum content length, so the server rejects the tool result as non-compliant.", "correct": false, "explanation": "The MCP protocol does not define a maximum content length that a generic error string like 'Operation failed' would exceed; metadata overhead is not a concern here because constant error messages do not increase message size beyond a trivial amount."}, {"letter": "C", "text": "The MCP specification requires every isError:true result to include a machine-parseable stack trace, so a generic text response without that structured data violates the protocol and is rejected by the platform.", "correct": false, "explanation": "The MCP specification does not mandate including a machine-parseable stack trace with isError:true results; it only requires that errors be indicated with isError:true in the content, without dictating the specific structured data format."}, {"letter": "D", "text": "Uniform error text prevents the server from ever setting isError:true because the MCP specification requires a unique diagnostic string to accompany the flag for each failure, so the tool cannot activate the error state.", "correct": false, "explanation": "The MCP specification does not require a unique diagnostic string for each failure; isError:true can be set with the same text repeatedly, and while this is poor design, it does not prevent the server from activating the error state."}], "correct": "A", "select": 1, "group": "J"}, {"id": "f2-017", "domain": 2, "task_id": "2.5", "objective": "Select and apply built-in tools (Read, Write, Edit, Bash, Grep, Glob) effectively", "situation": "An architect asks Claude Code to count how many files in the src/components directory tree currently have no matching test file at all, as a first step in planning test coverage work. Assume that the project follows a consistent and reliable naming convention that pairs each component file with its corresponding test file (for example, Button.tsx is tested by Button.test.tsx).", "question": "Which approach most directly answers this without unnecessary file reads?", "options": [{"letter": "A", "text": "Use Read to open every file under src/components and manually inspect each one for an associated describe block before counting", "correct": false, "explanation": "Reading every component file is precisely the unnecessary file-read cost the task seeks to avoid, and manual inspection is error-prone. A describe block is an implementation detail of test code; looking inside source files does not reliably establish whether a separate test file exists."}, {"letter": "B", "text": "Use Glob to list component files and test files separately by their naming pattern, then compare the two path lists for gaps", "correct": true, "explanation": "Because the question explicitly states a consistent and reliable naming convention, glob patterns can match file paths by name without reading file contents. Listing component files (e.g., src/components/**/*.tsx excluding test files) and test files (e.g., src/components/**/*.test.tsx) and comparing the two lists directly reveals components lacking a corresponding test file. This static approach avoids unnecessary file reads and is the most direct method for this task. Note: If the naming convention were not reliable, a coverage tool would be more appropriate."}, {"letter": "C", "text": "Use Grep to search each component file's own contents for the word test, and count the files where that word never appears", "correct": false, "explanation": "Grepping component source for the string \"test\" reads file contents and checks an irrelevant signal: a component may mention \"test\" in comments or code without having a test file, or have a test file without containing that word. It does not directly determine matching filenames."}, {"letter": "D", "text": "Use Bash to run the full test suite and count how many components report zero assertions executed against them", "correct": false, "explanation": "Running the full test suite executes code and measures assertion or coverage behavior, not the presence or absence of a matching test file. A component may have a test file that is skipped, broken, or otherwise reports zero assertions, leading to false positives. This approach is also more expensive and indirect than a static filename comparison."}], "correct": "B", "select": 1, "group": "F"}, {"id": "f2-018", "domain": 2, "task_id": "2.3", "objective": "Distribute tools appropriately across agents and configure tool choice", "situation": "A synthesis agent frequently needs to confirm a single numeric claim (e.g. a statistic cited in a source) before including it in a final answer, but it is not equipped to resolve deeper factual disputes between conflicting sources.", "question": "Following the guidance on scoped cross-role tools for high-frequency needs, how should the team design this?", "options": [{"letter": "A", "text": "Give the coordinator agent a verify_fact tool but not the synthesis agent, so that the synthesis agent must send every numeric claim to the coordinator for verification and wait for the result.", "correct": false, "explanation": "Providing the verify_fact tool only to the coordinator forces the synthesis agent to send every numeric claim for verification, adding overhead and defeating the purpose of a scoped cross-role tool. The high-frequency single-claim checks should be handled directly by the synthesis agent."}, {"letter": "B", "text": "Give the synthesis agent the full research agent tool set, enabling it to independently verify any numeric claim and resolve source discrepancies without coordinator intervention.", "correct": false, "explanation": "Giving the synthesis agent the full research tool set grants broad, out-of-role capabilities that undermine the principle of scoped access. Allowing it to independently resolve source discrepancies bypasses the coordinator's role in handling complex conflicts."}, {"letter": "C", "text": "Give the synthesis agent no verification tools at all, and require every numeric claim to be manually checked by the coordinator before inclusion in the final answer, regardless of complexity.", "correct": false, "explanation": "Requiring every numeric claim to be manually checked by the coordinator, even for simple high-frequency cases, introduces unnecessary overhead and latency. A narrow verification tool on the synthesis agent would handle these efficiently without overburdening the coordinator."}, {"letter": "D", "text": "Give the synthesis agent a narrow verify_fact tool for quick single-claim checks, and have it refer cases with conflicting sources to the coordinator for deeper resolution.", "correct": true, "explanation": "Providing the synthesis agent with a narrow verify_fact tool addresses the high-frequency need for quick single-claim verification, while routing cases with conflicting sources to the coordinator ensures complex disputes are handled by a more capable agent. This aligns with the scoped cross-role tool pattern."}], "correct": "D", "select": 1, "group": "B"}, {"id": "f2-019", "domain": 2, "task_id": "2.5", "objective": "Select and apply built-in tools (Read, Write, Edit, Bash, Grep, Glob) effectively", "situation": "A developer wants Claude Code to update a deprecated log statement logger.warn(\"legacy-path\") that appears twice in the same file, in two different functions, where only one of the two occurrences should change. Claude issues an Edit call with old_string set to exactly that log statement and the call fails.", "question": "What is the correct next step?", "options": [{"letter": "A", "text": "Switch to Grep with the multiline flag to rewrite the matching line directly, since Grep can modify file contents once a match is found", "correct": false, "explanation": "Grep is read-only and searches file contents; it has no capability to write or modify files, so it cannot perform the edit."}, {"letter": "B", "text": "Set replace_all to true on the same Edit call so both occurrences update identically, then manually revert whichever one should have stayed", "correct": false, "explanation": "replace_all rewrites every occurrence, which would also change the log statement the developer wants to keep, and reverting one afterward risks losing track of which change was intentional."}, {"letter": "C", "text": "Widen old_string to include enough surrounding context to uniquely identify the intended occurrence, then retry Edit with that string", "correct": true, "explanation": "Edit requires old_string to match exactly once; when a string occurs more than once, supplying additional surrounding context that appears only around the target occurrence restores uniqueness and lets Edit apply cleanly."}, {"letter": "D", "text": "Call Write with only the new log line as content, expecting Write to merge that single line into the correct spot in the existing file", "correct": false, "explanation": "Write overwrites the entire file with exactly the content provided; passing only a single line would delete the rest of the file's contents rather than merging into it."}], "correct": "C", "select": 1, "group": "F"}, {"id": "f2-020", "domain": 2, "task_id": "2.5", "objective": "Select and apply built-in tools (Read, Write, Edit, Bash, Grep, Glob) effectively", "situation": "Claude Code is asked to create a brand-new configuration file, config/feature-flags.json, that does not yet exist anywhere in the repository, with content fully specified by the architect in the request.", "question": "Which tool should Claude use to create this file?", "options": [{"letter": "A", "text": "Read, followed immediately by Write, on the assumption that Write always requires a preceding Read regardless of whether the target file exists", "correct": false, "explanation": "Read fails on a file that does not exist, and the prior-read requirement for Write only applies to overwriting an existing file, so a Read step is both impossible and unnecessary here."}, {"letter": "B", "text": "Write, providing the full file path and the complete specified content, since creating a brand-new file does not require a prior read", "correct": true, "explanation": "Write creates a new file or overwrites an existing one with the full content provided; the requirement to have previously read the file only applies when overwriting a file that already exists, so a brand-new file can be created directly."}, {"letter": "C", "text": "Bash, using a heredoc to populate the file, on the assumption that Write cannot create files that do not already exist", "correct": false, "explanation": "Write is fully capable of creating new files directly; there is no need to route file creation through a Bash heredoc when the built-in Write tool already handles this case."}, {"letter": "D", "text": "Edit, providing an old_string that matches the contents of an empty file and a new_string containing the specified content", "correct": false, "explanation": "Edit requires old_string to match existing text in a file that Claude has already read; there is no file yet to read or match against, so Edit cannot be used to create a new file."}], "correct": "B", "select": 1, "group": "F"}, {"id": "f2-021", "domain": 2, "task_id": "2.4", "objective": "Integrate MCP servers into Claude Code and agent workflows", "situation": "A team has ten MCP servers connected, but one small internal server exposes two tools that Claude needs on nearly every single turn, and the team has noticed occasional delay while tool search resolves them.", "question": "What configuration change addresses this for just that one server?", "options": [{"letter": "A", "text": "Set alwaysLoad: true on that server's entry so its tools load into context at session start instead of being deferred behind tool search", "correct": true, "explanation": "Setting alwaysLoad to true on a specific server's configuration exempts just that server from tool-search deferral, loading its tools into context at session start while other servers remain deferred."}, {"letter": "B", "text": "Set ENABLE_TOOL_SEARCH=false globally so every server's tools load upfront and none are deferred behind a search step", "correct": false, "explanation": "Disabling tool search globally loads every server's tools upfront, which solves the delay for the one important server but unnecessarily bloats context with all nine other servers' tool definitions too."}, {"letter": "C", "text": "Move that server's entry from project scope to user scope so its tools are prioritized ahead of the other nine connected servers", "correct": false, "explanation": "Scope (local/project/user) controls where a server's configuration is stored and who it's shared with; it has no effect on tool-search deferral or loading priority."}, {"letter": "D", "text": "Increase MAX_MCP_OUTPUT_TOKENS for just that server so its tool responses return noticeably faster once a call is actually made", "correct": false, "explanation": "MAX_MCP_OUTPUT_TOKENS controls how much output a tool call can return, not whether the tool's definition is deferred behind a tool-search step, so it doesn't address the described delay."}], "correct": "A", "select": 1, "group": "J"}, {"id": "f2-022", "domain": 2, "task_id": "2.4", "objective": "Integrate MCP servers into Claude Code and agent workflows", "situation": "An internal MCP server requires a Kerberos-derived token that must be freshly minted for every connection, and no OAuth authorization server is involved.", "question": "According to Anthropic documentation, which approach is recommended for handling this authentication scheme in Claude Code?", "options": [{"letter": "A", "text": "Do not attempt to configure Kerberos directly in Claude Code MCP authentication; use an intermediary service that authenticates via Kerberos and then uses Anthropic-supported credentials such as API keys or OAuth tokens.", "correct": false, "explanation": "This conflates authenticating to an internal MCP server with authenticating to Anthropic's API: Anthropic API keys and OAuth tokens are credentials for Anthropic services and do not satisfy an internal server's Kerberos requirement. The documentation names Kerberos explicitly as a scheme handled by a header-generating helper command, so no intermediary service is required."}, {"letter": "B", "text": "Configure a static headers entry with the token value hardcoded, then rotate the config file manually whenever the token expires.", "correct": false, "explanation": "A hardcoded static headers entry cannot satisfy the requirement that the token be freshly minted for every connection, because manual rotation only replaces the stored value after it has already expired. The documentation also notes that dynamic headers override any static header of the same name, and that a rejected static Authorization header makes Claude Code report the connection as failed."}, {"letter": "C", "text": "Configure headersHelper to run a script that generates the token and writes the resulting header JSON to stdout on each connection.", "correct": true, "explanation": "The Claude Code MCP documentation section \"Use dynamic headers for custom authentication\" states that when an MCP server uses an authentication scheme other than OAuth, such as Kerberos, short-lived tokens, or an internal SSO, headersHelper generates request headers at connection time. The helper must write a JSON object of string key-value pairs to stdout, and Claude Code runs it fresh on each connection, at session start and on reconnect, without caching the result, which is precisely what a per-connection Kerberos-derived token requires."}, {"letter": "D", "text": "Configure the oauth block with authServerMetadataUrl pointed at the internal Kerberos realm so Claude Code discovers the flow automatically.", "correct": false, "explanation": "authServerMetadataUrl inside the oauth block points Claude Code at an OAuth authorization server metadata document, alongside the default discovery chain of RFC 9728 protected resource metadata and RFC 8414 authorization server metadata. A Kerberos realm publishes no such metadata document, and the scenario states that no OAuth authorization server is involved, so discovery has nothing to resolve."}], "correct": "C", "select": 1, "group": "J"}, {"id": "f2-023", "domain": 2, "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "situation": "A subagent responsible for enriching customer records calls an internal lookup_address tool for 50 customers. For 47 the lookup succeeds; for 3 the tool returns isError: true with errorCategory: \"permission\" because the subagent's credentials lack access to those records' region. The subagent cannot obtain broader credentials itself.", "question": "What should it send back to the coordinator?", "options": [{"letter": "A", "text": "A retry loop that keeps calling lookup_address on the same 3 customers with the same credentials until the coordinator intervenes", "correct": false, "explanation": "Retrying with the same insufficient credentials will keep failing identically; a permission error of this kind cannot be resolved by repetition and should be escalated instead."}, {"letter": "B", "text": "An isError: true result for the entire batch, discarding the 47 successful lookups because the batch as a whole did not fully complete", "correct": false, "explanation": "Discarding 47 successful lookups because 3 others failed for an unrelated reason throws away valid completed work that the coordinator could otherwise use immediately."}, {"letter": "C", "text": "The 47 enriched records as partial results, plus a report naming the 3 unresolved customers, the permission error encountered, and what was attempted", "correct": true, "explanation": "Errors the subagent cannot resolve locally, like a credential-scope gap, should propagate to the coordinator along with the partial results already obtained and a summary of what was tried, so the coordinator can decide the next step."}, {"letter": "D", "text": "Only the 47 enriched records, silently dropping the 3 failures since local recovery attempts already reached their limit for this subagent", "correct": false, "explanation": "Silently dropping the failures hides a real, unresolved permission problem from the coordinator, which then has no way to know 3 records were never enriched."}], "correct": "C", "select": 1, "group": "J"}, {"id": "f2-024", "domain": 2, "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "situation": "A run_report MCP tool depends on a data warehouse that occasionally throttles requests with a 429 response. The tool author wants the agent to back off and retry automatically, but only up to a sensible limit, rather than retrying forever or giving up immediately.", "question": "Which structured error response best supports this behavior?", "options": [{"letter": "A", "text": "errorCategory: \"permission\", isRetryable: false, and a description stating the report requires elevated warehouse access to proceed", "correct": false, "explanation": "Rate limiting isn't an access-control failure; the caller already has permission to run the report, it's simply being throttled, so labeling this permission and non-retryable would wrongly stop the agent from ever succeeding."}, {"letter": "B", "text": "errorCategory: \"validation\", isRetryable: true, and a description asking the agent to reduce the report's date range before resubmitting the identical query", "correct": false, "explanation": "Nothing about the request's shape or the date range caused the 429; the request itself was valid, and telling the agent to change the query misdiagnoses a load-based throttle as a malformed-input problem."}, {"letter": "C", "text": "errorCategory: \"transient\", isRetryable: true, and a description noting the warehouse is rate-limiting requests, letting the agent apply its own bounded backoff strategy", "correct": true, "explanation": "Rate limiting is a transient, retryable condition tied to load rather than input correctness or access rights, so marking it transient with isRetryable true and a clear description lets the agent apply bounded backoff rather than guessing."}, {"letter": "D", "text": "No errorCategory or isRetryable field at all, relying on the phrase \"try again later\" in the text block to convey the retry semantics", "correct": false, "explanation": "Relying on free text alone forces the agent to parse natural language for retry semantics instead of reading structured fields, reintroducing the ambiguity that structured metadata is meant to eliminate."}], "correct": "C", "select": 1, "group": "J"}, {"id": "f2-025", "domain": 2, "task_id": "2.5", "objective": "Select and apply built-in tools (Read, Write, Edit, Bash, Grep, Glob) effectively", "situation": "A newly onboarded architect asks Claude Code to trace how a login request flows from the HTTP route handler through to the database call, in a codebase Claude has not explored yet.", "question": "To build this understanding efficiently while keeping context usage low, what is the best incremental strategy?", "options": [{"letter": "A", "text": "Use Read to open every file under the src directory up front, building a complete mental model of the whole codebase before looking for the login flow specifically.", "correct": false, "explanation": "Reading every file upfront is extremely inefficient and wastes context window. Claude Code has limited context capacity, and loading entire codebases leads to unnecessary token consumption and slower performance. The goal is to trace a specific flow, not understand the entire codebase at once."}, {"letter": "B", "text": "Use Glob to list every file in the repository sorted by modification time, then read the twenty most recently modified files on the assumption they relate to login.", "correct": false, "explanation": "Recent file modification does not guarantee relevance to the login flow. Other unrelated features might have been updated recently. This approach ignores project structure and the specific call chain, making it inefficient and prone to missing the actual login-related files."}, {"letter": "C", "text": "Start by reading CLAUDE.md or AGENTS.md if they exist to gain high-level architecture context, then use Grep to locate the route handler and its imports, and read files incrementally along the call chain.", "correct": true, "explanation": "According to Anthropic's documentation, CLAUDE.md and AGENTS.md are project context files that Claude Code reads at the start of every session, providing persistent high-level context without manual prompting. Leveraging these files first gives the architect necessary architectural overview, then using grep to pinpoint the route handler and progressively reading only relevant files keeps context usage low and targets the specific call chain efficiently. This incremental approach avoids overwhelming context with irrelevant files and aligns with recommended onboarding practices."}, {"letter": "D", "text": "Use Bash to run a full-text word count across the repository and read the files with the highest counts, on the assumption larger files hold core business logic.", "correct": false, "explanation": "File size does not reliably indicate where the login flow logic resides. This heuristic is arbitrary and likely includes many irrelevant large files (e.g., config, test fixtures, generated code). It wastes context and misses the targeted incremental exploration needed."}], "correct": "C", "select": 1, "group": "F"}, {"id": "f2-026", "domain": 2, "task_id": "2.5", "objective": "Select and apply built-in tools (Read, Write, Edit, Bash, Grep, Glob) effectively", "situation": "Claude Code needs to overwrite an existing configuration file, settings.local.json, with an updated version generated in response to the architect's request. Claude has not read this file at any point earlier in the current conversation.", "question": "What must Claude do before the Write call will succeed?", "options": [{"letter": "A", "text": "Delete the file first using Bash, since Write can only create files that do not already exist and cannot overwrite one directly", "correct": false, "explanation": "This approach is unnecessary and risky, as it could lead to data loss if something goes wrong between deletion and writing. Write is designed to handle both creation and overwriting of files. While the Files API requires delete-and-reupload for file modification, Claude Code's Write tool operates differently and can overwrite existing files when the proper safety prerequisites (such as prior reading) are met."}, {"letter": "B", "text": "Nothing extra is required, since Write can overwrite any existing file at any time regardless of whether it has been read in the conversation", "correct": false, "explanation": "This statement is inaccurate per Anthropic's official documentation. While Write can overwrite files, the recommended and default behavior includes safeguards—such as requiring the file to have been read first—to prevent unintended or destructive modifications. Arbitrary overwriting without prior context is not advised and may be restricted in practice."}, {"letter": "C", "text": "Read the existing file first, since Write requires it to have been read in this conversation before overwriting an existing file", "correct": true, "explanation": "Claude Code's Write tool includes safety mechanisms to prevent accidental overwrites. It typically requires that the target file has been read during the current conversation, ensuring Claude is aware of the current contents before making changes. Anthropic's documentation and best practices do not recommend arbitrary overwriting without such awareness, and discussions highlight the need for careful management of file modifications."}, {"letter": "D", "text": "Run Grep against the file to confirm its current contents, since Grep results satisfy the same prior-access requirement that Read would", "correct": false, "explanation": "Running Grep only provides partial, search-oriented content and does not fulfill the intent of the safety check. The Write tool's prerequisite is typically a full file read to guarantee Claude has seen the complete file. A tool like Grep may be used for specific content inspection, but it is not equivalent to a Read operation in the context of this safeguard."}], "correct": "C", "select": 1, "group": "F"}, {"id": "f2-027", "domain": 2, "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "situation": "Two tools, analyze_content (summarizes pasted text) and analyze_document (summarizes uploaded documents), share nearly identical one-line descriptions. Users report the model frequently calls analyze_document for pasted web article text instead of analyze_content.", "question": "What is the most effective fix?", "options": [{"letter": "A", "text": "Rewrite each description to state the expected input type and add a note on when the other tool applies.", "correct": true, "explanation": "Tool descriptions are the model's primary signal for tool selection; stating the expected input type and contrasting the two tools removes the ambiguity causing the misrouting."}, {"letter": "B", "text": "Increase the model's sampling temperature so it selects a more varied set of tools during response generation.", "correct": false, "explanation": "Sampling temperature affects token-level randomness in generated text, not the semantic reasoning the model uses to distinguish two functionally similar tool descriptions."}, {"letter": "C", "text": "Add a longer paragraph of promotional wording to analyze_document describing it as the more capable option.", "correct": false, "explanation": "Promotional wording adds length without adding differentiating information about input type or scope, so the ambiguity between the two tools remains."}, {"letter": "D", "text": "Remove analyze_content entirely from the tool list so the model always falls back to analyze_document for every case.", "correct": false, "explanation": "Removing a tool eliminates its intended use case rather than fixing the underlying description ambiguity, and pasted text would now be mishandled by a document-only tool."}], "correct": "A", "select": 1, "group": "J"}, {"id": "f2-028", "domain": 2, "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "situation": "A search_tickets MCP tool queries a support database for tickets matching a filter. For a particular customer, the query executes successfully but zero tickets match. The tool currently returns isError: true with the text \"No tickets found,\" and the coordinating agent responds to the user by apologizing for a system failure.", "question": "What is the correct fix?", "options": [{"letter": "A", "text": "Return isError: false but omit the results array entirely, letting the agent infer from the missing field that the search matched nothing", "correct": false, "explanation": "While omitting the results array might seem concise, it introduces ambiguity. Best practices recommend returning a structured empty array as meaningful context, making the outcome explicit to the agent. An absent field could be misinterpreted as an error or incomplete response."}, {"letter": "B", "text": "Keep isError: true but change errorCategory to transient so the agent automatically retries the identical search until a ticket eventually appears", "correct": false, "explanation": "Retrying a search that correctly returned no results is wasteful and will never yield matches. Transient error handling is for recoverable failures (e.g., timeouts), not for an empty result set. The correct approach is to return isError: false to indicate the tool completed successfully."}, {"letter": "C", "text": "Keep isError: true and add isRetryable false, since an empty result set should be treated the same as any other failed tool invocation", "correct": false, "explanation": "An empty result set from a valid query is not a failure; it is a successful execution with no data. Anthropic's guidelines explicitly state that returning an error flag for a successful query with no matches is an anti-pattern, as it misleads the agent into thinking something went wrong."}, {"letter": "D", "text": "Return isError: false with structured content showing an empty results array, since a successful query with no matches is not a tool execution error", "correct": true, "explanation": "Anthropic's best practices distinguish between a successful operation with no results and a tool execution error. Returning isError: false with an empty array is the recommended way to signal 'no matches found' because the tool executed successfully; treating it as an error would lead to inappropriate retries or misleading agent responses. For genuine errors, an explicit error flag and context should be returned."}], "correct": "D", "select": 1, "group": "J"}, {"id": "f2-029", "domain": 2, "task_id": "2.5", "objective": "Select and apply built-in tools (Read, Write, Edit, Bash, Grep, Glob) effectively", "situation": "An architect asks Claude Code to find every place in a large monorepo that calls a function named parseInvoice, including calls inside a minified bundle that is gitignored but still needs to be checked.", "question": "Which approach correctly locates all call sites?", "options": [{"letter": "A", "text": "Run Glob with the pattern **/*.js to list every JavaScript file, then judge from file names alone which ones likely reference parseInvoice", "correct": false, "explanation": "Glob only matches file paths by name pattern; it cannot see whether a file's contents reference parseInvoice, so this cannot identify call sites."}, {"letter": "B", "text": "Run Grep across the repo for parseInvoice, then Grep the gitignored bundle's path directly, since a direct path is still searched", "correct": true, "explanation": "Grep respects .gitignore and skips gitignored files by default, but passing a gitignored file's path directly still searches it, so a normal repo-wide Grep plus a targeted Grep on the bundle path covers both tracked and gitignored call sites."}, {"letter": "C", "text": "Run Grep once with the multiline flag enabled, assuming multiline mode makes Grep search gitignored files as a side effect of that flag", "correct": false, "explanation": "The multiline flag changes whether a pattern can match across line boundaries; it has no effect on whether gitignored files are included in the search."}, {"letter": "D", "text": "Run Glob with the pattern **/parseInvoice* to find files whose names contain the function, then treat that file list as the complete set of callers", "correct": false, "explanation": "Globbing for files named after the function only finds a file that happens to share the name, not the files where the function is called, so real call sites in other files are missed."}], "correct": "B", "select": 1, "group": "F"}, {"id": "f2-030", "domain": 2, "task_id": "2.5", "objective": "Select and apply built-in tools (Read, Write, Edit, Bash, Grep, Glob) effectively", "situation": "An architect asks Claude Code to identify every React test file in a codebase where naming mixes .test.tsx, .spec.tsx, and older files simply ending in Test.tsx, spread across many nested feature directories. Only a list of matching file paths is needed, with no content inspection.", "question": "Which tool is the most direct fit?", "options": [{"letter": "A", "text": "Grep, using output mode files_with_matches and a regex that matches the word test anywhere inside a file's contents", "correct": false, "explanation": "Grep searches file contents, not file names; searching for the word test inside file bodies would return unrelated files that mention testing and could miss test files that never use that literal word."}, {"letter": "B", "text": "Glob, using patterns such as **/*.test.tsx, **/*.spec.tsx, and **/*Test.tsx to match the naming conventions directly", "correct": true, "explanation": "Glob matches file paths by name pattern, including recursive ** matching, so running it with the three naming conventions directly returns exactly the matching file paths without needing to inspect any file contents."}, {"letter": "C", "text": "Read, pointed at the project root directory so it returns a recursive listing of every file that exists below it", "correct": false, "explanation": "Read loads the contents of a single file at a given path and does not list directories at all, so pointing it at the project root would not produce a recursive file listing."}, {"letter": "D", "text": "Bash, using a recursive directory listing command and then manually reading every returned file to check its extension", "correct": false, "explanation": "A recursive Bash listing followed by manually reading every file to check its name is far less direct than a purpose-built path-pattern match, and unnecessarily consumes context reading files that were never needed."}], "correct": "B", "select": 1, "group": "F"}, {"id": "f2-031", "domain": 2, "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "situation": "A deploy_service MCP tool requires an API token scoped to the deploy role. An agent calls it using a token scoped only to read.", "question": "The server returns isError: true with generic text \"Operation failed.\" Under a structured error design, how should this failure be categorized and handled differently from a network timeout on the same tool?", "options": [{"letter": "A", "text": "As errorCategory: \"permission\" with isRetryable: false, since resubmitting with the same token will fail identically until the caller's scope changes", "correct": true, "explanation": "An insufficiently scoped token is a permission error distinct from a timeout: it will never succeed on retry without a credential change, so isRetryable should be false and the category should reflect the access problem specifically."}, {"letter": "B", "text": "As errorCategory: \"validation\" with isRetryable: true, since the token itself is a malformed input field that a retry with backoff can resolve", "correct": false, "explanation": "The token isn't malformed input to be corrected by the caller's request shape; it's an authorization problem, and no amount of retrying with backoff changes what role the token carries."}, {"letter": "C", "text": "As errorCategory: \"transient\" with isRetryable: true, since both permission failures and timeouts stem from the deploy service being temporarily unreachable", "correct": false, "explanation": "A timeout is about transient unavailability that may clear on its own; a scope mismatch is a fixed condition tied to the caller's credentials and won't resolve by waiting or retrying."}, {"letter": "D", "text": "As errorCategory: \"business\" with isRetryable: false, since restricting deploy access is functionally the same kind of policy rule as a refund window", "correct": false, "explanation": "Business errors reflect policy rules about the request itself (like a return window), while this failure is about who is allowed to call the tool at all, which is the defining trait of a permission error."}], "correct": "A", "select": 1, "group": "J"}, {"id": "f2-032", "domain": 2, "task_id": "2.3", "objective": "Distribute tools appropriately across agents and configure tool choice", "situation": "A document-processing agent has both an extract_metadata tool and several enrichment tools (add_tags, link_related, generate_summary). The enrichment tools depend on fields that only extract_metadata produces, and in early tests the model sometimes calls an enrichment tool first with guessed values.", "question": "What is the recommended way to guarantee correct ordering for this first step?", "options": [{"letter": "A", "text": "Set tool_choice to {\"type\": \"tool\", \"name\": \"extract_metadata\"} on the turn where metadata is needed, then let the model choose from the enrichment tools with auto or any on subsequent turns.", "correct": true, "explanation": "Setting tool_choice to {\"type\": \"tool\", \"name\": \"extract_metadata\"} on the turn where metadata is needed forces the model to call that specific tool first, guaranteeing the required ordering. Once the metadata is in the conversation context, switching to auto or any on subsequent turns allows the model to freely select among the enrichment tools."}, {"letter": "B", "text": "Add a detailed description to each enrichment tool stating that they require metadata fields only extract_metadata can provide, and keep tool_choice set to auto throughout the conversation.", "correct": false, "explanation": "Adding detailed descriptions and relying on auto is exactly the type of nudge that has already proven unreliable in the team's tests. Description changes may influence the model but do not guarantee the required ordering, whereas forced tool_choice enforces it directly."}, {"letter": "C", "text": "Set tool_choice to {\"type\": \"any\"} for the entire conversation, so the model must select a tool on every turn, relying on its training to choose extract_metadata first when metadata is absent.", "correct": false, "explanation": "Setting tool_choice to any forces a tool call on every turn but does not specify which one; the model could still choose an enrichment tool first. This does not guarantee that extract_metadata will be called before the enrichment tools depend on its output."}, {"letter": "D", "text": "Remove the enrichment tools from the agent's tool list on the first turn and only provide extract_metadata, then after extract_metadata returns, re-add add_tags, link_related, and generate_summary for subsequent turns.", "correct": false, "explanation": "While removing enrichment tools and re-adding them after extract_metadata returns would work, it is a heavier-handed and less standard approach than using the tool_choice mechanism. The recommended way is to force the specific tool on the first turn, not to dynamically modify the tool list across turns."}], "correct": "A", "select": 1, "group": "B"}, {"id": "f2-033", "domain": 2, "task_id": "2.3", "objective": "Distribute tools appropriately across agents and configure tool choice", "situation": "A coordinator agent delegates tasks to a research agent, a coding agent, and a QA agent. Currently, every subagent is configured with the full union of all tools used anywhere in the pipeline, so each has around 15 tools available, and each agent's tool choice is left at {\"type\": \"auto\"} in every turn. The team reports agents occasionally reaching for tools clearly outside their remit, like the QA agent invoking a deploy tool.", "question": "What combination of changes best addresses this while preserving each agent's ability to decide when to act versus respond?", "options": [{"letter": "A", "text": "Restrict each subagent's tool set to only what its own role needs, and also force every subagent's tool_choice to a single named tool for all turns, such as search tool for research and test tool for QA.", "correct": false, "explanation": "While scoping tools to each role is a step in the right direction, forcing every subagent's tool_choice to a single named tool eliminates their flexibility to choose among appropriate tools as tasks vary. This conflicts with the requirement to preserve each agent's ability to decide when and which tool to use."}, {"letter": "B", "text": "Keep the shared 15-tool set for all subagents, but switch every subagent's tool_choice to {\"type\": \"none\"} so no subagent can call tools directly, requiring the coordinator to invoke tools on its behalf.", "correct": false, "explanation": "Setting tool_choice to {\"type\": \"none\"} for all subagents completely prevents them from calling any tools directly, breaking the delegation pipeline. It does not address the scoping problem and removes the agents' ability to decide per turn whether to act, which the scenario explicitly requires preserving."}, {"letter": "C", "text": "Restrict each subagent's tool set to only what its role needs, while keeping tool_choice at {\"type\": \"auto\"} so each agent still decides per turn whether to call a tool.", "correct": true, "explanation": "Restricting each subagent's tool set to only what its role needs directly removes the out-of-role tools (like the deploy tool from the QA agent) that caused misuse. Keeping tool_choice as {\"type\": \"auto\"} preserves each agent's ability to autonomously decide on every turn whether a tool call is needed, exactly as required."}, {"letter": "D", "text": "Keep the shared 15-tool set for all subagents, but switch every subagent's tool_choice to {\"type\": \"any\"} so a tool call is always forced, ensuring subagents cannot respond without selecting a tool.", "correct": false, "explanation": "This approach keeps the over-broad 15-tool set, so the QA agent still has access to the deploy tool. Forcing tool_choice to {\"type\": \"any\"} guarantees a tool call every turn, which actually increases the likelihood of inappropriate tool usage rather than resolving it."}], "correct": "C", "select": 1, "group": "B"}, {"id": "f2-034", "domain": 2, "task_id": "2.3", "objective": "Distribute tools appropriately across agents and configure tool choice", "situation": "A team is building an agent that uses manual extended thinking (thinking: {type: \"enabled\"}) to reason before acting, and they want to force it to always call a tool rather than answer directly. They set tool_choice to {\"type\": \"any\"} while manual extended thinking is enabled, and the request fails.", "question": "What is the correct explanation and recommended remedy?", "options": [{"letter": "A", "text": "When manual extended thinking is enabled, the tool_choice values {\"type\": \"any\"} and {\"type\": \"tool\", ...} are not supported; set it to {\"type\": \"auto\"} or {\"type\": \"none\"} instead. To force a tool call while still using thinking, migrate to adaptive thinking (supported on newer models), or disable manual extended thinking.", "correct": true, "explanation": "Anthropic's documentation states that tool use with manual extended thinking (thinking: {type: \"enabled\"}) supports only tool_choice: {\"type\": \"auto\"} (the default) or tool_choice: {\"type\": \"none\"}. Attempting {\"type\": \"any\"} or {\"type\": \"tool\", ...} forces tool use and causes an error with manual extended thinking. Adaptive thinking on newer models does support forced tool use, so if forced tool calling is required, use adaptive thinking or disable manual extended thinking."}, {"letter": "B", "text": "The request failed because extended thinking disables all tool_choice options; to fix this, remove all tools from the request and let the model output its reasoning steps as text before acting.", "correct": false, "explanation": "Extended thinking does not disable all tool_choice options; manual extended thinking supports auto and none, while adaptive thinking supports forced tool use. Removing tools is not required and would prevent tool use entirely."}, {"letter": "C", "text": "The request failed because {\"type\": \"any\"} is deprecated; replace it with {\"type\": \"forced\"} and specify a tool name like search to satisfy the forced tool choice requirement.", "correct": false, "explanation": "tool_choice: {\"type\": \"any\"} is not deprecated; it is a valid option that forces the model to choose one of the provided tools. There is no {\"type\": \"forced\"} option; the named forced tool choice is {\"type\": \"tool\", \"name\": \"...\"}."}, {"letter": "D", "text": "The request failed because {\"type\": \"any\"} requires at least two tools to be defined; adding a second tool, such as a calculator, resolves the incompatibility with extended thinking.", "correct": false, "explanation": "tool_choice: {\"type\": \"any\"} does not require any minimum number of tools beyond having at least one tool available; it means the model must select from the provided tools. The incompatibility is with manual extended thinking, not tool count."}], "correct": "A", "select": 1, "group": "B"}, {"id": "f2-035", "domain": 2, "task_id": "2.5", "objective": "Select and apply built-in tools (Read, Write, Edit, Bash, Grep, Glob) effectively", "situation": "Claude needs to reorganize a file by moving several scattered export statements into one grouped block near the top.", "question": "Which tool sequence should Claude use?", "options": [{"letter": "A", "text": "Read the file to load its full contents, then call Write with the complete restructured file content back over that same path.", "correct": true, "explanation": "Read loads the current file contents into context, and Write overwrites the target file with the complete restructured version, including the reordered and consolidated export block. This is the appropriate approach when reshaping overall layout rather than editing a single contiguous span."}, {"letter": "B", "text": "Issue one Edit call per export statement, each targeting a short unique snippet, relying on the accumulated edits to produce the new grouped layout.", "correct": false, "explanation": "A dozen separate Edit calls would be more complex and error-prone than a single Read/Write rewrite, especially because reordering and consolidation cannot be achieved cleanly through independent contiguous snippet replacements."}, {"letter": "C", "text": "Call Glob for the file's own path to confirm it exists, then call Edit with old_string set to the whole file's text and new_string as the new version.", "correct": false, "explanation": "Glob only confirms path existence and does not modify content. Edit is intended for short, unique, contiguous snippets, so using the entire file as old_string is inefficient and error-prone for whole-file replacement."}, {"letter": "D", "text": "Call Grep with output mode content to retrieve the matching export lines, treating the returned text as already written back to the file.", "correct": false, "explanation": "Grep with output mode content only retrieves matching lines and does not write them back to the file. Treating returned text as already written would leave the file unchanged and the export statements still scattered."}], "correct": "A", "select": 1, "group": "F"}, {"id": "f2-036", "domain": 2, "task_id": "2.5", "objective": "Select and apply built-in tools (Read, Write, Edit, Bash, Grep, Glob) effectively", "situation": "Claude Code needs to find all Python files that define a class inheriting from BaseHandler, in a codebase where class definitions may span multiple lines when the base class list is long, such as class OrderHandler(\\n BaseHandler, LoggingMixin\\n):. A single-line Grep search for BaseHandler only catches some of these definitions.", "question": "What is the most direct fix?", "options": [{"letter": "A", "text": "Run Grep once per file using Bash to invoke it individually on each path, on the assumption that Grep cannot be scoped to one language repo-wide", "correct": false, "explanation": "Grep already supports scoping by language with the type parameter or by file pattern with the glob parameter across an entire repository in a single call, so invoking it per file through Bash is unnecessary."}, {"letter": "B", "text": "Run Grep scoped to Python files with the type parameter set to py, and enable multiline mode so wrapped class headers are matched", "correct": true, "explanation": "Scoping the search to Python files with the type parameter narrows the search appropriately, and enabling multiline mode allows the BaseHandler pattern to match even when the class header wraps across multiple lines."}, {"letter": "C", "text": "Run Glob with the pattern **/*.py to list every Python file, then read each file completely to visually check for BaseHandler", "correct": false, "explanation": "Reading every Python file in full to manually check for BaseHandler is far less efficient than fixing the search pattern, and does not scale well across a large codebase."}, {"letter": "D", "text": "Run Grep with output mode files_with_matches only, on the assumption that this mode searches content more thoroughly than content mode does", "correct": false, "explanation": "files_with_matches only changes what information is returned about matches, not how thoroughly the content is searched; it would not help match a pattern spanning multiple lines any more than content mode does."}], "correct": "B", "select": 1, "group": "F"}, {"id": "f2-037", "domain": 2, "task_id": "2.3", "objective": "Distribute tools appropriately across agents and configure tool choice", "situation": "A classification subagent must always emit a structured label using one of its provided tools (e.g. tag_urgent, tag_normal, tag_spam) and must never return free-text commentary instead of a call, though which specific tag applies depends on the message content.", "question": "Which tool_choice setting guarantees this behavior?", "options": [{"letter": "A", "text": "tool_choice: {\"type\": \"tool\", \"name\": \"tag_normal\"}, which forces the same tag every time regardless of message content", "correct": false, "explanation": "Forcing a single named tool would apply the same tag on every message regardless of content, which contradicts the need for the correct tag to vary by message."}, {"letter": "B", "text": "tool_choice: {\"type\": \"any\"}, which requires the model to call one of the provided tools without pinning it to a specific one", "correct": true, "explanation": "\"any\" tells the model it must use one of the provided tools without forcing a particular one, guaranteeing a structured tool call while still letting the model pick the tag that matches the message."}, {"letter": "C", "text": "tool_choice: {\"type\": \"auto\"}, which lets the model decide whether calling a tag tool or replying in prose better fits the message", "correct": false, "explanation": "\"auto\" leaves the choice of calling a tool at all up to the model, which does not guarantee a structured tag is always produced, undermining the requirement for consistent structured output."}, {"letter": "D", "text": "tool_choice: {\"type\": \"none\"}, which stops the model from calling any tag tool and relies on prompt wording instead", "correct": false, "explanation": "\"none\" blocks tool calls entirely, so the model could only respond in prose, which is the opposite of the required structured-output guarantee."}], "correct": "B", "select": 1, "group": "B"}, {"id": "f2-038", "domain": 2, "task_id": "2.5", "objective": "Select and apply built-in tools (Read, Write, Edit, Bash, Grep, Glob) effectively", "situation": "A module dateUtils.ts re-exports several functions under different names, such as export { formatDate as fmt, parseDate as pd }. Claude needs to change the parameter list that formatDate accepts, so every call site must be updated to match the new signature — whether it invokes the function by its original name or through a re-exported alias like fmt.", "question": "Which approach correctly accounts for the re-exports?", "options": [{"letter": "A", "text": "Run Glob for **/dateUtils* to find files related to date utilities, then assume every caller also lives inside a file matching that same pattern", "correct": false, "explanation": "Callers of a utility module are typically scattered across unrelated feature files that do not share dateUtils in their own file names, so this pattern would miss most or all real call sites."}, {"letter": "B", "text": "Run a single Grep search for the literal string formatDate, assuming any caller that uses the alias fmt will also match that same pattern", "correct": false, "explanation": "A caller that imports and uses the alias fmt never has the literal text formatDate in its own file, so a single search for formatDate would miss every call site that uses the re-exported name — and every one of those call sites needs updating for the new parameter list."}, {"letter": "C", "text": "Read dateUtils.ts first to identify every exported name including aliases, then Grep for each exported name across the codebase", "correct": true, "explanation": "Reading the module first reveals all exported names, including aliases created by the re-export; searching for each exported name individually with Grep finds every call site that must be updated for the new signature, whether it uses the original name or an alias."}, {"letter": "D", "text": "Run Bash to count how many times the word export appears in the repository, then read only the files where that count exceeds a fixed threshold", "correct": false, "explanation": "Counting occurrences of the word export across the repository has no relationship to which files call formatDate or its alias, so this does not help locate the call sites that need updating."}], "correct": "C", "select": 1, "group": "F"}, {"id": "f2-039", "domain": 2, "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "situation": "A tool description reads \"Updates a customer record with given fields.\" In practice, the model frequently passes fields that don't exist on the schema, and it's unclear whether partial updates are supported.", "question": "What documentation gap explains this failure mode?", "options": [{"letter": "A", "text": "The description omits the expected input format and boundary behavior, such as which fields are valid and whether partial updates work.", "correct": true, "explanation": "Anthropic's best practices stress that tool descriptions must specify valid fields, expected formats, and boundary conditions (like partial update support) to prevent model hallucination. Without these details, the model cannot reliably infer correct inputs, leading to the observed failures. Official documentation emphasizes clear input schemas to avoid such ambiguity."}, {"letter": "B", "text": "The model cannot reliably handle optional parameters, so the documentation must require every field to be marked required, such as all fields being mandatory.", "correct": false, "explanation": "Claude can handle optional parameters if they are properly defined in the tool's schema. Requiring all fields as mandatory is not a documentation gap but an overly restrictive design choice that deviates from standard API practices and limits flexibility. The core issue is insufficient description, not optional parameter support."}, {"letter": "C", "text": "The tool's JSON schema uses camelCase field names, which conflicts with the model's snake_case training, causing invalid field passing.", "correct": false, "explanation": "Anthropic models are trained on diverse text and can handle camelCase field names without issue. There is no documented conflict with snake_case that would cause the model to pass invalid fields. The failure stems from unclear tool specification, not casing conventions."}, {"letter": "D", "text": "The customer record tool was declared after read-only tools, causing the model to deprioritize it and pass fields not defined in the schema.", "correct": false, "explanation": "Tool declaration order does not influence tool prioritization or cause the model to pass invalid fields. The model selects tools based on user intent and tool descriptions, not declaration order. No evidence suggests order impacts field validity."}], "correct": "A", "select": 1, "group": "J"}, {"id": "f2-040", "domain": 2, "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "situation": "A coordinator dispatches the same document-indexing task to three subagents in parallel, each covering a different folder. Subagent 1 finishes cleanly. Subagent 2 hits a permission error on one file it cannot resolve locally and reports partial results plus that failure. Subagent 3's process crashes with no output at all.", "question": "How should the coordinator's downstream handling differ between subagent 2 and subagent 3?", "options": [{"letter": "A", "text": "For subagent 2, the coordinator should ignore the reported permission error and mark the folder complete, since most files were indexed successfully, and for subagent 3, the coordinator should also mark its folder complete because no error was reported.", "correct": false, "explanation": "Ignoring a reported permission gap and marking the folder complete hides a real indexing deficiency from downstream consumers. Additionally, treating subagent 3's folder as complete when no output was received masks a total processing failure."}, {"letter": "B", "text": "The coordinator should treat both subagent 2 and subagent 3 identically by retrying files with permission errors, assuming subagent 3's crash was also due to a permission issue on some file, since that is the most common cause of non-completion in such tasks.", "correct": false, "explanation": "Assuming a crash with no output was caused by a permission error is speculative and unfounded; crashes can stem from many root causes. Guessing without evidence risks misdirecting recovery efforts and may not resolve the actual issue."}, {"letter": "C", "text": "For subagent 2, the coordinator can use the partial results and address the specific reported permission gap; for subagent 3, lacking completed work or diagnostic detail, it must treat the entire folder as unprocessed.", "correct": true, "explanation": "Subagent 2 returns partial results and a specific permission error, allowing the coordinator to integrate successes and precisely address the failure. Subagent 3 provides no output, leaving no basis for targeted recovery, so the entire folder must be considered unprocessed."}, {"letter": "D", "text": "The coordinator should treat both subagent 2 and subagent 3 identically by discarding any partial results from subagent 2 and marking both folders for indexing as unprocessed, since neither subagent fully completed its assigned indexing task successfully.", "correct": false, "explanation": "Discarding subagent 2's partial results wastes completed indexing work and ignores diagnostic detail that could guide precise recovery. Treating a structured partial failure the same as a total, silent crash is wasteful and overlooks actionable information."}], "correct": "C", "select": 1, "group": "J"}, {"id": "f2-041", "domain": 2, "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "situation": "A tool's description reads only \"Get weather for a location.\" Users report the model sometimes garbles multi-word city names and invokes this tool when the user actually wants a historical climate report.", "question": "What change would most directly reduce these failures?", "options": [{"letter": "A", "text": "Change the tool's name to a random unique string so it no longer semantically collides with any other tool at all.", "correct": false, "explanation": "Changing the tool's name to a random unique string removes semantic meaning, making it harder for the model to understand when to invoke the tool. Tool names should be descriptive and align with their function; a meaningless name increases ambiguity rather than resolving collisions."}, {"letter": "B", "text": "Move all input validation into the tool's runtime error handler instead of describing constraints up front.", "correct": false, "explanation": "Moving input validation to the runtime error handler only catches mistakes after the tool has already been invoked, so it does not prevent the model from calling the tool incorrectly in the first place. Describing constraints up front in the tool description is what helps the model make correct invocation decisions."}, {"letter": "C", "text": "Expand the description to state the input format, include an example query, and note historical lookups are out of scope.", "correct": true, "explanation": "Anthropic's documentation emphasizes reducing ambiguity by being explicit about input and output formats, providing examples, and noting scope limitations. Expanding the description to include the expected input format (e.g., a single city name), an example query, and a statement that historical climate reports are out of scope directly addresses both garbled multi-word names and incorrect invocation for historical data. This aligns with best practices for tool descriptions, which should clearly specify the tool's purpose, inputs, and boundaries."}, {"letter": "D", "text": "Shorten the description further so the model spends less time interpreting it before invoking the tool.", "correct": false, "explanation": "Shortening the description further would increase ambiguity, as vague prompts cause the model to fill in assumptions that often mismatch user intent. Anthropic's guidance explicitly warns that vague prompts produce inconsistent outputs, so more detail—not less—is needed to reduce failures."}], "correct": "C", "select": 1, "group": "J"}, {"id": "f2-042", "domain": 2, "task_id": "2.3", "objective": "Distribute tools appropriately across agents and configure tool choice", "situation": "A single generalist agent handling an entire content pipeline (research, outline, draft, edit, publish) has 20 tools and shows declining tool-selection accuracy as more tools were added over time. The architect proposes splitting it into five specialized subagents, each scoped to roughly 4 tools for its stage, coordinated by a lightweight orchestrator.", "question": "What is the main reliability benefit of this redesign, based on the relationship between tool count and selection accuracy?", "options": [{"letter": "A", "text": "Each subagent now runs on a smaller context window, which forces the model to think more carefully before selecting a tool", "correct": false, "explanation": "Subagent context size is unrelated to the tool-selection problem described; the issue was the number of competing tools, not the amount of context available to reason over."}, {"letter": "B", "text": "Splitting into subagents reduces the total number of API calls made across the whole pipeline, which is the main source of the earlier selection errors", "correct": false, "explanation": "Splitting into subagents does not inherently reduce total API calls across the pipeline, and call count was never identified as the cause of the selection errors, so this doesn't explain the reliability benefit."}, {"letter": "C", "text": "The orchestrator caches tool results across subagents, which eliminates the need for most subagents to make tool calls at all", "correct": false, "explanation": "The scenario describes no caching mechanism, and eliminating tool calls entirely is not the goal; each subagent still needs to call its own scoped tools to do its job."}, {"letter": "D", "text": "Each subagent now discriminates among a much smaller candidate set, which directly reduces the decision complexity that was degrading tool selection at 20 tools", "correct": true, "explanation": "The core principle is that decision complexity scales with the number of tools an agent must choose among; scoping each subagent to roughly 4 tools directly shrinks that candidate set and restores the selection reliability that degraded at 20 tools."}], "correct": "D", "select": 1, "group": "B"}, {"id": "f2-043", "domain": 2, "task_id": "2.4", "objective": "Integrate MCP servers into Claude Code and agent workflows", "situation": "Priya has built a brand-new, still-unstable MCP server while prototyping in one repository on her workstation. She wants Claude Code to load the server only in that repository, keep it out of the other projects on her machine, and never expose it to her teammates.", "question": "Which MCP installation scope satisfies all of these constraints?", "options": [{"letter": "A", "text": "User scope, because it also lives in ~/.claude.json but then loads in every project on her machine", "correct": false, "explanation": "User scope writes to the same ~/.claude.json file, but the documentation says user-scoped servers provide cross-project accessibility and are available across all projects on your machine, so the still-unstable server would follow Priya into every other project she opens."}, {"letter": "B", "text": "Project scope, because it writes .mcp.json in the repository root for the whole team to share once committed", "correct": false, "explanation": "Project scope stores the configuration in a .mcp.json file at the project root, and the documentation says to check that file into version control so everyone on the team gets the same MCP tools and services. That is precisely the teammate exposure Priya wants to avoid."}, {"letter": "C", "text": "Local scope, because Claude Code stores it in ~/.claude.json under that project's path and keeps it private", "correct": true, "explanation": "The scope table in Claude Code's MCP documentation lists Local as loading in the current project only and not shared with the team, and the prose adds that Claude Code stores it in ~/.claude.json under that project's path, so the same server will not appear in your other projects. Sitting in the home-directory file does not make the entry user-scoped: it is nested under projects, then that project's path, then mcpServers, which is what confines it to the one repository."}, {"letter": "D", "text": "Enterprise managed settings, because organization policy is the only way to confine a server to one named project", "correct": false, "explanation": "Managed settings are an organization policy tier that reaches every project on every machine the organization deploys them to, which is the wrong granularity for one person's single prototype, and the MCP installation scopes the documentation defines are Local, Project and User."}], "correct": "C", "select": 1, "group": "J"}, {"id": "f2-044", "domain": 2, "task_id": "2.4", "objective": "Integrate MCP servers into Claude Code and agent workflows", "situation": "A connected MCP server named docs exposes a resource for the authentication guide. A developer wants to have Claude directly analyze that specific document as part of their prompt, the same way they would reference a local file.", "question": "What is the correct way to do this?", "options": [{"letter": "A", "text": "Ask the agent in plain language to \"open the docs server\" and trust that it infers which specific resource is relevant without any reference syntax", "correct": false, "explanation": "Vague natural-language requests don't reliably resolve to one specific resource; the @ mention syntax exists precisely so a specific resource can be referenced unambiguously."}, {"letter": "B", "text": "Add a resources field naming the document inside .mcp.json so it loads automatically into context at the start of every session", "correct": false, "explanation": "There is no .mcp.json resources field for auto-loading specific documents at session start; resources are discovered from the connected server and referenced on demand via @ mentions."}, {"letter": "C", "text": "Call the server's list_resources tool manually first, then paste the raw JSON result from that call directly into the next prompt", "correct": false, "explanation": "Claude Code already provides tools to list and read MCP resources automatically when referenced; manually invoking a list call and pasting raw output is unnecessary and bypasses the intended @ mention workflow."}, {"letter": "D", "text": "Type an @ mention in the prescribed form, such as @docs:file://api/authentication, to reference that exact resource inline in the prompt", "correct": true, "explanation": "Resources are referenced with @ mentions in the form @server:protocol://resource/path, letting a developer point at a specific resource inline just like referencing a file."}], "correct": "D", "select": 1, "group": "J"}, {"id": "f2-045", "domain": 2, "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "situation": "A finance agent has generate_summary_report and generate_detailed_report, both described only as \"Generates a report for the given account.\" An analyst asks for \"a quick overview,\" and the model sometimes invokes the detailed variant instead.", "question": "What long-term fix best prevents this pattern from recurring on new tools?", "options": [{"letter": "A", "text": "Require the analyst to always type the exact internal tool name into every request instead of relying on the model to choose.", "correct": false, "explanation": "Requiring the analyst to name the tool defeats the purpose of natural-language tool selection and does not scale to future tools with similar ambiguity."}, {"letter": "B", "text": "Rename both tools to have identical names distinguished only by a numeric version suffix, so a routing layer disambiguates them.", "correct": false, "explanation": "Identical names with only a version suffix give the model no semantic information about which one produces a summary versus a detailed report."}, {"letter": "C", "text": "Establish a description template requiring tools to state output granularity, an example query, and how they differ from similar ones.", "correct": true, "explanation": "A description template that mandates output granularity, an example query, and explicit differentiation from similarly named tools prevents this ambiguity from recurring as new tools are added."}, {"letter": "D", "text": "Cap the number of tools available to the agent at two, so the model always faces a simple binary choice between them.", "correct": false, "explanation": "Limiting the tool count to two does not resolve the description ambiguity between the two report tools already causing this issue."}], "correct": "C", "select": 1, "group": "J"}, {"id": "f2-046", "domain": 2, "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "situation": "A coordinator agent delegates a three-step data migration to a subagent: extract, transform, and load, but the load step fails twice on a database connection reset, a known transient condition, before finally succeeding on the third attempt inside the subagent's own execution.", "question": "What should the subagent report back to the coordinator?", "options": [{"letter": "A", "text": "An isError: true result describing both connection resets in detail, so the coordinator can decide independently whether the migration should be retried", "correct": false, "explanation": "The MCP specification defines isError: true for reporting a tool execution that failed, giving actionable feedback the model can use to self-correct and retry — an outcome that still needs to be acted on. By the time the subagent reports back, the load step has already succeeded, so marking the result as an error would misrepresent the actual, successful outcome. The Agent SDK's subagent design also keeps intermediate tool calls and results inside the subagent's own context; only its final message returns to the parent, so surfacing every internal retry to the coordinator defeats the purpose of delegating the step in the first place."}, {"letter": "B", "text": "A partial-results payload listing only the extract and transform steps as done, omitting the load step entirely since it initially failed twice", "correct": false, "explanation": "The load step finished successfully on the third attempt, so a report that omits it entirely would understate what was actually accomplished and could cause the coordinator to unnecessarily re-run a step that already completed."}, {"letter": "C", "text": "An escalation asking the coordinator to obtain new database credentials, since two consecutive connection resets indicate the credentials have expired", "correct": false, "explanation": "A connection reset is a transient network-level failure, not an authentication failure; expired credentials would produce a consistent authentication error on every attempt rather than a reset that clears on retry. Escalating for new credentials here targets the wrong cause and would delay a migration that had already completed."}, {"letter": "D", "text": "A success result summarizing the completed migration, since the transient failures were resolved locally and never needed to surface above the subagent", "correct": true, "explanation": "The Claude Agent SDK's subagent design keeps intermediate tool calls and results inside the subagent's own execution: only the subagent's final message returns to the parent, so a subagent can work through a transient failure without that internal detail flowing up to the coordinator's context. Because the load step ultimately succeeded, reporting a success result reflects the actual final state accurately. Flagging the tool result as failed here would misrepresent a step that succeeded, since MCP's isError flag is defined for reporting an execution that still needs correction, not one that already recovered before returning."}], "correct": "D", "select": 1, "group": "J"}, {"id": "f2-047", "domain": 2, "task_id": "2.3", "objective": "Distribute tools appropriately across agents and configure tool choice", "situation": "A data-pipeline monitoring agent must always respond with a structured action (acknowledge_alert, escalate_alert, or suppress_alert) for every incoming alert, since downstream automation parses only tool calls and cannot handle free-text replies.", "question": "Which tool_choice value directly guarantees the model will not return plain conversational text?", "options": [{"letter": "A", "text": "Leaving tool_choice unset while adding a system prompt instruction to \"always respond with a tool call\"", "correct": false, "explanation": "Leaving tool_choice unset defaults to \"auto\" and relies on prompt wording alone, which is not a guarantee; the model can still ignore the instruction and reply in text."}, {"letter": "B", "text": "{\"type\": \"any\"}, since it requires the model to call one of the provided tools rather than reply in prose", "correct": true, "explanation": "Setting tool_choice to {\"type\": \"any\"} guarantees the model calls one of the available tools rather than returning conversational text, which is exactly what downstream automation that only parses tool calls requires."}, {"letter": "C", "text": "{\"type\": \"auto\"}, since it is the default and applies whenever any tools are present in the request", "correct": false, "explanation": "\"auto\" is the default behavior and explicitly allows the model to respond directly in text when it judges a tool call unnecessary, which does not meet the requirement to always produce a structured action."}, {"letter": "D", "text": "{\"type\": \"none\"}, since it disables prose generation and forces structured output by default", "correct": false, "explanation": "\"none\" prevents any tool calls at all, which is the opposite of what's needed; it would force the model to reply only in prose, breaking the downstream automation."}], "correct": "B", "select": 1, "group": "B"}, {"id": "f2-048", "domain": 2, "task_id": "2.3", "objective": "Distribute tools appropriately across agents and configure tool choice", "situation": "An architect is designing tool access for a three-agent pipeline: an intake agent, a processing agent, and a delivery agent. The intake agent occasionally needs to check processing status, which is normally a processing-agent operation.", "question": "Following the principle of scoped tool access with limited cross-role tools, how should the architect handle this?", "options": [{"letter": "A", "text": "Remove status checking from the pipeline, redesigning the three-agent workflow so the intake agent never requires a tool outside its core responsibility of accepting intakes.", "correct": false, "explanation": "This is an overly rigid interpretation of role separation. Status checking may be a legitimate operational need for the intake agent (e.g., to provide user updates). Eliminating the capability entirely could degrade the user experience. Anthropic's patterns support limited cross-role tool access when justified, rather than forcing pure isolation at all costs."}, {"letter": "B", "text": "Give the intake agent only a narrow check_status tool for that specific occasional need, while routing deeper processing operations through the processing agent.", "correct": true, "explanation": "This follows the principle of scoped, limited cross-role tool access: rather than routing every occasional status check through the processing agent, the intake agent is given a single narrow, specialized tool for that specific occasional need, while deeper processing operations stay with the processing agent. Granting this narrowly scoped, purpose-built access maintains the principle of least privilege and avoids the token overhead of a broader tool grant. Refer to Anthropic's guidance on building effective agents and tool design best practices."}, {"letter": "C", "text": "Give the processing agent a copy of the intake agent's status check tool, so that either agent can independently perform status checks without routing through the processing agent's operations.", "correct": false, "explanation": "This option is redundant and misaligned with the scenario. The status check is already a processing-agent operation; duplicating the tool on the processing agent serves no purpose. The requirement is for the intake agent to occasionally have query capability, not to equip the processing agent with additional tools. The correct approach is to give the intake agent a narrow, read-only tool, not to mirror tools between agents."}, {"letter": "D", "text": "Give the intake agent the processing agent's full tool set, enabling it to directly perform status checks and execute any processing operation as part of its intake workflow.", "correct": false, "explanation": "This violates the principle of scoped tool access. Granting the full processing tool set to the intake agent blurs agent responsibilities, increases the risk of unintended actions, and leads to unnecessary token consumption. Anthropic recommends a deny-by-default tool configuration and granting only the minimum necessary tools to each agent."}], "correct": "B", "select": 1, "group": "B"}, {"id": "f2-049", "domain": 2, "task_id": "2.4", "objective": "Integrate MCP servers into Claude Code and agent workflows", "situation": "An architect is scoping the tool integration work for a new project. The team needs to connect Claude Code to their Jira instance for standard issue read/write operations, and separately needs a way to trigger their in-house deployment pipeline, which has no public equivalent.", "question": "How should the architect approach these two needs?", "options": [{"letter": "A", "text": "Build custom MCP servers for both Jira and the deployment pipeline, since only in-house servers can be trusted with production credentials", "correct": false, "explanation": "Building a custom server for a standard integration like Jira duplicates effort that an existing, maintained community server already covers well, and trust is established through review, not exclusively through in-house authorship."}, {"letter": "B", "text": "Adopt an existing community or vendor Jira MCP server for the standard integration, and reserve custom server work for the team-specific pipeline", "correct": true, "explanation": "Standard integrations like Jira are well served by existing community or vendor MCP servers, while custom server development is best reserved for team-specific workflows that have no standard equivalent, like an in-house deployment pipeline."}, {"letter": "C", "text": "Adopt a community MCP server for the deployment pipeline, since third-party servers already understand internal tooling, and build a custom one for Jira instead", "correct": false, "explanation": "This reverses the guidance: no generic community server exists for a team's proprietary deployment pipeline, and Jira is exactly the kind of standard integration a community server already handles."}, {"letter": "D", "text": "Skip MCP for both needs and instead give the agent direct Bash access to curl the Jira API and run the deployment scripts by hand", "correct": false, "explanation": "Bypassing MCP with raw Bash/curl calls loses the structured tool discovery, credential handling, and reusability that MCP servers provide, and is not the recommended integration pattern."}], "correct": "B", "select": 1, "group": "J"}, {"id": "f2-050", "domain": 2, "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "situation": "A subagent handling document translation calls a translate_text tool that fails with a connection timeout on the first attempt. The subagent's local retry policy allows up to 2 automatic retries for transient errors before escalating. The second retry also times out.", "question": "What should the subagent do next?", "options": [{"letter": "A", "text": "Stop retrying and propagate the failure to the coordinator, noting the errorCategory, that timeouts persisted across the allowed local attempts, and what text remained untranslated.", "correct": true, "explanation": "After exhausting the local retry budget for a transient timeout error, the subagent should stop retrying and escalate the failure to the coordinator. It must include the errorCategory, note that timeouts persisted across the allowed local attempts, and specify what text remained untranslated. This structured context enables the coordinator to make an informed higher-level decision."}, {"letter": "B", "text": "Immediately reclassify the error as a permission failure and propagate it to the coordinator, including the permission flag and that repeated timeouts indicate an access-control issue, while aborting further retries.", "correct": false, "explanation": "Repeated connection timeouts strongly suggest a transient network or service issue, not an access-control failure. Reclassifying the error as a permission problem without any supporting evidence would mislead the coordinator, likely triggering irrelevant permission checks while ignoring the actual transient unavailability."}, {"letter": "C", "text": "Silently return a fabricated translation to the coordinator, using a cached fallback response from a prior successful call to avoid workflow interruption, despite the translate_text tool timeouts.", "correct": false, "explanation": "Silently returning a fabricated or cached translation misrepresents the outcome of a failed tool call as successful, potentially corrupting downstream workflows with invalid data. An honest, structured escalation is always preferable to masking failures with unverified fallbacks."}, {"letter": "D", "text": "Continue retrying indefinitely at the subagent level, logging each attempt and resetting the retry count, ensuring that transient errors like timeouts are never escalated to the coordinator.", "correct": false, "explanation": "Local recovery is explicitly bounded by the retry policy; retrying indefinitely after the cap is reached violates that policy and wastes resources. It also forecloses the coordinator's ability to respond with broader strategies, such as switching to an alternative service or notifying an operator."}], "correct": "A", "select": 1, "group": "J"}, {"id": "f2-051", "domain": 2, "task_id": "2.5", "objective": "Select and apply built-in tools (Read, Write, Edit, Bash, Grep, Glob) effectively", "situation": "An architect wants Claude Code to find every place in a codebase that throws an error containing the phrase \"insufficient permissions\", including messages built by concatenating a string literal across two lines with a plus sign. A single-line search pattern is not catching the split-line cases.", "question": "What should Claude do?", "options": [{"letter": "A", "text": "Re-run Glob with the recursive ** pattern to search deeper into subdirectories where the missed messages might be hiding", "correct": false, "explanation": "The recursive ** pattern controls how deep Glob searches directory structure for file names; it has nothing to do with whether Grep can match content spanning multiple lines, and Glob does not search file contents at all."}, {"letter": "B", "text": "Re-run Grep with multiline mode enabled so the pattern can match text that spans across the line boundary", "correct": true, "explanation": "Grep matches within a single line by default; enabling multiline mode allows the pattern to match text that spans across line boundaries, which is exactly what is needed to catch a message literal split across two lines."}, {"letter": "C", "text": "Switch to Read on every file in the repository so the full text of each file is visible instead of only matching lines", "correct": false, "explanation": "Reading every file in the repository to visually scan for a phrase is far less efficient than fixing the search itself, and does not address the actual reason the split-line messages were missed."}, {"letter": "D", "text": "Re-run Grep with output mode count instead of content, since count mode is described as searching more thoroughly than content mode", "correct": false, "explanation": "Output mode only changes what information Grep returns about matches (counts versus content versus file paths); it does not change which matches are found, so switching modes would not fix the multiline matching problem."}], "correct": "B", "select": 1, "group": "F"}, {"id": "f2-052", "domain": 2, "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "situation": "A get_customer_orders MCP tool is called for a customer who exists in the system but has placed zero orders. Separately, the same tool is called with a customer ID that does not exist in the database at all.", "question": "How should these two outcomes be reported so the agent can respond correctly in each case?", "options": [{"letter": "A", "text": "Both cases return isError:false with an empty orders array, and the agent should proceed without error handling for both scenarios, as the tool correctly reports the absence of orders in each case.", "correct": false, "explanation": "This conflates a valid empty result with a missing customer record. A nonexistent customer ID indicates the requested resource was not found and should be reported as an error so the agent can ask for a corrected ID rather than silently treating it as having no orders."}, {"letter": "B", "text": "Both cases return isError:true with errorCategory transient, and the agent should schedule periodic retries for these calls to account for possible delayed order placement or customer record creation.", "correct": false, "explanation": "A customer with zero orders is a valid successful response, not an error. A nonexistent customer ID is also not transient; retrying will not make the ID appear if it is invalid. The correct distinction is empty success versus not_found_error."}, {"letter": "C", "text": "The zero-orders case returns isError:false with an empty order array; the nonexistent-ID case returns isError:true with errorCategory not_found_error and a description that the customer ID was not found.", "correct": true, "explanation": "This is the recommended split. A successful lookup with no orders is a valid empty collection, so it should be reported as isError:false with an empty array. A customer ID that does not exist is a missing resource, not an empty result; Anthropic's resource-not-found guidance and REST best practice use a not_found_error type or HTTP 404 with a clear message. Reporting it as a validation error is less precise than not_found_error."}, {"letter": "D", "text": "The zero-orders case returns isError:true with errorCategory suspicious and a description noting the empty result; the nonexistent-ID case returns isError:false with an empty orders array, treating the missing ID as an empty result.", "correct": false, "explanation": "This reverses the correct outcomes. Zero orders should be isError:false with an empty array, while a missing customer ID should be reported as isError:true with a not_found_error category so the agent can distinguish an empty order history from an invalid customer reference."}], "correct": "C", "select": 1, "group": "J"}, {"id": "f2-053", "domain": 2, "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "situation": "A charge_card MCP tool receives a request with an amount field formatted as \"$45.00\" instead of a numeric type as specified by its input schema.", "question": "Before the tool's handler logic even runs, how does the MCP client typically surface this failure, and how should that differ from the tool later reporting a declined charge?", "options": [{"letter": "A", "text": "Both failures are reported identically as JSON-RPC protocol errors with code -32602 (Invalid params), because the client validates the request against the schema before calling the tool handler.", "correct": false, "explanation": "The malformed argument does trigger a JSON-RPC error, but a declined charge is not a protocol error; it's an execution-time business failure reported via isError: true in the tool result. MCP separates these concerns."}, {"letter": "B", "text": "Both failures are reported inside a tool result with isError:true, because protocol errors are reserved for unknown tool names; all other issues, like malformed arguments or declined charges, appear as tool-level errors.", "correct": false, "explanation": "Invalid arguments cause JSON-RPC protocol errors, not tool-level errors with isError: true. Protocol errors include invalid params and unknown methods, while declined charges are business outcomes reported inside the result, not as JSON-RPC errors."}, {"letter": "C", "text": "The malformed argument triggers a JSON-RPC protocol error from schema validation before the tool executes, while a declined charge is reported inside the tool result with isError:true.", "correct": true, "explanation": "The malformed argument violates the tool's input schema, causing a JSON-RPC invalid params error before handler execution. A declined charge, however, occurs during tool execution and is reported through the tool result with isError: true, keeping protocol and execution errors distinct."}, {"letter": "D", "text": "The client silently coerces the malformed argument to a number before invocation, so neither a schema validation error nor a declined charge occurs; the handler receives a valid amount, and any decline is a business result.", "correct": false, "explanation": "MCP clients validate input against the schema strictly; a string like \"$45.00\" when a number is expected will cause a validation error, not silent coercion. The handler would not receive a valid amount, and no business result from a declined charge would occur."}], "correct": "C", "select": 1, "group": "J"}, {"id": "f2-054", "domain": 2, "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "situation": "A QA engineer notices that a coding assistant consistently picks run_linter instead of run_type_checker when a user says \"check my code for issues,\" even though the user meant type errors.", "question": "Both tools' descriptions read \"Checks code for issues.\" What is the most targeted fix?", "options": [{"letter": "A", "text": "Instruct users to say \"lint\" or \"type check\" explicitly every time, since tool descriptions cannot influence this kind of ambiguity.", "correct": false, "explanation": "This approach shifts the burden entirely onto users and contradicts the premise that clear tool descriptions can resolve this kind of ambiguity."}, {"letter": "B", "text": "Merge the two tools' outputs into a single combined report and always run both regardless of what the user actually asked for.", "correct": false, "explanation": "Always running both tools wastes execution on the category the user did not ask about and avoids fixing the underlying description ambiguity."}, {"letter": "C", "text": "Specify in each description the exact category of issue each tool detects, along with an example phrase associated with each one.", "correct": true, "explanation": "Naming the specific issue category each tool detects and giving an associated example phrase gives the model the differentiating signal needed to route type-error requests correctly."}, {"letter": "D", "text": "Reduce the number of code-checking tools available to the assistant to one, since offering more than one inherently causes confusion.", "correct": false, "explanation": "Removing one of the two distinct capabilities eliminates functionality rather than clarifying which tool applies to which category of issue."}], "correct": "C", "select": 1, "group": "J"}, {"id": "f2-055", "domain": 2, "task_id": "2.4", "objective": "Integrate MCP servers into Claude Code and agent workflows", "situation": "A project has three MCP servers configured: a Postgres server, a Sentry server, and a Slack server, all healthy and connected. Mid-session, the agent is asked to pull a database schema, cross-reference a recent Sentry error, and post a summary to Slack in one request.", "question": "What should the architect expect about tool availability for this request?", "options": [{"letter": "A", "text": "Tools from all three connected servers are available to the agent simultaneously, so it can call across Postgres, Sentry, and Slack tools within the same turn as needed", "correct": true, "explanation": "Tools from every configured and connected MCP server are discovered and made available together, so the agent can freely call tools across multiple servers within one request."}, {"letter": "B", "text": "Only the tools from the server whose name the prompt most closely matches will be loaded, so the agent must be told explicitly which server to prioritize", "correct": false, "explanation": "Claude Code does not restrict tool availability to a single best-matching server; all connected servers' tools are simultaneously available regardless of prompt wording."}, {"letter": "C", "text": "Only one MCP server can hold an active connection at any given time, so the developer must disconnect Postgres before Sentry's tools will respond to a call", "correct": false, "explanation": "Multiple MCP servers can be connected and active at once; there is no single-active-connection restriction that forces disconnecting one server to use another."}, {"letter": "D", "text": "The agent must complete every Postgres tool call and receive its results before Sentry and Slack tools become selectable at all during that same turn", "correct": false, "explanation": "There is no requirement to sequence tool calls strictly by server; the agent can interleave calls to different servers' tools as the task demands."}], "correct": "A", "select": 1, "group": "J"}, {"id": "f2-056", "domain": 2, "task_id": "2.3", "objective": "Distribute tools appropriately across agents and configure tool choice", "situation": "An onboarding agent walks a new user through account setup and needs to call create_profile immediately as the very first action of the conversation, before considering any other tool such as send_welcome_email or assign_default_settings. After that first call, the agent must be free either to call one of the remaining tools or to reply to the user in plain text, depending on what the user says.", "question": "What is the best way to configure this?", "options": [{"letter": "A", "text": "Use tool_choice: {\"type\": \"any\"} for the entire conversation so the model is always forced to pick from create_profile, send_welcome_email, and assign_default_settings", "correct": false, "explanation": "The any setting requires the model to use one of the provided tools without forcing a particular one, so it never guarantees that create_profile runs first. It also prefills a tool call on every turn, so the agent could never respond to the user in plain text."}, {"letter": "B", "text": "Use tool_choice: {\"type\": \"none\"} for the first turn so the model cannot call any tool, then switch to {\"type\": \"auto\"} afterward", "correct": false, "explanation": "The none setting prevents Claude from using any tools, so blocking tool calls on the first turn also blocks create_profile, the very action that is required to happen first."}, {"letter": "C", "text": "Order create_profile first in the tools array and leave tool_choice at {\"type\": \"auto\"} for the whole conversation", "correct": false, "explanation": "Position in the tools array is not a documented mechanism for controlling which tool is called first. Under auto the model decides whether to call a tool at all and which one, so create_profile is not guaranteed to be the opening action."}, {"letter": "D", "text": "Use tool_choice: {\"type\": \"tool\", \"name\": \"create_profile\"} for the first turn only, then switch to tool_choice: {\"type\": \"auto\"} for subsequent turns", "correct": true, "explanation": "Naming the tool in tool_choice forces that specific tool, so create_profile is guaranteed to run as the opening action. Switching to auto afterward is what the later turns need: auto lets Claude decide whether to call any provided tool or not, so the agent can invoke a remaining tool when the user's reply calls for one and answer in plain text when it does not."}], "correct": "D", "select": 1, "group": "B"}, {"id": "f2-057", "domain": 2, "task_id": "2.4", "objective": "Integrate MCP servers into Claude Code and agent workflows", "situation": "An agent working against a large issue tracker keeps issuing many exploratory search-tool calls just to figure out which issues exist before it can act on any of them. The team wants to cut down on this exploratory overhead.", "question": "What MCP capability addresses this directly?", "options": [{"letter": "A", "text": "Increase the tool-call timeout on the issue tracker server so that each search call retrieves more issues and the agent gets a complete view with fewer queries.", "correct": false, "explanation": "Increasing the tool-call timeout only affects how long a call can run; it does not change the number of exploratory calls the agent must make. The agent would still need multiple searches to discover issues, failing to address the root cause."}, {"letter": "B", "text": "Have the MCP server expose an issue-summary catalog as an MCP resource, so the agent can see available issues up front without needing any repeated searches.", "correct": true, "explanation": "An MCP resource can serve as a catalog of issue summaries, giving the agent upfront visibility into the issue tracker without needing to perform any exploratory tool calls. This directly reduces overhead by eliminating repeated searches."}, {"letter": "C", "text": "Switch the issue tracker server's transport from stdio to HTTP so that each query completes faster, and the agent can gather all necessary information with fewer calls.", "correct": false, "explanation": "Switching transport from stdio to HTTP may improve latency, but it does not provide the agent with upfront knowledge of available issues. The agent would still need to explore through search calls, so the exploratory overhead remains."}, {"letter": "D", "text": "Add several more search-related tools to the server so the agent can use targeted queries to find relevant issues directly and avoid unnecessary exploratory searches.", "correct": false, "explanation": "Adding more search tools enables more targeted queries but does not eliminate the need to discover what issues exist in the first place. The agent would still have to exploratively search, whereas a resource catalog provides immediate visibility."}], "correct": "B", "select": 1, "group": "J"}, {"id": "f2-058", "domain": 2, "task_id": "2.3", "objective": "Distribute tools appropriately across agents and configure tool choice", "situation": "A code-review subagent was originally scoped to Read, Grep, and Glob so it could inspect a codebase without changing it. During a refactor, an engineer also grants it Bash and a deploy_service tool because \"it might be handy.\" Shortly after, the review agent begins running deploy_service mid-review on branches that haven't been approved.", "question": "What principle explains this outcome and what is the correct fix?", "options": [{"letter": "A", "text": "Tools beyond an agent's specialization tend to get misused when available, so Bash and deploy_service should be removed, restricting the review agent to the read-only tools its role requires.", "correct": true, "explanation": "This is a direct instance of an agent misusing a tool outside its specialization. The fix is to remove Bash and deploy_service, restricting the review agent to the read-only tools (Read, Grep, Glob) that match its role, following the principle of least privilege."}, {"letter": "B", "text": "The deploy_service tool's input schema is likely malformed, so the fix is to add stricter JSON schema validation on its parameters so that the agent only submits deployment requests for branches that have passed review.", "correct": false, "explanation": "Schema validation only ensures that tool call parameters are well-formed, but it does not address whether the agent should have access to the deploy tool at all. The underlying problem is scoping, not input validation."}, {"letter": "C", "text": "The review agent should be given even more tools, such as advanced file-search and environment-status utilities, so that deploy_service becomes just one option among many and is therefore less likely to be selected inadvertently.", "correct": false, "explanation": "Adding more tools increases complexity and does not reduce the likelihood of mis‑selection; it may actually make the agent more likely to use an inappropriate tool. The correct approach is to limit tools to only those needed for the task."}, {"letter": "D", "text": "The review agent's system prompt needs a stronger instruction telling it never to deploy, while keeping all five tools available so that it can still use Bash for code inspection but is prevented from running deploy_service.", "correct": false, "explanation": "Relying solely on prompt instructions to prevent misuse is fragile because the agent may still accidentally or adversarially invoke the tool. The scoped-access principle dictates that tools outside the agent's role should be removed, not just discouraged."}], "correct": "A", "select": 1, "group": "B"}, {"id": "f2-059", "domain": 2, "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "situation": "A tool named analyze_content, originally built for summarizing pasted text, has its description copied almost verbatim onto a newer tool, analyze_document, meant for uploaded PDFs.", "question": "Following the pattern of renaming a tool to eliminate overlap, what should the team do?", "options": [{"letter": "A", "text": "Delete analyze_document and require users to always paste document contents in as plain text from now on.", "correct": false, "explanation": "This approach would remove the ability to leverage the Files API for structured document handling and goes against Anthropic's recommendation to differentiate tools by purpose, not eliminate necessary functionality."}, {"letter": "B", "text": "Rename analyze_content to a more specific name, such as summarize_pasted_text, and limit its description to pasted text only, while analyze_document handles uploaded PDFs.", "correct": true, "explanation": "Per CCAR-F guidance, tools must have distinct, non-overlapping descriptions to prevent model confusion. The study guide explicitly warns: \"Avoid identical or overlapping descriptions across tools.\" Renaming and scoping the tool to its specific input type eliminates overlap with the PDF-focused analyze_document, following the recommended pattern of clarifying tool boundaries."}, {"letter": "C", "text": "Give both tools identical descriptions but different internal function names so backend routing can disambiguate them.", "correct": false, "explanation": "Anthropic's tool design best practices warn against identical or overlapping descriptions because the model may confuse them; internal function names are not used for selection, so this fails to prevent misrouting."}, {"letter": "D", "text": "Add a disclaimer inside each tool's error message stating the wrong tool may have been called, shown after execution fails.", "correct": false, "explanation": "This is a reactive measure that does not prevent misrouting at the point of tool selection. Proactively eliminating description overlap is the correct and recommended approach."}], "correct": "B", "select": 1, "group": "J"}, {"id": "f2-060", "domain": 2, "task_id": "2.3", "objective": "Distribute tools appropriately across agents and configure tool choice", "situation": "A developer is building an assistant that must force tool use on every user turn. All requests require either current data or an action, so the model must never answer directly from its own knowledge.", "question": "Which tool_choice configuration matches this requirement?", "options": [{"letter": "A", "text": "tool_choice: {\"type\": \"tool\", \"name\": \"answer_directly\"}, which forces that one fixed response-generation tool to run on every turn", "correct": false, "explanation": "{\"type\": \"tool\", \"name\": \"...\"} forces a specific named tool, but answer_directly is not a standard Anthropic tool and this option would not let Claude choose from the available tools. Forcing an arbitrary tool does not fulfill the requirement to select an appropriate tool on each turn."}, {"letter": "B", "text": "tool_choice: {\"type\": \"auto\"}, letting the model decide per turn whether to call a tool or respond directly", "correct": false, "explanation": "auto is the default and allows Claude to answer directly when it judges the request does not require a tool. The requirement is to force a tool call on every turn, so allowing direct responses does not meet the requirement."}, {"letter": "C", "text": "tool_choice: {\"type\": \"any\"}, requiring the model to call one of the provided tools on every turn", "correct": true, "explanation": "{\"type\": \"any\"} forces Claude to use at least one of the provided tools on every turn while letting it choose the most appropriate tool. This matches the requirement that the model must never answer directly and must always call a tool for current data or an action."}, {"letter": "D", "text": "tool_choice: {\"type\": \"none\"}, so the model can only describe what it would do but never actually call a tool", "correct": false, "explanation": "{\"type\": \"none\"} prevents all tool calls. This directly contradicts the requirement that the model must force tool usage and call a tool on every turn."}], "correct": "C", "select": 1, "group": "B"}, {"id": "f2-061", "domain": 2, "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "situation": "A book_meeting_room MCP tool fails because the requested room is already reserved for the requested time slot. The tool author wants the agent to be able to explain the conflict to the user in natural language and suggest picking a different time, without the agent needing to parse a raw exception message.", "question": "Which element of the structured error response most directly enables this?", "options": [{"letter": "A", "text": "Omitting the description field entirely, since errorCategory alone is sufficient for the agent to generate an accurate, context-specific explanation", "correct": false, "explanation": "errorCategory alone (e.g., \"business\") tells the agent the type of failure but not the specific detail, like which room or time was in conflict, that a customer-facing explanation actually needs."}, {"letter": "B", "text": "Returning the raw exception stack trace from the scheduling library so the agent can extract the room name using text parsing", "correct": false, "explanation": "Forcing the agent to parse a raw stack trace to extract meaning is precisely the brittle, unstructured approach that a dedicated human-readable description field is meant to replace."}, {"letter": "C", "text": "Setting isRetryable to true so the agent automatically resubmits the identical booking request until the room becomes free on its own", "correct": false, "explanation": "A double-booking conflict for the same fixed time slot won't resolve itself by retrying the identical request; the room stays booked, so automatic resubmission wastes attempts rather than helping the user pick a new time."}, {"letter": "D", "text": "A human-readable description field stating the room is already booked for that slot, separate from any machine-oriented errorCategory or isRetryable flags", "correct": true, "explanation": "A clear, human-readable description alongside the machine-oriented fields is exactly what lets the agent relay an accurate, natural-language explanation to the user without needing to reverse-engineer meaning from category codes or exception internals."}], "correct": "D", "select": 1, "group": "J"}, {"id": "f2-062", "domain": 2, "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "situation": "An MCP server's create_invoice tool calls a downstream billing API that returns a 503 while the service is deploying. The tool wraps this in a result with isError: true and a text block reading only \"Operation failed.\" The agent retries the same call five times in a row, each time failing the same way, before giving up.", "question": "What is the most direct cause of the wasted retries?", "options": [{"letter": "A", "text": "The result carries no structured metadata distinguishing transient from non-retryable failures. The agent thus has no basis for deciding whether retrying is worthwhile.", "correct": true, "explanation": "The tool returned only a generic text message with isError: true, providing no structured metadata such as an error category or a retryability flag. Without that information, the agent cannot distinguish a transient 503 from a permanent failure and therefore defaults to blind retries, wasting calls."}, {"letter": "B", "text": "The downstream billing API returned an HTTP status code rather than a JSON-RPC error object, so the MCP client could not parse the response and defaulted to retrying repeatedly.", "correct": false, "explanation": "The MCP client only receives the tool result content that the server returns; the raw downstream HTTP status code is not exposed. Once the server translates the 503 into a result with isError: true and a text block, the original status code becomes irrelevant to the client's retry logic."}, {"letter": "C", "text": "The agent's context window ran out of space to store the error text, so it could not remember that the same request had just failed and therefore repeated the call as if it were a new attempt.", "correct": false, "explanation": "There is no indication in the scenario that the agent's context window is exhausted. The agent is retrying the same call immediately in succession, which points to a lack of structured error metadata rather than a failure to remember previous attempts."}, {"letter": "D", "text": "The tool set isError to true instead of false, which signals to the agent that unlimited retries are the correct response and prevents it from recognizing that the error is transient.", "correct": false, "explanation": "Setting isError to true simply marks the tool call as failed; it does not prescribe any retry behavior. The agent does not interpret isError: true as a signal for unlimited retries, so this is not the direct cause of the repeated attempts."}], "correct": "A", "select": 1, "group": "J"}, {"id": "f2-063", "domain": 2, "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "situation": "A developer is writing a new tool description and wants to minimize the chance the model confuses it with an existing similar tool.", "question": "Which combination of elements most reliably differentiates the tools for selection purposes?", "options": [{"letter": "A", "text": "A catchy tool name paired with an emoji so it appears visually distinct within the full list of tools.", "correct": false, "explanation": "A memorable name or emoji provides no functional information about input, scope, or how the tool differs from a similar one."}, {"letter": "B", "text": "Expected input format, one or two example queries the tool handles, and a note on when to prefer the other tool instead.", "correct": true, "explanation": "Input format, representative example queries, and an explicit contrast with the similar tool together give the model the differentiating signal needed for reliable selection."}, {"letter": "C", "text": "The name of the engineer who implemented the tool along with its internal implementation language choice.", "correct": false, "explanation": "Implementation metadata such as author or programming language is irrelevant to how the model reasons about which tool matches a given request."}, {"letter": "D", "text": "An exhaustive list of every possible parameter combination the tool accepts, without any narrative explanation of intended use.", "correct": false, "explanation": "A raw parameter list without narrative context fails to convey when the tool should be chosen over a similar alternative."}], "correct": "B", "select": 1, "group": "J"}, {"id": "f2-064", "domain": 2, "task_id": "2.5", "objective": "Select and apply built-in tools (Read, Write, Edit, Bash, Grep, Glob) effectively", "situation": "Claude Code needs to update a version string inside a nested JSON field in package.json, where the same string coincidentally also appears as a substring inside an unrelated dependency name elsewhere in the file. Claude wants Edit to target only the version field.", "question": "Which old_string design correctly ensures the edit applies to the intended location only?", "options": [{"letter": "A", "text": "Use a regular expression in old_string that matches the version field specifically, since Edit interprets old_string as regex when quotes are present", "correct": false, "explanation": "Edit performs exact string replacement and does not support regex matching in old_string, so a regex pattern would simply be treated as literal text and likely fail to match at all."}, {"letter": "B", "text": "Use only the bare version string as old_string, since Edit automatically infers which occurrence is semantically the version field", "correct": false, "explanation": "Edit does not infer semantic meaning from context; if the bare string appears more than once in the file, the match is not unique and the edit will fail regardless of which occurrence a human would consider the real version field."}, {"letter": "C", "text": "Set replace_all to true so every occurrence of the version string throughout the whole file is updated to the same new value", "correct": false, "explanation": "Setting replace_all would also rewrite the unrelated occurrence inside the dependency name, corrupting a value that was never meant to change."}, {"letter": "D", "text": "Include the surrounding JSON key and adjacent structure, such as the \"version\": prefix and its line, so the string becomes unique to that field", "correct": true, "explanation": "Edit performs exact string matching and requires old_string to appear exactly once; including the \"version\": key alongside the value gives enough surrounding context to disambiguate the field from an unrelated occurrence of the same substring elsewhere in the file."}], "correct": "D", "select": 1, "group": "F"}, {"id": "f2-065", "domain": 2, "task_id": "2.3", "objective": "Distribute tools appropriately across agents and configure tool choice", "situation": "A research agent has a generic fetch_url tool that can retrieve content from any URL, including URLs the model hallucinates or pulls from untrusted parts of a document. This has led to the agent fetching malformed or irrelevant pages. The team wants a scoped alternative that only allows retrieval of legitimate source documents.", "question": "Which change best follows the guidance on replacing generic tools with constrained alternatives?", "options": [{"letter": "A", "text": "Remove URL retrieval from the agent entirely and require a human to paste document contents into the conversation", "correct": false, "explanation": "Eliminating automated retrieval entirely removes a needed capability rather than constraining it appropriately, which is a heavier-handed response than the scoped-tool pattern calls for."}, {"letter": "B", "text": "Keep fetch_url unchanged and add a second, identical tool named fetch_url_v2 as a fallback for failed retrievals", "correct": false, "explanation": "Adding a duplicate tool increases the tool count and decision complexity without adding any validation, so it does not constrain what URLs can be fetched."}, {"letter": "C", "text": "Replace fetch_url with a load_document tool that validates the URL against an allowed document source before retrieving it", "correct": true, "explanation": "This mirrors the documented pattern of replacing a generic, unconstrained tool with a narrower one that validates its input, such as swapping fetch_url for a load_document tool that checks the URL against known-good sources before fetching."}, {"letter": "D", "text": "Keep fetch_url but rename it to get_document so its purpose is clearer to the model when selecting a tool", "correct": false, "explanation": "A rename alone changes the label but not the underlying behavior; the tool still accepts arbitrary URLs without validation, so the same misuse pattern would continue."}], "correct": "C", "select": 1, "group": "B"}, {"id": "f2-066", "domain": 2, "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "situation": "An agent has both translate_text and localize_content, where the latter also adjusts currency, dates, and cultural references beyond direct translation.", "question": "Both descriptions currently read \"Converts text between languages.\" Which fix best restores reliable routing?", "options": [{"letter": "A", "text": "Merge both tools into one and let the model pass a boolean localization flag, keeping the merged description as short as possible.", "correct": false, "explanation": "Merging reduces tool count but does not directly fix the root cause of identical descriptions. Anthropic recommends clear, specific descriptions to avoid ambiguity; a boolean flag inside a single merged tool can still lead to routing confusion if the description isn’t detailed. The option adds unnecessary complexity and isn’t the simplest best practice per research."}, {"letter": "B", "text": "Rename translate_text to localize_content_v1 so the model treats it as a deprecated but still valid alternative.", "correct": false, "explanation": "Renaming without clarifying functional differences fails to resolve the routing issue. The model might avoid the deprecated tool or interpret it incorrectly without distinct descriptions. Anthropic guidance prioritizes precise, descriptive tool definitions over superficial naming changes."}, {"letter": "C", "text": "Remove translate_text from the toolset entirely so only localize_content remains for any language task at all.", "correct": false, "explanation": "Removing the direct‑translation tool forces the agent to always apply localization, which may be unnecessary and reduces flexibility. While the model can theoretically perform translation via the remaining tool, the design loses the ability to explicitly request simple translations. Anthropic’s approach supports keeping distinct tools with clear, separate descriptions for reliable routing."}, {"letter": "D", "text": "Revise localize_content's description to mention currency and culture, and state translate_text does direct translation only.", "correct": true, "explanation": "Differentiating tool descriptions aligns with Anthropic's guidance for tool use in agentic systems. Research emphasizes providing detailed, distinct descriptions so the model can correctly route requests. By clarifying that translate_text performs only direct translation and localize_content handles cultural and currency adaptation, the agent reliably selects the appropriate tool based on user intent."}], "correct": "D", "select": 1, "group": "J"}, {"id": "f2-067", "domain": 2, "task_id": "2.4", "objective": "Integrate MCP servers into Claude Code and agent workflows", "situation": "A developer wants to try out an experimental local MCP server that queries their personal Notion workspace. They do not want it to appear for any other teammate, and they want it available whenever they open any project on their own machine.", "question": "Which configuration achieves this?", "options": [{"letter": "A", "text": "Add the server with project scope so the entry is written to .mcp.json and stays private until the developer marks it as personal-only", "correct": false, "explanation": "Project scope writes to .mcp.json specifically so the server is shared with the whole team via version control; there is no per-server flag to keep a project-scoped entry personal."}, {"letter": "B", "text": "Add the server with user scope so the entry is written to ~/.claude.json and loads across every project on that machine without being shared", "correct": true, "explanation": "User scope stores the server in ~/.claude.json and makes it available across all of that user's projects while remaining private to their account, matching cross-project personal use."}, {"letter": "C", "text": "Add the server directly inside .claude/settings.json so it inherits the personal visibility rules of local project settings", "correct": false, "explanation": "MCP servers are configured through mcpServers entries in .mcp.json or ~/.claude.json, not through the general settings.json permission/settings file."}, {"letter": "D", "text": "Add the server with local scope so the entry is written to .mcp.json but excluded from git tracking via a .gitignore rule", "correct": false, "explanation": "Local-scoped servers are stored in ~/.claude.json under that specific project's path, not in .mcp.json, and they only load in the project where they were added rather than across all projects."}], "correct": "B", "select": 1, "group": "J"}, {"id": "f2-068", "domain": 2, "task_id": "2.5", "objective": "Select and apply built-in tools (Read, Write, Edit, Bash, Grep, Glob) effectively", "situation": "Claude Code is asked to rename an environment variable from API_TIMEOUT_MS to REQUEST_TIMEOUT_MS everywhere it is referenced across a codebase of several hundred files, with each occurrence sitting in different surrounding code.", "question": "Which approach best discovers the full scope of the change before applying it?", "options": [{"letter": "A", "text": "Run Glob with the pattern **/API_TIMEOUT_MS to locate files whose names contain the variable, then edit only those matching files", "correct": false, "explanation": "Glob matches file names, not file contents; a source file referencing the environment variable in its code would not be named after that variable, so this would find few or no relevant files."}, {"letter": "B", "text": "Run Write on the project's environment configuration file first, then rely on Claude to infer other affected files from that one change afterward", "correct": false, "explanation": "Writing one file first does not reveal the other files that also reference the variable, and Claude cannot reliably infer scope from a single unrelated edit rather than actually searching the codebase."}, {"letter": "C", "text": "Run Grep with output mode content and a glob scope to list every file and line referencing API_TIMEOUT_MS, then review that list before editing", "correct": true, "explanation": "Grep in content mode returns the matching file and line for every occurrence of the pattern, giving a complete, reviewable inventory of every reference to the variable before any files are modified, scoped further with a glob parameter if needed."}, {"letter": "D", "text": "Run Bash to open every file in an interactive editor, since a variable rename of this kind must be reviewed visually rather than located programmatically", "correct": false, "explanation": "Bash has no built-in mechanism to open files in an interactive editor within this workflow, and manually inspecting hundreds of files is far less reliable than a programmatic content search."}], "correct": "C", "select": 1, "group": "F"}, {"id": "f2-069", "domain": 2, "task_id": "2.4", "objective": "Integrate MCP servers into Claude Code and agent workflows", "situation": "A team member tries to add a custom MCP server named computer-use to give the agent a specialized screenshot tool.", "question": "What should the architect expect to happen?", "options": [{"letter": "A", "text": "Claude Code loads the custom server normally and simply hides the built-in server named computer-use for the rest of that session", "correct": false, "explanation": "Claude Code does not silently substitute a custom server for a built-in one under a reserved name; it skips the conflicting custom entry and warns instead."}, {"letter": "B", "text": "Claude Code merges the tools from the custom server directly into the built-in server's own tool set under the shared reserved name", "correct": false, "explanation": "There is no field-level merging of a custom server's tools into a built-in server; a reserved-name conflict simply causes the custom entry to be skipped."}, {"letter": "C", "text": "Claude Code loads the custom server but silently strips its screenshot tool since that capability is presumed reserved for the built-in server", "correct": false, "explanation": "Claude Code does not selectively strip individual tools from a reserved-name server; the entire conflicting server entry is skipped at load time."}, {"letter": "D", "text": "Claude Code rejects or skips the server because computer-use is a reserved built-in name, so the team member needs to pick a different name", "correct": true, "explanation": "Names like workspace, claude-in-chrome, computer-use, and the Claude Preview/Browser names are reserved for Claude Code's built-in servers; a configuration using one of these names is skipped at load time with a warning, and claude mcp add rejects it outright."}], "correct": "D", "select": 1, "group": "J"}, {"id": "f2-070", "domain": 2, "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "situation": "A tool intended only for retrieving publicly available stock prices is described as \"Gets financial data.\" Occasionally, the model calls it hoping to retrieve a user's private account balance, which it cannot do.", "question": "What description change best sets the correct boundary?", "options": [{"letter": "A", "text": "Remove all wording from the description and rely solely on the tool's name to communicate its scope to the model.", "correct": false, "explanation": "Removing description text eliminates the primary signal the model uses for selection, leaving the scope boundary even less clear than before."}, {"letter": "B", "text": "State explicitly that the tool returns public market data only and cannot access private account or user-specific information.", "correct": true, "explanation": "Explicitly stating the public-only scope and the private-data exclusion gives the model the boundary information needed to avoid attempting unsupported requests."}, {"letter": "C", "text": "Rename the tool to get_financial_data_v2 so the model recognizes it as an updated, more capable version.", "correct": false, "explanation": "A version-number rename implies expanded capability rather than clarifying scope, which could worsen the misunderstanding about what the tool can access."}, {"letter": "D", "text": "Add a private boolean parameter to the input schema without altering any of the existing description text.", "correct": false, "explanation": "Adding a parameter without updating the description leaves the model unaware of the boundary until after it has already attempted the wrong call."}], "correct": "B", "select": 1, "group": "J"}, {"id": "f2-071", "domain": 2, "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "situation": "A team names two tools process_data and process_information, and both share the description \"Processes data provided by the user.\" The model routes roughly half of matching tasks to the wrong tool.", "question": "What is the most likely root cause?", "options": [{"letter": "A", "text": "The model's context window is too small to hold both full tool definitions during a single request's generation pass, leading to ambiguous routing.", "correct": false, "explanation": "Modern context windows are large enough to accommodate many tool definitions; two short tool descriptions would not approach the limit, so context window size is not the cause of the ambiguous routing."}, {"letter": "B", "text": "The near-duplicate descriptions give the model no differentiating signal, since descriptions are its main basis for choosing similar tools.", "correct": true, "explanation": "Because the descriptions are nearly identical, the model cannot reliably distinguish between the two tools; descriptions are the primary signal used for tool selection when tools are otherwise similar."}, {"letter": "C", "text": "The user's request phrasing was too vague to reliably distinguish between the tools, so the fix belongs in the user's prompt rather than the tool definitions.", "correct": false, "explanation": "The problem is that both tools have identical descriptions, so the model cannot differentiate regardless of user phrasing; the root cause is the tool definitions, not the user's prompt. Fixing the descriptions is the correct solution."}, {"letter": "D", "text": "The tools were declared in the wrong order in the tools array, and reordering them so that the intended tool is listed first will correct the routing behavior.", "correct": false, "explanation": "The order of tools in the array does not influence the model's selection; the model bases its decision on the content of tool descriptions, not their position. Reordering will not resolve the ambiguous routing."}], "correct": "B", "select": 1, "group": "J"}, {"id": "f2-072", "domain": 2, "task_id": "2.4", "objective": "Integrate MCP servers into Claude Code and agent workflows", "situation": "A custom MCP server dynamically adds a new tool partway through a long-running session, based on state changes on its own backend. The server sends the appropriate MCP notification for this.", "question": "What should the architect expect Claude Code to do, without any manual reconnect?", "options": [{"letter": "A", "text": "It will automatically refresh the tools from that server after receiving its list_changed notification, making the new tool usable without a disconnect.", "correct": true, "explanation": "Claude Code supports the MCP list_changed notification, which signals that a server's tool list has been updated. Upon receiving this notification, Claude automatically refreshes its internal tool list for that server, making the new tool immediately usable without any manual intervention or session restart."}, {"letter": "B", "text": "Ignore the list_changed notification and continue with the initial tool set, as tools are only loaded at session start and no dynamic refresh is supported.", "correct": false, "explanation": "Tool lists are not fixed for the entire session; the list_changed notification is specifically designed to allow dynamic updates to a server's capabilities mid-session. Claude Code handles this notification and updates its tool set accordingly, so the initial tool set is not the final one."}, {"letter": "C", "text": "Disconnect from the server and silently reconnect in the background, discarding any in-flight tool calls, then rely on the reconnect to pick up the new tool list.", "correct": false, "explanation": "Handling a list_changed notification does not require disconnecting from the MCP server or discarding in-flight tool calls. The tool list is refreshed in-place, preserving the existing connection and any ongoing tool executions."}, {"letter": "D", "text": "Require the user to run /mcp and manually select \"Refresh tools\" before the newly added tool becomes available, since the tool list updates only on manual refresh.", "correct": false, "explanation": "The refresh of the tool list is fully automatic when the list_changed notification is received; no manual action, such as running /mcp or selecting 'Refresh tools', is necessary for the update to take effect."}], "correct": "A", "select": 1, "group": "J"}, {"id": "f2-073", "domain": 2, "task_id": "2.4", "objective": "Integrate MCP servers into Claude Code and agent workflows", "situation": "A project-scoped stdio server's .mcp.json entry sets \"args\": [\"--root\", \"${CLAUDE_PROJECT_DIR}\"] with no fallback value.", "question": "What happens when a teammate runs Claude Code from a shell where this variable happens not to be set in their own environment?", "options": [{"letter": "A", "text": "The expansion falls back automatically to the user's home directory, since that is treated as the implicit default whenever no fallback is written in .mcp.json", "correct": false, "explanation": "Claude Code documents no implicit home-directory default for any .mcp.json variable reference. Without an explicit fallback such as ${CLAUDE_PROJECT_DIR:-.}, an unset reference resolves to the literal unexpanded text plus a missing-variable warning, never to the home directory."}, {"letter": "B", "text": "The expansion fails outright, because ${CLAUDE_PROJECT_DIR} can only ever be read from the invoking shell's own environment, never from a value Claude Code injects itself", "correct": false, "explanation": "A missing environment variable in .mcp.json does not stop the server from starting or force a hard failure. Claude Code's documentation says the config still loads, that it reports a missing-variable warning in claude mcp list, and that it uses the unexpanded text as-is. The reference is also not confined to the invoking shell's environment alone, since Claude Code does inject the value, just too late for this particular expansion."}, {"letter": "C", "text": "It resolves correctly, since Claude Code injects CLAUDE_PROJECT_DIR into the spawned server's own environment, though a default like ${CLAUDE_PROJECT_DIR:-.} remains the safer practice", "correct": false, "explanation": "Claude Code does inject CLAUDE_PROJECT_DIR into the spawned server's own environment, but expansion inside .mcp.json's command and args happens earlier, while Claude Code is still parsing the file, before that injection exists. Anthropic's MCP documentation states the variable is set in the server's environment, not in Claude Code's own environment, so a project-scoped entry with no fallback does not resolve correctly on its own."}, {"letter": "D", "text": "The expansion leaves the literal ${CLAUDE_PROJECT_DIR} text and logs a missing-variable warning, since the value exists only in the server's environment, not in Claude Code's own", "correct": true, "explanation": "Anthropic's Claude Code MCP documentation states this variable is set in the server's environment, not in Claude Code's own environment, so referencing it in a project-scoped .mcp.json entry, or a local- or user-scoped server in ~/.claude.json, needs a default such as ${CLAUDE_PROJECT_DIR:-.}. Without one, the reference is treated as unset: the config still loads, a missing-variable warning is reported, and the unexpanded text is used as-is. Plugin-provided configurations are the documented exception, since they substitute the value directly."}], "correct": "D", "select": 1, "group": "J"}, {"id": "f2-074", "domain": 2, "task_id": "2.2", "objective": "Implement structured error responses for MCP tools", "situation": "A submit_refund MCP tool rejects a request because the customer's order is past the 30-day return window. The engineer designing the error response wants the agent to explain the rejection to the customer in plain language and never attempt this exact call again.", "question": "Which response body best achieves this?", "options": [{"letter": "A", "text": "A text block plus structured content with errorCategory: \"permission\" and a description asking the customer to contact their account administrator.", "correct": false, "explanation": "The permission error category signals an authorization or access‑control issue, not a business‑policy rejection. The refund was denied because the order exceeded the return window, not because of insufficient permissions. The agent would be misled into guiding the customer to contact an administrator instead of explaining the policy. The correct category is business, with a description that clearly states the 30‑day window has passed."}, {"letter": "B", "text": "A bare text block reading \"Refund denied\" with no structured metadata.", "correct": false, "explanation": "Without structured metadata like errorCategory, the agent cannot distinguish a business-rule violation from a transient failure, nor can it reliably decide whether to retry. Best practices require explicit categorization so the agent makes appropriate recovery decisions and clearly explains the reason to the customer, as documented in Anthropic's certification guide."}, {"letter": "C", "text": "A text block explaining the rejection, along with structured metadata including errorCategory: \"business\" to indicate a business rule violation and a description stating the order fell outside the 30-day return window.", "correct": true, "explanation": "Per Anthropic's Claude Certification Guide, business rule violations like an order being past the return window should be communicated with a structured error response. The recommended approach includes errorCategory: \"business\", which tells the agent that the error is permanent and the request must not be retried. The plain-text description can be relayed to the customer. Although MCP tools do not provide a standard isRetryable field, the errorCategory field serves the same purpose: a business error implies that retries will never succeed, guiding the agent toward an alternative workflow."}, {"letter": "D", "text": "A text block plus structured content with errorCategory: \"transient\" and a description telling the agent to wait 30 seconds and resubmit.", "correct": false, "explanation": "This response misclassifies a permanent business-rule rejection as a transient error. Transient errors suggest temporary issues (e.g., network timeouts) that may resolve on retry. Here, the 30‑day return window has definitively expired, so retrying will always fail. The correct approach uses errorCategory: \"business\" to signal non‑retryable failure."}], "correct": "C", "select": 1, "group": "J"}, {"id": "f2-075", "domain": 2, "task_id": "2.3", "objective": "Distribute tools appropriately across agents and configure tool choice", "situation": "An architect is evaluating a customer-support triage agent whose only job is to read an incoming ticket and route it to billing, technical, or account-management queues. The agent currently has route_ticket, plus process_refund, reset_password, and close_account tools \"in case they're needed later.\" Logs show the agent occasionally calls process_refund directly instead of routing the ticket to billing.", "question": "What is the most appropriate fix?", "options": [{"letter": "A", "text": "Leave all four tools in place but add a system prompt line that forces refunds to go through billing first, so the triage agent only calls process_refund after routing.", "correct": false, "explanation": "Relying on a system prompt to enforce routing before refunds is a weak control compared to removing out-of-role tools. The logs already show that the agent bypasses such soft guidance in practice."}, {"letter": "B", "text": "Add a human approval step before process_refund executes so the triage agent can still invoke it but only after manual confirmation, preventing direct automated refunds.", "correct": false, "explanation": "Adding a human approval step still leaves process_refund available to an agent whose role is routing only; this does not fix the underlying scoping problem, merely adding a mitigation that could be bypassed or mishandled."}, {"letter": "C", "text": "Remove process_refund, reset_password, and close_account from the triage agent so it only has route_ticket, matching its single core routing function.", "correct": true, "explanation": "Removing process_refund, reset_password, and close_account leaves only route_ticket, which enforces scoped access matching the triage agent’s single core routing function. This directly addresses the root cause (the agent having out-of-role tools that were misused) rather than adding mitigations."}, {"letter": "D", "text": "Rename process_refund to billing_queue_helper so the triage agent treats it as a routing alias rather than a direct action, avoiding unintended refund calls.", "correct": false, "explanation": "Renaming process_refund to billing_queue_helper only changes its surface label, not its functionality or availability; the agent can still invoke the refund logic directly regardless of naming."}], "correct": "C", "select": 1, "group": "B"}, {"id": "f2-076", "domain": 2, "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "situation": "An MCP server exposes get_user_profile and get_user_permissions with identical one-line descriptions: \"Returns user account information.\" Client agents integrating this server report frequent misrouting when asked for either a display name or an access level.", "question": "As the MCP server author, what is the most effective fix?", "options": [{"letter": "A", "text": "Add authentication scopes to the MCP server configuration, since permission errors are the actual cause of the reported misrouting.", "correct": false, "explanation": "The reported issue is agents calling the wrong tool, not being denied access, so authentication scopes do not address the described symptom."}, {"letter": "B", "text": "Update each tool's description to name the specific fields it returns, such as display name versus roles and access scopes.", "correct": true, "explanation": "Naming the specific fields each tool returns gives connecting agents the information needed to distinguish profile data from permission data without inspecting the schema."}, {"letter": "C", "text": "Combine both endpoints into a single MCP resource instead of a tool, since resources are inherently immune to selection ambiguity.", "correct": false, "explanation": "Converting to a resource does not inherently resolve ambiguity; resources still need clear, distinguishing descriptions of what they expose."}, {"letter": "D", "text": "Instruct every client application connecting to the server to hardcode which tool to call for each intent, bypassing tool selection.", "correct": false, "explanation": "Hardcoding tool choice in every client application defeats the purpose of exposing a general-purpose MCP tool interface and does not scale across integrators."}], "correct": "B", "select": 1, "group": "J"}, {"id": "f2-077", "domain": 2, "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "situation": "A tool named analyze_content was originally built to summarize any pasted content, including emails and internal notes, but its description was never updated after a web-specific successor tool was introduced. The model now sometimes calls the general tool for web-only tasks.", "question": "What is the recommended remediation?", "options": [{"letter": "A", "text": "Rename the general tool to reflect its remaining scope, and update its description to exclude the case the newer tool now handles.", "correct": true, "explanation": "Updating the general tool's name and description to reflect its narrower remaining scope removes the overlap with the newer, web-specific tool."}, {"letter": "B", "text": "Increase the priority weight of the newer tool in the backend routing configuration without touching either tool's description.", "correct": false, "explanation": "A backend priority weight influences routing outside of the model's own reasoning and does not resolve the description staleness that caused the confusion."}, {"letter": "C", "text": "Leave both tools with their original names, but ask users to specify which tool they want by internal ID every time.", "correct": false, "explanation": "Requiring users to reference internal tool IDs defeats the purpose of natural-language tool selection and does not address the stale description."}, {"letter": "D", "text": "Delete the description text from both tools so the model relies entirely on tool names for disambiguation.", "correct": false, "explanation": "Removing description text eliminates useful information entirely rather than correcting the specific stale claim about scope."}], "correct": "A", "select": 1, "group": "J"}, {"id": "g16", "domain": 3, "scenario": "Claude Code for Continuous Integration", "situation": "Your CI pipeline runs the Claude Code CLI (in `--print` mode) using CLAUDE.md to provide project context for code review, and developers generally find the reviews substantive. However, they report that integrating findings into the workflow is difficult—Claude outputs narrative paragraphs that must be manually copied into PR comments. The team wants to automatically post each finding as a separate inline PR comment at the relevant place in code, which requires structured data with file path, line number, severity level, and suggested fix. Which approach is most effective?", "question": "Which approach is most effective?", "options": [{"letter": "A", "text": "Add an “Output Format for Review” section to CLAUDE.md with examples of structured findings so Claude learns the expected format from project context.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Use the CLI flags `--output-format json` and `--json-schema` to enforce structured findings, then parse the output to post inline comments via the GitHub API.", "correct": true, "explanation": "Using `--output-format json` with `--json-schema` enforces structured output at the CLI level, guaranteeing well-formed JSON with the required fields (file path, line number, severity, suggested fix) that can be reliably parsed and posted as inline PR comments via the GitHub API. It leverages built-in CLI capabilities designed specifically for structured output."}, {"letter": "C", "text": "Include explicit formatting instructions in the review prompt requiring each finding to follow a parseable template like `[FILE:path] [LINE:n] [SEVERITY:level] ...`.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Keep narrative review format but add a summarization step that uses Claude to generate a structured JSON summary of findings.", "correct": false, "explanation": ""}], "correct": "B", "task_id": "3.6", "objective": "Integrate Claude Code into CI/CD pipelines", "group": "D"}, {"id": "g26", "domain": 3, "scenario": "Claude Code for Continuous Integration", "situation": "Your pipeline script runs `claude \"Analyze this pull request for security issues\"`, but the job hangs indefinitely. Logs show Claude Code is waiting for interactive input. What is the correct approach to run Claude Code in an automated pipeline?", "question": "What is the correct approach?", "options": [{"letter": "A", "text": "Add a `--batch` flag: `claude --batch \"Analyze this pull request for security issues\"`.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Add the `-p` flag: `claude -p \"Analyze this pull request for security issues\"`.", "correct": true, "explanation": "The `-p` (or `--print`) flag is the documented way to run Claude Code non-interactively. It processes the prompt, prints the result to stdout, and exits without waiting for user input—ideal for CI/CD pipelines."}, {"letter": "C", "text": "Redirect stdin from `/dev/null`: `claude \"Analyze this pull request for security issues\" < /dev/null`.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Set the environment variable `CLAUDE_HEADLESS=true` before running the command.", "correct": false, "explanation": ""}], "correct": "B", "task_id": "3.6", "objective": "Integrate Claude Code into CI/CD pipelines", "group": "D"}, {"id": "g32", "domain": 3, "scenario": "Code Generation with Claude Code", "situation": "You need to add Slack as a new notification channel. The existing codebase has clear, established patterns for email, SMS, and push channels. However, Slack’s API offers fundamentally different integration approaches—incoming webhooks (simple, one-way), bot tokens (support delivery confirmation and programmatic control), or Slack Apps (two-way events, requires workspace approval). Your task says “add Slack support” without specifying integration method or requiring advanced features like delivery tracking.", "question": "How should you approach this task?", "options": [{"letter": "A", "text": "Start in direct execution mode using incoming webhooks to match the existing one-way notification pattern.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Switch to planning mode to explore integration options and architectural implications, then present a recommendation before implementation.", "correct": true, "explanation": "Slack integration has multiple valid approaches with significantly different architectural implications, and requirements are ambiguous. Planning mode lets you evaluate trade-offs among webhooks, bot tokens, and Slack Apps and align on an approach before implementation."}, {"letter": "C", "text": "Start in direct execution mode by scaffolding a Slack channel class using existing patterns, deferring the integration method decision.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Start in direct execution mode using a bot-token approach to ensure delivery confirmation is possible.", "correct": false, "explanation": ""}], "correct": "B", "task_id": "3.4", "objective": "Determine when to use plan mode vs direct execution", "group": "C"}, {"id": "g33", "domain": 3, "scenario": "Code Generation with Claude Code", "situation": "Your CLAUDE.md file has grown to 400+ lines containing coding standards, testing conventions, a detailed PR review checklist, deployment instructions, and database migration procedures. You want Claude to always follow coding standards and testing conventions, but apply PR review, deploy, and migration guidance only when doing those tasks.", "question": "Which restructuring approach is most effective?", "options": [{"letter": "A", "text": "Move all guidance into separate Skills files organized by workflow type, leaving only a brief project description in CLAUDE.md.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Keep everything in CLAUDE.md but use `@import` syntax to organize into separately maintained files by category.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Split CLAUDE.md into files under `.claude/rules/` with path-bound glob patterns so each rule loads only for the relevant file types.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Keep universal standards in CLAUDE.md and create Skills for workflow-specific guidance (PR review, deploy, migrations) with trigger keywords.", "correct": true, "explanation": "CLAUDE.md content loads in every session, ensuring coding standards and testing conventions always apply, while Skills are invoked on demand when Claude detects trigger keywords—ideal for workflow-specific guidance like PR review, deployment, and migrations."}], "correct": "D", "task_id": "3.1", "objective": "Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization", "group": "D"}, {"id": "g34", "domain": 3, "scenario": "Code Generation with Claude Code", "situation": "You’re tasked with restructuring your team’s monolithic application into microservices. This impacts changes across dozens of files and requires decisions about service boundaries and module dependencies.", "question": "Which approach should you choose?", "options": [{"letter": "A", "text": "Switch to planning mode to explore the codebase, understand dependencies, and design the implementation approach before making changes.", "correct": true, "explanation": "Planning mode is the right strategy for complex architectural restructuring like splitting a monolith: it allows safe exploration and informed decisions about boundaries before committing to potentially expensive changes across many files."}, {"letter": "B", "text": "Start in direct execution mode and switch to planning only after encountering unexpected complexity during implementation.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Start in direct execution mode and make incremental changes, letting implementation reveal natural service boundaries.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Use direct execution with detailed upfront instructions that specify each service structure.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "3.4", "objective": "Determine when to use plan mode vs direct execution", "group": "C"}, {"id": "g35", "domain": 3, "scenario": "Code Generation with Claude Code", "situation": "Your team created a `/analyze-codebase` skill that performs deep code analysis—dependency scanning, test coverage counts, and code quality metrics. After running the command, team members report Claude becomes less responsive in the session and loses the context of the original task.", "question": "How do you most effectively fix this while keeping full analysis capabilities?", "options": [{"letter": "A", "text": "Add `context: fork` in the skill frontmatter to run the analysis in an isolated subagent context.", "correct": true, "explanation": "`context: fork` runs the analysis in an isolated subagent context so the large output does not pollute the main session’s context window and Claude does not lose track of the original task. It preserves full analysis capability while keeping the main session responsive."}, {"letter": "B", "text": "Add `model: haiku` in frontmatter to use a faster, cheaper model for analysis.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Split the skill into three smaller skills, each producing less output.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Add instructions to the skill to compress all results into a short summary before displaying them.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "3.2", "objective": "Create and configure custom slash commands and skills", "group": "D"}, {"id": "g36", "domain": 3, "scenario": "Code Generation with Claude Code", "situation": "Your team uses a `/commit` skill in `.claude/skills/commit/SKILL.md`. A developer wants to customize it for their personal workflow (different commit message format, extra checks) without affecting teammates.", "question": "What do you recommend?", "options": [{"letter": "A", "text": "Create a personal version under `~/.claude/skills/` with a different name, e.g., `/my-commit`.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Add conditional logic based on username in the project skill frontmatter.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Create a personal version at `~/.claude/skills/commit/SKILL.md` with the same name.", "correct": true, "explanation": "Personal skills take precedence over project skills with the same name. A personal skill at `~/.claude/skills/commit/SKILL.md` will override the team’s project skill, allowing the developer to customize their workflow while maintaining the familiar `/commit` command name for their personal use. This approach is better than option A because it preserves the original command name, improving the developer’s workflow without affecting teammates."}, {"letter": "D", "text": "Set `override: true` in the personal skill frontmatter to prioritize it over the project version.", "correct": false, "explanation": ""}], "correct": "C", "task_id": "3.2", "objective": "Create and configure custom slash commands and skills", "group": "D"}, {"id": "g37", "domain": 3, "scenario": "Code Generation with Claude Code", "situation": "Your team has used Claude Code for months. Recently, three developers report Claude follows the guidance “always include comprehensive error handling,” but a fourth developer who just joined says Claude does not follow it. All four work in the same repo and have up-to-date code.", "question": "What is the most likely cause and fix?", "options": [{"letter": "A", "text": "The guidance lives in the original developers’ user-level `~/.claude/CLAUDE.md` files, not in the project `.claude/CLAUDE.md`. Move the instruction to the project-level file so all team members receive it.", "correct": true, "explanation": "If the guidance was added only to the original developers’ user-level configs and not to the project-level `.claude/CLAUDE.md`, new team members won’t receive it. Moving it to the project-level configuration ensures all current and future team members automatically get the guidance."}, {"letter": "B", "text": "The new developer’s `~/.claude/CLAUDE.md` contains conflicting instructions overriding project settings; they should delete the conflicting section.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Claude Code learns per-user preferences over time; the new developer must repeat the requirement until Claude “remembers” it.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Claude Code caches CLAUDE.md after first read; original developers use cached versions. Everyone should clear the Claude Code cache.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "3.1", "objective": "Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization", "group": "D"}, {"id": "g38", "domain": 3, "scenario": "Code Generation with Claude Code", "situation": "You find that including 2–3 full endpoint implementation examples as context significantly improves consistency when generating new API endpoints. However, this context is useful only when creating new endpoints—not when debugging, reviewing code, or other work in the API directory.", "question": "Which configuration approach is most effective?", "options": [{"letter": "A", "text": "Add endpoint examples and pattern documentation to the project CLAUDE.md so they are always available.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Manually reference endpoint examples in every generation request by copying code into the prompt.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Configure path-specific rules in `.claude/rules/api/` that include endpoint examples and activate when working in the API directory.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Create a skill that references the endpoint examples and contains pattern-following instructions, invoked on demand via a slash command.", "correct": true, "explanation": "A skill invoked on demand loads the example context only when generating new endpoints, not during unrelated tasks like debugging or review. This keeps the main context clean while preserving high-quality generation when needed."}], "correct": "D", "task_id": "3.1", "objective": "Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization", "group": "D"}, {"id": "g39", "domain": 3, "scenario": "Code Generation with Claude Code", "situation": "Your team created a `/migration` skill that generates database migration files. It takes the migration name via `$ARGUMENTS`. In production you observe three issues: (1) developers often run the skill without arguments, causing poorly named files, (2) the skill sometimes uses database schema details from unrelated prior conversations, and (3) a developer accidentally ran destructive test cleanup when the skill had broad tool access.", "question": "Which configuration approach fixes all three problems?", "options": [{"letter": "A", "text": "Use positional parameters `$1` and `$2` instead of `$ARGUMENTS` to enforce specific inputs, include explicit schema file references via `@` syntax for context control, and add a frontmatter description warning about destructive operations.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Add `argument-hint` in frontmatter to request required parameters, use `context: fork` to isolate execution, and restrict `allowed-tools` to file-write operations.", "correct": true, "explanation": "This uses three separate configuration features to address each problem: `argument-hint` improves argument entry and reduces missing arguments, `context: fork` prevents context leakage from prior conversations, and `allowed-tools` constrains the skill to safe file-writing operations, preventing destructive actions."}, {"letter": "C", "text": "Split into `/migration-create` and `/migration-apply` skills, add validation instructions to request migration name if missing, and use different `allowed-tools` scopes for each.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Add validation instructions in the skill SKILL.md to ensure `$ARGUMENTS` is a valid name, add prompts to ignore prior conversation context, and list prohibited operations to avoid.", "correct": false, "explanation": ""}], "correct": "B", "task_id": "3.2", "objective": "Create and configure custom slash commands and skills", "group": "D"}, {"id": "g40", "domain": 3, "scenario": "Code Generation with Claude Code", "situation": "Your codebase contains areas with different coding conventions: React components use functional style with hooks, API handlers use async/await with specific error handling, and database models follow the repository pattern. Test files are distributed across the codebase next to the code under test (e.g., `Button.test.tsx` next to `Button.tsx`), and you want all tests to follow the same conventions regardless of location.", "question": "What is the most supported way to ensure Claude automatically applies the correct conventions when generating code?", "options": [{"letter": "A", "text": "Put all conventions in the root CLAUDE.md under headings for each area and rely on Claude to infer which section applies.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Create skills in `.claude/skills/` for each code type, embedding conventions in each SKILL.md.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Place a separate CLAUDE.md file in each subdirectory containing conventions for that area.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Create rule files under `.claude/rules/` with YAML frontmatter specifying glob patterns to conditionally apply conventions based on file paths.", "correct": true, "explanation": "`.claude/rules/` files with YAML frontmatter and glob patterns (e.g., `**/*.test.tsx`, `src/api/**/*.ts`) enable deterministic, path-based convention application regardless of directory structure. This is the most supported approach for cross-cutting patterns like distributed test files."}], "correct": "D", "task_id": "3.1", "objective": "Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization", "group": "D"}, {"id": "g41", "domain": 3, "scenario": "Code Generation with Claude Code", "situation": "You want to create a custom slash command `/review` that runs your team’s standard code review checklist. It should be available to every developer when they clone or update the repository.", "question": "Where should you create the command file?", "options": [{"letter": "A", "text": "In `~/.claude/commands/` in each developer’s home directory.", "correct": false, "explanation": ""}, {"letter": "B", "text": "In the project repository under `.claude/commands/`.", "correct": true, "explanation": "Putting custom slash commands under `.claude/commands/` inside the project repository ensures they are version-controlled and automatically available to every developer who clones or updates the repo. This is the intended location for project-level custom commands in Claude Code."}, {"letter": "C", "text": "In `.claude/config.json` as an array of commands.", "correct": false, "explanation": ""}, {"letter": "D", "text": "In the root project CLAUDE.md.", "correct": false, "explanation": ""}], "correct": "B", "task_id": "3.2", "objective": "Create and configure custom slash commands and skills", "group": "D"}, {"id": "g42", "domain": 3, "scenario": "Code Generation with Claude Code", "situation": "Your team’s CLAUDE.md grew beyond 500 lines mixing TypeScript conventions, testing guidance, API patterns, and deployment procedures. Developers find it hard to locate and update the right sections.", "question": "What approach does Claude Code support to organize project-level instructions into focused topical modules?", "options": [{"letter": "A", "text": "Define a `.claude/config.yaml` mapping file patterns to specific sections inside CLAUDE.md.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Create separate Markdown files in `.claude/rules/`, each covering one topic (e.g., `testing.md`, `api-conventions.md`).", "correct": true, "explanation": "Claude Code supports a `.claude/rules/` directory where you can create separate Markdown files for topical guidance (e.g., `testing.md`, `api-conventions.md`), allowing teams to organize large instruction sets into focused, maintainable modules."}, {"letter": "C", "text": "Split instructions into README.md files in relevant subdirectories that Claude automatically loads as instructions.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Create multiple files named CLAUDE.md at different levels of the directory tree, each overriding parent instructions.", "correct": false, "explanation": ""}], "correct": "B", "task_id": "3.1", "objective": "Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization", "group": "D"}, {"id": "g43", "domain": 3, "scenario": "Code Generation with Claude Code", "situation": "You create a custom skill `/explore-alternatives` that your team uses to brainstorm and evaluate implementation approaches before choosing one. Developers report that after running the skill, subsequent Claude responses are influenced by the alternatives discussion—sometimes referencing rejected approaches or retaining exploration context that interferes with actual implementation.", "question": "How should you most effectively configure this skill?", "options": [{"letter": "A", "text": "Use the `!` prefix in the skill to run exploration logic as a bash subprocess.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Add `context: fork` in the skill frontmatter.", "correct": true, "explanation": "`context: fork` runs the skill in an isolated subagent context so exploration discussions do not pollute the main conversation history. This prevents rejected approaches and brainstorming context from influencing subsequent implementation work."}, {"letter": "C", "text": "Split into two skills—`/explore-start` and `/explore-end`—to mark boundaries when exploration context should be discarded.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Create the skill in `~/.claude/skills/` instead of `.claude/skills/`.", "correct": false, "explanation": ""}], "correct": "B", "task_id": "3.2", "objective": "Create and configure custom slash commands and skills", "group": "D"}, {"id": "g44", "domain": 3, "scenario": "Code Generation with Claude Code", "situation": "Your team wants to add a GitHub MCP server for searching PRs and checking CI status via Claude Code. Each of six developers has their own personal GitHub access token. You want consistent tooling across the team without committing credentials to version control.", "question": "Which configuration approach is most effective?", "options": [{"letter": "A", "text": "Have each developer add the server in user scope via `claude mcp add --scope user`.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Create an MCP server wrapper that reads tokens from a `.env` file and proxies GitHub API calls, then add the wrapper to the project `.mcp.json`.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Add the server to the project `.mcp.json` using environment variable substitution (`${GITHUB_TOKEN}`) for auth and document the required environment variable in the project README.", "correct": true, "explanation": "A project `.mcp.json` with environment variable substitution is idiomatic: it provides a single version-controlled source of truth for MCP configuration while letting each developer supply credentials via environment variables. Documenting the variable makes onboarding easy without committing secrets."}, {"letter": "D", "text": "Configure the server in project scope with a placeholder token, then tell developers to override it in their local config.", "correct": false, "explanation": ""}], "correct": "C", "task_id": "2.4", "objective": "Integrate MCP servers into Claude Code and agent workflows", "group": "J"}, {"id": "m19", "domain": 3, "scenario": null, "situation": "An engineer used `Claude Code` yesterday to investigate authentication flows in a legacy monolith, building up significant context over a 2-hour session. Today she wants to continue that specific investigation. She's worked on three other codebases since then and knows the session was named \"auth-deep-dive\".", "question": "How should she resume?", "options": [{"letter": "A", "text": "Start fresh and re-read the same files", "correct": false, "explanation": "Wastes the two hours of accumulated context from yesterday."}, {"letter": "B", "text": "Use `--session-id` with the UUID from yesterday's session transcript file", "correct": false, "explanation": "Possible but cumbersome — she'd have to hunt for the UUID when she already has the session name."}, {"letter": "C", "text": "Use `--continue` to pick up where the most recent conversation left off", "correct": false, "explanation": "`--continue` resumes the most recent session in this repo. She's worked on three other codebases since then, so the most recent is not auth-deep-dive."}, {"letter": "D", "text": "Use `--resume` auth-deep-dive to load that specific session by name", "correct": true, "explanation": "Correct. `--resume` with the session name is designed for exactly this: pick a specific prior session out of many, by the name you gave it."}], "correct": "D", "task_id": "1.7", "objective": "Manage session state, resumption, and forking", "group": "A"}, {"id": "f3-001", "domain": 3, "task_id": "3.3", "objective": "Apply path-specific rules for conditional convention loading", "situation": "A frontend team's component tests use the .test.tsx extension but live in many different locations: some are under src/ (e.g., src/components/__tests__/, src/pages/tests/, and files colocated next to source components), and others are outside src/ (e.g., a root-level tests/ folder).", "question": "Which paths pattern in a rules file frontmatter ensures the rule applies to every one of these test files regardless of which folder it lives in?", "options": [{"letter": "A", "text": "src/components/*.test.tsx", "correct": false, "explanation": "This pattern matches only .test.tsx files directly inside src/components/ without recursion. It would not match files in nested folders like src/components/__tests__/, nor any files in other locations such as src/pages/tests/, colocated files, or files outside src/. A recursive glob like **/*.test.tsx is required to cover all possible depths and locations."}, {"letter": "B", "text": "**/*.test.tsx", "correct": true, "explanation": "The ** glob segment matches any number of directories, so this pattern discovers .test.tsx files at any depth, including those under src/ and those outside it, such as a root-level tests/ folder. This is a standard discovery pattern: Vitest's default test file patterns explicitly include **/*.test.{ts,js,mjs,cjs,tsx,jsx}. Because it is not anchored to src/, it applies regardless of the folder location."}, {"letter": "C", "text": "**/__tests__/*.tsx", "correct": false, "explanation": "This pattern matches any .tsx file directly inside a folder named __tests__, but it does not enforce the .test.tsx naming convention and would also match non-test .tsx files in such folders. Moreover, it would miss test files in src/pages/tests/ (folder not named __tests__) and colocated test files that are not inside a __tests__ directory."}, {"letter": "D", "text": "src/**/*.test.tsx", "correct": false, "explanation": "This pattern only matches .test.tsx files located under the src/ directory. It would miss test files outside src/, such as the root-level tests/ folder mentioned in the scenario. Since the requirement is to cover all test files regardless of location, a pattern that restricts the search to src/ is insufficient."}], "correct": "B", "select": 1, "group": "D"}, {"id": "f3-002", "domain": 3, "task_id": "3.1", "objective": "Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization", "situation": "An engineer launches Claude Code from the services/billing/ directory inside a larger repository. The repository contains three CLAUDE.md files: one at the repository root, one at services/billing/CLAUDE.md, and one at services/billing/reports/CLAUDE.md. The engineer has not yet read or edited any files in the services/billing/reports/ subdirectory.", "question": "At the moment the session starts, which files are loaded into context?", "options": [{"letter": "A", "text": "Only services/billing/CLAUDE.md loads, since directory-level files never combine with ancestor files unless explicitly imported first.", "correct": false, "explanation": "Official documentation states that when Claude Code operates within a subdirectory, it loads both the root CLAUDE.md and the CLAUDE.md present in that specific subdirectory. The root file always provides shared constraints, so it is included alongside the working directory's instructions."}, {"letter": "B", "text": "All three CLAUDE.md files load immediately, because Claude Code always preloads every CLAUDE.md found anywhere under the working directory.", "correct": false, "explanation": "Claude Code does not pre‑scan and load every CLAUDE.md in the tree. Subdirectory files follow an on‑demand loading pattern—only the root and the immediate working directory CLAUDE.md are loaded at launch. This is an intentional design to avoid bloating the initial context window."}, {"letter": "C", "text": "The root and services/billing/CLAUDE.md load at launch; reports/CLAUDE.md loads later, only when Claude reads a file in that subdirectory.", "correct": true, "explanation": "Anthropic's hierarchical loading behavior is: the root CLAUDE.md is always loaded at session start, and any CLAUDE.md in the immediate working directory is loaded as well. Deeper subdirectory files like reports/CLAUDE.md are loaded on demand, only when Claude interacts with files inside their directory. This progressive disclosure minimizes context size."}, {"letter": "D", "text": "Only the root CLAUDE.md loads at launch; services/billing/ and the reports subdirectory file both wait until Claude reads a file there.", "correct": false, "explanation": "While the root CLAUDE.md does load at launch, the CLAUDE.md inside the working directory services/billing/ is also loaded immediately. Only deeper subdirectories, like reports/, are deferred until their files are accessed."}], "correct": "C", "select": 1, "group": "D"}, {"id": "f3-003", "domain": 3, "task_id": "3.3", "objective": "Apply path-specific rules for conditional convention loading", "situation": "In a monorepo, Team A's rules live under packages/team-a/.claude/rules/ and Team B's rules live under packages/team-b/.claude/rules/. An engineer on Team A, working exclusively in packages/team-a/, wants to stop Team B's rules from entering context by changing only their own local Claude Code configuration — without editing any file inside packages/team-b/ — while still letting Team A's own path-scoped rules load conditionally as normal.", "question": "What should the engineer configure?", "options": [{"letter": "A", "text": "Ensure that all rule files in packages/team-b/.claude/rules/ include a paths frontmatter key with glob patterns scoped to packages/team-b/. Path-scoped rules only load when Claude works on matching files, so Team B's rules will not be triggered while the engineer works exclusively in packages/team-a/.", "correct": false, "explanation": "Adding paths frontmatter to every rule in packages/team-b/.claude/rules/ would keep those rules from loading outside matching paths, but doing so requires editing files inside packages/team-b/, which the scenario rules out for this engineer. Claude Code's own guidance for excluding another team's instructions without touching their files is the claudeMdExcludes setting."}, {"letter": "B", "text": "Add a claudeMdExcludes entry in settings pointing to packages/team-b/.claude/rules/**, which skips Team B's rules files without affecting how Team A's path-scoped rules conditionally load based on their own paths frontmatter.", "correct": true, "explanation": "claudeMdExcludes lets an engineer skip specific files or directories — including a .claude/rules/ directory — from their own settings, without editing the excluded team's files. Claude Code's documentation gives this exact monorepo pattern: exclude a rules directory such as .../other-team/.claude/rules/** from .claude/settings.local.json, keeping the exclusion local to the engineer's own machine. Team A's own rules keep their paths frontmatter, so they still load conditionally as before."}, {"letter": "C", "text": "Delete the paths frontmatter from every rule in packages/team-a/.claude/rules/ so those rules load unconditionally at launch, which the engineer hopes also stops packages/team-b/.claude/rules/ files from ever being discovered.", "correct": false, "explanation": "Removing paths frontmatter makes Team A's own rules load unconditionally rather than conditionally, the opposite of the desired behavior, and it has no effect on whether Team B's .claude/rules/ files are discovered or loaded."}, {"letter": "D", "text": "Set autoMemoryEnabled to false in the engineer's local project settings, which the engineer expects stops Claude Code from discovering any .claude/rules/ directory located outside the current package.", "correct": false, "explanation": "autoMemoryEnabled controls whether Claude Code saves its own auto-generated memory notes under ~/.claude/projects/<project>/memory/; it has no effect on whether .claude/rules/ files are discovered or loaded. Excluding another team's rules without editing their files is done with the claudeMdExcludes setting instead."}], "correct": "B", "select": 1, "group": "D"}, {"id": "f3-004", "domain": 3, "task_id": "3.5", "objective": "Apply iterative refinement techniques for progressive improvement", "situation": "", "question": "An architect is refining a text-summarization endpoint and finds Claude's summaries alternate unpredictably between one-sentence summaries and multi-paragraph summaries depending on the wording of the prompt, even though the prompt states \"keep summaries concise.\" What is the most direct fix for this inconsistency?", "options": [{"letter": "A", "text": "Replace 'concise' with a detailed description such as 'write a summary that is short, tight, succinct, and to the point, avoiding any unnecessary words' to provide the model with explicit length guidance.", "correct": false, "explanation": "A pile of synonyms like \"short, tight, succinct\" is still ambiguous prose; it doesn't specify an explicit length or format. The model still lacks a concrete anchor, so inconsistency will persist."}, {"letter": "B", "text": "Remove the word 'concise' from the prompt and instead let the model determine the appropriate summary length for each input article, resulting in stable and predictable behavior across all requests.", "correct": false, "explanation": "Removing the length constraint altogether does not guarantee stable behavior; without any guidance, the model may still vary summary length unpredictably. It abandons the goal of concise summaries, opposite of the architect's intent."}, {"letter": "C", "text": "Instruct Claude to produce exactly one sentence summaries regardless of input length, using a system prompt like 'Your response must be exactly one sentence long' and a stop sequence to truncate any further text.", "correct": false, "explanation": "Mandating exactly one sentence for all inputs is overly rigid and may produce poor summaries for complex articles. It addresses inconsistency only by sacrificing quality and flexibility, which is not a desirable fix."}, {"letter": "D", "text": "Provide 2-3 sample input articles paired with example summaries at the exact target length and style, so \"concise\" is anchored to a concrete example rather than left open to interpretation.", "correct": true, "explanation": "Few-shot examples (2-3 pairs) demonstrate the desired length and style concretely, teaching the model what \"concise\" means in context. This anchors the instruction and resolves inconsistency more effectively than rephrasing or adding rigidity."}], "correct": "D", "select": 1, "group": "I"}, {"id": "f3-005", "domain": 3, "task_id": "3.5", "objective": "Apply iterative refinement techniques for progressive improvement", "situation": "An architect asks Claude Code to \"normalize phone numbers to a standard format\" across a customer database. After two rounds of correction, Claude is still inconsistently formatting extensions and international codes.", "question": "What should the architect do to communicate the transformation more effectively?", "options": [{"letter": "A", "text": "Split the normalization rules into separate sections for domestic numbers, international codes, and extensions, each described with exhaustive prose detail rather than examples.", "correct": false, "explanation": "Exhaustive prose detail, even if split into sections, does not remove the interpretive gaps that caused the previous failures; examples directly specify the transformation without relying on ambiguous wording."}, {"letter": "B", "text": "Ask Claude to research five phone-number libraries, extract the formatted output each produces for sample international numbers, and adopt the most frequent convention as the standard.", "correct": false, "explanation": "Surveying external libraries introduces an irrelevant research step and does not define the specific formatting required by the architect's database; it also fails to address extension handling directly."}, {"letter": "C", "text": "Provide 2-3 concrete input/output pairs covering domestic, international, and extension cases so Claude can infer the exact rule from examples alone, not prose.", "correct": true, "explanation": "Concrete input/output examples covering domestic, international, and extension cases allow Claude to infer the exact transformation rule directly, resolving edge cases far more reliably than prose descriptions."}, {"letter": "D", "text": "Repeat the original instruction but prepend each rule with all-caps modifiers like IMPORTANT and MUST so that Claude gives the normalization task greater weight.", "correct": false, "explanation": "Adding emphasis modifiers like IMPORTANT or MUST may help Claude prioritize a task, but they do not clarify the desired output format for ambiguous edge cases like extensions and international codes; the underlying ambiguity remains."}], "correct": "C", "select": 1, "group": "I"}, {"id": "f3-006", "domain": 3, "task_id": "3.1", "objective": "Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization", "situation": "A large repository has a shared set of code-review rule files that several separate repositories across the organization should reuse verbatim, updated from one source of truth so all repositories stay in sync automatically when the source changes. An architect proposes placing symlinks inside each repository's .claude/rules/ directory pointing back to a central rules folder maintained outside those repositories.", "question": "Is this a supported way to organize rules, and what should the architect verify?", "options": [{"letter": "A", "text": "It is supported; .claude/rules/ resolves symlinks normally with circular detection, so verify each repo's symlinks target the shared source", "correct": true, "explanation": "The .claude/rules/ directory supports symlinks with normal resolution and circular symlink detection, so linking to a shared external source is a valid pattern; the architect just needs to confirm the symlink targets resolve correctly on each developer's machine."}, {"letter": "B", "text": "It is supported, but only when the shared rules sit in the same git repo as a submodule; symlinks to a separate repository never resolve", "correct": false, "explanation": "Symlink resolution for rules is not restricted to targets inside a git submodule of the same repository; it works for any resolvable filesystem path the symlink points to."}, {"letter": "C", "text": "It is not supported; .claude/rules/ only discovers regular files, so symlinked entries are silently ignored, forcing physical copies", "correct": false, "explanation": "This is incorrect; .claude/rules/ explicitly supports symlinks, so this claim about them being silently ignored is false."}, {"letter": "D", "text": "It is supported only for whole symlinked directories, not individual files, so a single shared file like security.md cannot be linked alone", "correct": false, "explanation": "Both entire shared directories and individual shared files (such as a single security.md) can be symlinked into .claude/rules/; it isn't restricted to whole-directory linking only."}], "correct": "A", "select": 1, "group": "D"}, {"id": "f3-007", "domain": 3, "task_id": "3.1", "objective": "Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization", "situation": "During a long working session, an engineer triggers a context compaction. Afterward, they notice Claude has stopped following an instruction that was in a CLAUDE.md file located deep inside a subdirectory the engineer had already worked in earlier in the session, while an instruction from the project-root CLAUDE.md is still being followed correctly.", "question": "What explains this difference in behavior?", "options": [{"letter": "A", "text": "Nested CLAUDE.md files are deleted from context permanently after compaction and can only be restored by starting an entirely new session", "correct": false, "explanation": "The instruction is not permanently lost; it becomes available again as soon as Claude reads a file in that subdirectory, without requiring a brand-new session."}, {"letter": "B", "text": "Compaction only preserves instructions in the first 200 lines of a session, so the subdirectory file's later position caused it to drop", "correct": false, "explanation": "This describes the unrelated 200-line/25KB limit that applies to auto memory's MEMORY.md file, not to CLAUDE.md compaction behavior, and mischaracterizes how compaction treats nested files."}, {"letter": "C", "text": "Compaction corrupts nested CLAUDE.md files on disk, so the subdirectory file must be re-created before its instructions work again", "correct": false, "explanation": "Compaction summarizes conversation context; it does not modify or corrupt files on disk, and no file recreation is needed."}, {"letter": "D", "text": "Root CLAUDE.md is re-read and re-injected after compaction, but nested CLAUDE.md files reload only when Claude next reads a file there", "correct": true, "explanation": "Project-root CLAUDE.md is specifically re-read and re-injected after /compact, while nested subdirectory CLAUDE.md files are not automatically re-injected and only reload the next time Claude reads a file in that subdirectory."}], "correct": "D", "select": 1, "group": "D"}, {"id": "f3-008", "domain": 3, "task_id": "3.5", "objective": "Apply iterative refinement techniques for progressive improvement", "situation": "Before implementing a multi-region deployment strategy for a service Claude Code has not worked on before, an architect wants to surface failure-mode considerations such as what happens during a regional outage, how state is reconciled after a network partition heals, and which region is authoritative during a split.", "question": "What is the recommended way to elicit this input before implementation?", "options": [{"letter": "A", "text": "Write out every failure scenario, including regional outages and partition recovery, in exhaustive detail personally beforehand, so that no interview or clarifying questions from Claude are necessary to surface failure-mode considerations.", "correct": false, "explanation": "Exhaustively writing scenarios alone relies solely on the architect's existing foresight, missing the benefit of Claude's questioning to uncover unanticipated failure modes. The interview approach complements the architect's knowledge, ensuring more thorough coverage."}, {"letter": "B", "text": "Give Claude a brief description of the deployment goal and have it interview the architect in detail using the AskUserQuestion tool, focusing on failure modes and tradeoffs rather than obvious questions.", "correct": true, "explanation": "Using the AskUserQuestion tool allows Claude to interview the architect, probing for failure-mode considerations specific to this service. This approach surfaces tradeoffs and scenarios the architect may not have anticipated, making it ideal for an unfamiliar area before implementation."}, {"letter": "C", "text": "Have Claude begin writing the deployment configuration immediately, and rely on production incidents to surface failure modes such as regional outages and partition recovery, adjusting the configuration iteratively as issues arise.", "correct": false, "explanation": "Deferring to production incidents to discover failure modes is risky and costly, as it can lead to outages and data loss. The interview pattern is designed to proactively surface such concerns during planning, avoiding reactive fixes."}, {"letter": "D", "text": "Ask Claude to summarize a generic multi-region deployment tutorial, and then adopt its default failure-handling mechanisms, such as failover and consistency, without tailoring them to this service's specific regional constraints.", "correct": false, "explanation": "Adopting generic tutorial defaults without discussing this service's unique constraints risks overlooking critical failure scenarios like authoritative region selection during a split. Tailoring is essential for a reliable multi-region deployment."}], "correct": "B", "select": 1, "group": "I"}, {"id": "f3-009", "domain": 3, "task_id": "3.3", "objective": "Apply path-specific rules for conditional convention loading", "situation": "A codebase currently has separate CLAUDE.md files inside src/api/tests/, src/web/tests/, and packages/shared/tests/, each repeating nearly identical testing conventions. These per-directory CLAUDE.md files only load when Claude happens to read a file inside that specific subdirectory. The team wants one definition of testing conventions that reliably applies to every test file, no matter which directory it is added to in the future.", "question": "What is the best solution?", "options": [{"letter": "A", "text": "Keep the three subdirectory CLAUDE.md files as the definitive source for testing rules because subdirectory CLAUDE.md files automatically apply to any test file created within those directories and will cascade to new test subdirectories over time.", "correct": false, "explanation": "Subdirectory CLAUDE.md files only load when Claude reads a file inside that particular directory; they do not cascade or automatically apply to newly created subdirectories. Maintaining separate files leads to duplication and does not guarantee that conventions will apply to future directories unless explicitly added. The official guidance favors path-scoped rules over per-directory files for cross-cutting concerns like testing conventions."}, {"letter": "B", "text": "Consolidate the three files into a single top-level CLAUDE.md so the shared testing conventions load into every session and are applied proactively whenever a test-related context is detected, regardless of directory structure.", "correct": false, "explanation": "While a top-level CLAUDE.md would make the conventions available globally, it unnecessarily loads them into every session, consuming valuable context window tokens even when no testing work is being done. Anthropic recommends keeping root CLAUDE.md files concise (under 200 lines) and using modular rules in .claude/rules/ for conditional, path-specific activation, which is more efficient and contextual."}, {"letter": "C", "text": "Replace the three subdirectory CLAUDE.md files with one .claude/rules/testing.md file using a paths glob that matches the project's test file naming (e.g., \"**/*.test.*\" or \"**/*.spec.*\"), so the convention applies by file type across the repo rather than depending on directory structure.", "correct": true, "explanation": "This is the recommended approach per Anthropic's official documentation. The .claude/rules/ directory supports modular, path-scoped rules that load only when Claude interacts with matching files. By using a glob pattern tailored to your actual test file naming (e.g., **/*.test.* for projects using .test. suffixes, or **/*.spec.* for .spec. conventions), the testing conventions are automatically applied to all test files across the entire repository, regardless of their directory location. This method conserves context window tokens and ensures consistent, scalable application of rules without depending on directory hierarchy."}, {"letter": "D", "text": "Rename the three subdirectory CLAUDE.md files to CLAUDE.local.md so they are excluded from version control while still loading automatically whenever Claude navigates into those specific testing directories, providing consistent test conventions locally.", "correct": false, "explanation": "CLAUDE.local.md files are designed for personal, uncommitted preferences and are not intended for shared team conventions. They would not provide a single source of truth for the whole team and would still only apply scoped to those exact directories. The correct solution is to commit a path-scoped rule in .claude/rules/ so that all team members automatically adhere to the same test conventions across the entire codebase."}], "correct": "C", "select": 1, "group": "D"}, {"id": "f3-010", "domain": 3, "task_id": "3.6", "objective": "Integrate Claude Code into CI/CD pipelines", "situation": "A team configures a single long-running Claude Code session that both implements a feature and then reviews its own diff before opening the PR. QA notices this workflow catches noticeably fewer bugs than an independent reviewer would.", "question": "What is the underlying reason a self-review in the same session tends to be weaker?", "options": [{"letter": "A", "text": "Claude Code caches tool call results per session, causing the review step to reuse stale file contents from before the edits were made", "correct": false, "explanation": "Claude Code does not silently serve stale cached file contents for tool calls; each Read reflects the current file state, so this is not the mechanism behind weaker self-review."}, {"letter": "B", "text": "The --print flag disables the Read tool after the first turn, so the reviewing pass cannot re-open files it already edited", "correct": false, "explanation": "The -p/--print flag does not disable tools after the first turn; Claude Code retains full tool access for the duration of a print-mode invocation."}, {"letter": "C", "text": "The session retains the reasoning and assumptions used while writing the code, so it tends to re-apply the same blind spots when checking its own output instead of evaluating it fresh", "correct": true, "explanation": "Because the same session already committed to an approach, it carries forward the mental model and assumptions from implementation, making it prone to confirming its own reasoning rather than critically re-examining the change the way an independent session would."}, {"letter": "D", "text": "The session runs out of available context window tokens by the time review begins, forcing Claude to skip reading large portions of the diff", "correct": false, "explanation": "Context exhaustion can happen in long sessions, but it is not the general reason self-review is weaker; the effect occurs even in sessions well within context limits."}], "correct": "C", "select": 1, "group": "D"}, {"id": "f3-011", "domain": 3, "task_id": "3.2", "objective": "Create and configure custom slash commands and skills", "situation": "A developer sets context: fork and agent: Explore on a pr-summary skill that fetches PR data and summarizes it. They notice the summary never reflects conventions written in the project's CLAUDE.md.", "question": "Why not?", "options": [{"letter": "A", "text": "The pr-summary skill requires the allowed-tools: Read permission to include CLAUDE.md in the forked subagent's context, and without that permission, the file is not loaded even when the project has conventions.", "correct": false, "explanation": "Anthropic's documentation does not indicate that a Read tool permission is a prerequisite for loading CLAUDE.md. CLAUDE.md is typically provided as context automatically, not via a tool invocation. The absence of conventions here is due to the Explore agent's design, not a missing permission."}, {"letter": "B", "text": "The built-in Explore agent skips loading CLAUDE.md at startup to keep its context small, so a forked skill using that agent only sees the skill content and the agent's own system prompt.", "correct": true, "explanation": "Anthropic's recommended practices for context engineering emphasize progressive disclosure and minimizing initial context. The Explore agent is designed to keep its context lean by omitting broad, project-level files like CLAUDE.md at startup. This allows it to function efficiently with only the skill content and its system prompt, loading additional context only when needed."}, {"letter": "C", "text": "CLAUDE.md conventions are only loaded when a skill is invoked without arguments, and since the pr-summary skill receives PR data as an argument, the conventions are omitted from the forked subagent's context.", "correct": false, "explanation": "Anthropic's Agent Skills and CLAUDE.md loading do not depend on whether arguments are passed to a skill. CLAUDE.md is generally read at conversation start or during context bootstrap, regardless of arguments. This option invents an unsupported conditional-loading rule."}, {"letter": "D", "text": "context: fork enforces strict isolation by removing project-level files like CLAUDE.md from every subagent's context, so even when using agent: Explore, the agent never sees the conventions and cannot apply them.", "correct": false, "explanation": "There is no official documentation stating that context: fork universally strips CLAUDE.md from all subagent contexts. While fork may isolate certain elements, the behavior described is not the primary reason CLAUDE.md conventions are missing here. The absence is specifically tied to the Explore agent's context-management design."}], "correct": "B", "select": 1, "group": "D"}, {"id": "f3-012", "domain": 3, "task_id": "3.2", "objective": "Create and configure custom slash commands and skills", "situation": "A developer on a shared repository wants a personal /my-standup slash command that formats their daily standup notes and should not appear in teammates' / menus or be included in pull requests.", "question": "According to current Anthropic Claude Code skill best practices, where should they create this personal slash command?", "options": [{"letter": "A", "text": "In .claude/commands/my-standup.md, then add a note in the PR description asking reviewers to ignore that file.", "correct": false, "explanation": ".claude/commands/ is inside the project repository, so the file would be committed and visible to teammates who pull the repo. A note asking reviewers to ignore it does not prevent Claude Code from loading it or stop it from appearing as a project slash command."}, {"letter": "B", "text": "In .claude/commands/my-standup.md with a leading underscore in the filename so teammates' clients skip loading it.", "correct": false, "explanation": "A leading underscore has no documented effect that causes Claude Code to skip a project command. Because .claude/commands/ is inside the repository, the command would still be committed and available to teammates' clients."}, {"letter": "C", "text": "In ~/.claude/commands/my-standup.md, which is a legacy personal command location that still works but is not the current recommended skill-based structure.", "correct": false, "explanation": "for current best practices. Legacy personal slash commands can be placed under ~/.claude/commands/, but Anthropic now recommends the directory-based skill structure ~/.claude/skills/<name>/SKILL.md as the canonical approach. The skills format supports additional files and automatic discovery, so the question's current-best-practice requirement points to that structure instead."}, {"letter": "D", "text": "In ~/.claude/skills/my-standup/SKILL.md, which is a personal skill not committed to the repository and available as a slash command only to that developer.", "correct": true, "explanation": "Anthropic documents custom skills in Claude Code as filesystem-based directories containing SKILL.md; personal skills belong in ~/.claude/skills/, while project skills belong in .claude/skills/. A personal skill placed here is discovered automatically, becomes available as /my-standup, and is not committed to the shared repository, so teammates' / menus and pull requests are unaffected."}], "correct": "D", "select": 1, "group": "D"}, {"id": "f3-013", "domain": 3, "task_id": "3.6", "objective": "Integrate Claude Code into CI/CD pipelines", "situation": "A CI job asks Claude Code to generate new unit tests for a modified module. The generated suite repeatedly proposes tests that duplicate scenarios already covered by the existing test file for that module.", "question": "What should the pipeline change to reduce this redundant output?", "options": [{"letter": "A", "text": "Run the job with --no-session-persistence so no test history carries over between CI runs, forcing each run to generate tests without reference to past test patterns.", "correct": false, "explanation": "Disabling session persistence only stops Claude from referencing past CI run history; it does not supply the current test file as context. Without that context, Claude remains unaware of existing scenarios and may still produce duplicates."}, {"letter": "B", "text": "Pass the existing test files for the module into Claude's context alongside the modified source, so generation is scoped against what is already covered.", "correct": true, "explanation": "By passing the existing test files alongside the modified source into Claude's context, the model can identify which scenarios are already covered and generate only new tests. This directly addresses the root cause of duplicate output."}, {"letter": "C", "text": "Raise --max-budget-usd for the job so Claude can afford to write a larger, more exhaustive test suite that covers every case twice to double-check each scenario for complete coverage.", "correct": false, "explanation": "A larger budget allows Claude to produce more content but does not provide it with the current test suite, so redundant tests would still be generated. Additionally, intentionally covering every case twice would increase duplication rather than reduce it."}, {"letter": "D", "text": "Switch the job to --output-format stream-json and use a client-side script to deduplicate generated test events against the module's current test file.", "correct": false, "explanation": "Switching to stream-json output and using a script to deduplicate events may remove duplicates after the fact, but it does not prevent Claude from generating redundant tests in the first place, wasting time and compute. The streaming format itself does not give Claude knowledge of existing coverage."}], "correct": "B", "select": 1, "group": "D"}, {"id": "f3-014", "domain": 3, "task_id": "3.1", "objective": "Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization", "situation": "In a monorepo, an engineer working exclusively on packages/web finds that Claude Code's context is cluttered at startup with CLAUDE.md content from packages/admin-dashboard and several packages/legacy-* packages. This happens because the monorepo's root CLAUDE.md file uses import statements to load all subpackage CLAUDE.md files. The engineer wants these excluded from their own sessions without affecting other teammates who might work in those packages.", "question": "What should they do?", "options": [{"letter": "A", "text": "Delete the CLAUDE.md files from packages/admin-dashboard and packages/legacy-* directly, since unused files should be removed from the repository", "correct": false, "explanation": "Deleting shared CLAUDE.md files would remove project documentation and instructions for teammates who do work in those packages, causing team-wide regressions. The use case calls for a personal, local exclusion, not a repository-level deletion."}, {"letter": "B", "text": "Add claudeMdExcludes patterns for those packages to the committed .claude/settings.json at the repository root, since exclusions can only be configured at the project scope", "correct": false, "explanation": "A committed .claude/settings.json is shared project configuration and would apply team-wide, not just to this engineer. Local exclusions belong in .claude/settings.local.json, while project-level settings are for durable shared configuration."}, {"letter": "C", "text": "Add claudeMdExcludes patterns for those packages to .claude/settings.local.json, since local settings apply only to that engineer's machine", "correct": true, "explanation": "Anthropic's documentation describes claudeMdExcludes as a way to exclude rules in monorepos or prevent irrelevant instructions from loading into Claude Code sessions. .claude/settings.local.json is intended for personal overrides within a project and is typically excluded from version control via .gitignore, so the exclusions affect only that engineer and not other teammates."}, {"letter": "D", "text": "Ask the admin-dashboard and legacy package owners to move their CLAUDE.md files into .claude/rules/, since only root-level CLAUDE.md files are loaded across package boundaries", "correct": false, "explanation": "The documented mechanism for this scenario is claudeMdExcludes to keep specific CLAUDE.md files from loading, not moving files into .claude/rules/. Also, changing file locations in shared packages would impact other teammates and is not a local, personal override."}], "correct": "C", "select": 1, "group": "D"}, {"id": "f3-015", "domain": 3, "task_id": "3.4", "objective": "Determine when to use plan mode vs direct execution", "situation": "A platform team is restructuring a monolith into three microservices, touching shared data models, service boundaries, and deployment topology, with several viable ways to split the code.", "question": "Before any files are changed, which approach should the architect direct Claude Code to take?", "options": [{"letter": "A", "text": "Run the restructuring through a CI pipeline with bypassPermissions enabled so the extraction finishes without any manual review step", "correct": false, "explanation": "bypassPermissions removes prompts and safety checks entirely, which is the opposite of the deliberate, reviewable investigation an architectural restructuring calls for."}, {"letter": "B", "text": "Start in the default accept-edits mode so Claude can begin extracting services immediately and adjust the boundaries as issues surface", "correct": false, "explanation": "Accept-edits mode lets Claude write files immediately, so boundary mistakes get baked into multiple files before anyone reviews the overall design."}, {"letter": "C", "text": "Enter plan mode so Claude explores the codebase, evaluates the competing service boundaries, and proposes a design before any edit is approved", "correct": true, "explanation": "Plan mode restricts Claude to read-only exploration and requires an approved plan before edits happen, which is exactly what a multi-file architectural decision with several valid approaches needs to avoid costly rework."}, {"letter": "D", "text": "Skip investigation and have Claude pick one service split at random right away, then rely on integration tests to catch structural mistakes later", "correct": false, "explanation": "Picking a split without investigation discards the architectural analysis the task needs and pushes discovery of structural errors to test time, after the damage is already spread across files."}], "correct": "C", "select": 1, "group": "C"}, {"id": "f3-016", "domain": 3, "task_id": "3.4", "objective": "Determine when to use plan mode vs direct execution", "situation": "A developer asks Claude to add a check that rejects an expiration date earlier than today's date inside one existing form-validation function. The function and its surrounding file are already well understood by the team.", "question": "What is the most efficient way to handle this request?", "options": [{"letter": "A", "text": "Treat the request as a multi-file migration and draft a phased plan covering all forms in the application", "correct": false, "explanation": "The request only concerns one function, not every form in the application, so treating it as a multi-file migration inflates the scope beyond what was asked."}, {"letter": "B", "text": "Proceed with direct execution, since the change is confined to one function with a clear, well-scoped requirement", "correct": true, "explanation": "Adding a single validation conditional to one already-understood function is a simple, well-scoped change, which direct execution handles efficiently without added process."}, {"letter": "C", "text": "Delegate the task to an Explore subagent to catalog every date-handling function in the repository before writing the conditional", "correct": false, "explanation": "There is no broad discovery need for a change scoped to one known function, so routing it through the Explore subagent adds overhead without benefit."}, {"letter": "D", "text": "Enter plan mode first to explore whether other validation functions in the codebase should also be restructured", "correct": false, "explanation": "The request does not ask for restructuring other validation functions; expanding scope into plan mode here investigates a problem that was never posed."}], "correct": "B", "select": 1, "group": "C"}, {"id": "f3-017", "domain": 3, "task_id": "3.5", "objective": "Apply iterative refinement techniques for progressive improvement", "situation": "A payments engineer asks Claude Code to implement a discount calculation function and wants to minimize back-and-forth correction cycles from the outset. Order totals must round to the nearest cent using banker's rounding, and stacking discounts must apply in a specific order that is easy to get wrong.", "question": "What should the engineer do before Claude writes any implementation code?", "options": [{"letter": "A", "text": "Describe the rounding and stacking rules in a single long paragraph of prose and trust that a sufficiently detailed paragraph will be interpreted correctly the first time", "correct": false, "explanation": "A single prose paragraph describing rules that are easy to get wrong is exactly the kind of description prone to inconsistent interpretation; a test suite provides a much more reliable, checkable target."}, {"letter": "B", "text": "Write a test suite first that encodes the rounding and stacking rules with concrete cases, then have Claude implement against it and iterate on any failures", "correct": true, "explanation": "Writing a test suite that encodes the specific rounding and stacking behavior first turns an easy-to-get-wrong set of rules into a concrete, executable check that Claude can implement against and use to drive focused iteration on any failures."}, {"letter": "C", "text": "Ask Claude to implement the function using whatever rounding and stacking order seems most common in e-commerce systems, then correct it if the checkout totals look wrong later", "correct": false, "explanation": "Guessing at common e-commerce conventions risks not matching the engineer's actual banker's-rounding and stacking-order requirements, and catching the mismatch later in checkout totals is a costlier way to find the same bugs a test suite would catch immediately."}, {"letter": "D", "text": "Have Claude implement the function without tests, then manually spot-check three or four representative orders by hand in a spreadsheet", "correct": false, "explanation": "Manual spreadsheet spot-checks of a few orders provide a weaker, non-repeatable signal than an executable test suite that can be rerun automatically as Claude iterates."}], "correct": "B", "select": 1, "group": "I"}, {"id": "f3-018", "domain": 3, "task_id": "3.1", "objective": "Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization", "situation": "A repository already has an AGENTS.md file used by several other AI coding tools, containing conventions the team wants Claude Code to follow as well, plus a short list of Claude-specific instructions like 'use plan mode for changes under src/billing/'. The team wants to avoid maintaining the same conventions in two places.", "question": "What is the recommended way to structure CLAUDE.md?", "options": [{"letter": "A", "text": "Leave CLAUDE.md absent entirely, since Claude Code silently falls back to reading AGENTS.md whenever no CLAUDE.md file is present", "correct": false, "explanation": "Claude Code does not natively read AGENTS.md files; there is no silent fallback. Without CLAUDE.md, Claude Code will not pick up any project‑specific conventions from AGENTS.md."}, {"letter": "B", "text": "Create CLAUDE.md starting with @AGENTS.md as an import, followed by the Claude-specific instructions such as the plan-mode rule underneath", "correct": true, "explanation": "This is the officially recommended method. The @AGENTS.md import loads the shared conventions, and you can append Claude‑specific instructions below it. This avoids duplication and keeps everything in one file that Claude Code reads automatically at the start of every session. While a symbolic link (symlink) is an alternative workaround, it is not the recommended approach because it may require administrator permissions or Developer Mode on Windows; the import method works cross‑platform without extra privileges."}, {"letter": "C", "text": "Create a symbolic link (symlink) named CLAUDE.md pointing to AGENTS.md so both files always share the same content", "correct": false, "explanation": "Although a symlink is technically possible and sometimes mentioned as a workaround, it is not the officially recommended method. Symlinks can require administrator permissions or Developer Mode on Windows, making the import method more universally compatible. Anthropic's documentation recommends the @AGENTS.md import approach as the primary way to include AGENTS.md content in CLAUDE.md while still allowing Claude‑specific additions."}, {"letter": "D", "text": "Manually copy the full contents of AGENTS.md into CLAUDE.md today, and remember to re-copy it by hand every time AGENTS.md changes", "correct": false, "explanation": "Manually copying and re-syncing creates a maintenance burden and is error-prone. Anthropic's documentation explicitly recommends using an @-import to avoid duplication and keep instructions in sync automatically."}], "correct": "B", "select": 1, "group": "D"}, {"id": "f3-019", "domain": 3, "task_id": "3.5", "objective": "Apply iterative refinement techniques for progressive improvement", "situation": "A data engineer is iterating with Claude Code on a migration script that transforms legacy customer records into a new schema. After the first pass, records with a null middle-name field are being dropped instead of migrated with an empty string.", "question": "What is the most effective way to fix this specific edge case handling?", "options": [{"letter": "A", "text": "Provide a specific test case showing a record with a null middle-name field as input and the expected output record with an empty string in that field", "correct": true, "explanation": "Supplying a specific test case with the exact input containing the null field and the exact expected output pins down precisely how that edge case should be handled, directly resolving the ambiguity."}, {"letter": "B", "text": "Ask Claude to add a generic try/except block around the entire transformation function so that any record causing an error is skipped silently", "correct": false, "explanation": "Silently skipping records that cause errors would hide the null middle-name records instead of migrating them, which does not fix the underlying requirement to include them with an empty string."}, {"letter": "C", "text": "Instruct Claude to rewrite the entire migration script from scratch using a different scripting language in the hope that the new implementation avoids the issue", "correct": false, "explanation": "Rewriting the script in a different language does not address the specific null-handling logic and risks introducing unrelated new defects instead of fixing the identified edge case."}, {"letter": "D", "text": "Tell Claude the migration script has a bug involving null values without specifying which field or what the corrected output should look like", "correct": false, "explanation": "A vague description of \"a bug involving null values\" leaves Claude to guess which field and what the correct output should be, which is far less reliable than a concrete example."}], "correct": "A", "select": 1, "group": "I"}, {"id": "f3-020", "domain": 3, "task_id": "3.6", "objective": "Integrate Claude Code into CI/CD pipelines", "situation": "A reviewer bot re-runs after every new commit on a PR, and previously-posted inline comments about a missing null check keep reappearing as duplicate comments each time, even though the author never addressed them.", "question": "What change to the re-run invocation would most directly stop the duplicate postings?", "options": [{"letter": "A", "text": "Include the prior review findings in the new run's context and instruct Claude to only report findings not already present in those prior findings, thereby avoiding duplicate comments.", "correct": true, "explanation": "Anthropic's context engineering guidance recommends providing prior review findings in the new run's context and explicitly instructing Claude to report only new findings not already present. This can be implemented with XML tags or persisted notes, and it directly prevents duplicate comments without relying on the bot to check PR comments on its own."}, {"letter": "B", "text": "Add --no-session-persistence to the re-run invocation, causing each run to start with no memory of prior findings and thus preventing duplicate comments about issues like the missing null check.", "correct": false, "explanation": "--no-session-persistence would clear prior context or memory between runs, meaning the bot would start fresh without knowledge of prior findings. This is likely to cause duplicate comments, not prevent them. The recommended fix is to include prior review findings in the new run's context and instruct Claude to report only new findings."}, {"letter": "C", "text": "Switch the re-run's --output-format from json to stream-json so the bot streams findings, checks each against existing PR comments, and only posts those not already present, avoiding duplicate comments.", "correct": false, "explanation": "Changing output format from json to stream-json only affects how results are delivered, not how findings are deduplicated. The bot still lacks a mechanism to fetch and compare existing PR comments, so streaming output alone does not prevent duplicate posts."}, {"letter": "D", "text": "Increase --max-turns on the re-run to a higher value so the bot can review its previous comments on the PR, recognize the null check was already reported, and avoid posting it again.", "correct": false, "explanation": "--max-turns controls how many agentic turns or iterations the bot can execute, not whether it retrieves and compares prior PR comments. A longer run does not make the bot check existing comments; Anthropic's documented approach for avoiding duplicate findings is to explicitly include prior findings in the context."}], "correct": "A", "select": 1, "group": "D"}, {"id": "f3-021", "domain": 3, "task_id": "3.5", "objective": "Apply iterative refinement techniques for progressive improvement", "situation": "A product architect wants Claude Code to build a real-time collaborative document editor, a domain the architect has not built in before. Before any implementation starts, the architect wants to surface cache invalidation strategy, conflict resolution approach, and failure handling considerations that might not be obvious up front.", "question": "What technique should the architect use?", "options": [{"letter": "A", "text": "Search for an open-source collaborative editor, instruct Claude to replicate its architecture, and assume that following its design patterns will automatically surface appropriate cache invalidation and conflict resolution strategies.", "correct": false, "explanation": "Copying an open-source architecture does not automatically surface the right design considerations for a different product context; those patterns may hide assumptions or fail under different requirements. The recommended approach is to elicit and examine tradeoffs with the architect before implementation, not to rely on a copied architecture's implicit choices."}, {"letter": "B", "text": "Give Claude a brief description of the feature and ask it to interview the architect using the AskUserQuestion tool, exploring technical implementation, edge cases, and tradeoffs before writing a spec.", "correct": true, "explanation": "This leverages an interactive requirements-gathering interview before any implementation starts, which is the recommended way to surface non-obvious design considerations such as cache invalidation, conflict resolution, and failure handling. The provided research does not independently verify the AskUserQuestion tool name, but it does confirm that thoroughly exploring technical implementation, edge cases, and tradeoffs with an architect before writing a specification is a fundamental recommended practice."}, {"letter": "C", "text": "Write a complete technical specification personally covering every design decision, including cache invalidation, conflict resolution, and failure handling, then hand it to Claude as a fixed set of implementation instructions.", "correct": false, "explanation": "Writing a complete fixed specification up front often misses considerations that are not obvious to the architect, especially in an unfamiliar domain. The recommended practice is to clarify constraints, gather feedback, and explore tradeoffs collaboratively before finalizing the spec, not to lock all decisions in beforehand."}, {"letter": "D", "text": "Ask Claude to implement a minimal collaborative editor, then use every bug and edge case found through manual testing as the primary source to surface cache invalidation, conflict resolution, and failure handling design considerations.", "correct": false, "explanation": "This approach starts implementation before surfacing design considerations, which defeats the \"before any implementation starts\" requirement. Iterative testing and bug review are valuable later, but they are not a substitute for upfront architectural exploration of cache invalidation, conflict resolution, and failure handling."}], "correct": "B", "select": 1, "group": "I"}, {"id": "f3-022", "domain": 3, "task_id": "3.6", "objective": "Integrate Claude Code into CI/CD pipelines", "situation": "A team's release-notes generator asks Claude to extract structured data and passes the --json-schema flag with a schema string that contains a typo (an unescaped brace).", "question": "What should they expect to happen?", "options": [{"letter": "A", "text": "claude silently falls back to unstructured plain-text output with no indication that the schema was invalid", "correct": false, "explanation": "Anthropic's documentation specifies that invalid JSON schemas cause the request to fail with an explicit error, not a silent fallback. The purpose of structured outputs is to guarantee schema-compliant output; falling back to unstructured text would defeat this guarantee and is not implemented."}, {"letter": "B", "text": "claude exits with an error stating the value is not a valid JSON Schema, along with the validator's diagnostic, instead of producing any output", "correct": true, "explanation": "Anthropic's Structured Outputs feature enforces strict JSON Schema Draft 2020-12 validation on input schemas. When a malformed schema (e.g., unescaped brace) is provided, the API returns a 400 status code with an 'invalid_request_error' type and a diagnostic message. Claude Code, which wraps the API, surfaces this error and exits without generating any output. This aligns with official documentation stating that such client-side errors must be fixed before retrying."}, {"letter": "C", "text": "claude retries the request against the API up to three times using --max-retries before giving up and printing the raw text response", "correct": false, "explanation": "Automatic retries in Anthropic SDKs are designed for transient errors (e.g., rate limits, server errors). A 400 error due to an invalid JSON schema is a permanent, client-side error and is explicitly not retried per documentation. The schema must be corrected before resubmitting."}, {"letter": "D", "text": "claude ignores the malformed schema and instead of failing, returns an empty structured_output object with no fields", "correct": false, "explanation": "The system does not partially apply or ignore an invalid schema. Constrained decoding relies on a valid grammar compiled from the schema; if the schema cannot be parsed, the request is rejected with an error. An empty structured_output object would still be a response, but no response is generated at all."}], "correct": "B", "select": 1, "group": "D"}, {"id": "f3-023", "domain": 3, "task_id": "3.2", "objective": "Create and configure custom slash commands and skills", "situation": "A project ships a shared .claude/skills/commit/SKILL.md skill that writes commit messages in a style one developer finds too terse for their own habits. The developer wants a personal richer version of that skill while continuing to invoke the same /commit slash command themselves. Teammates should continue to see the original project skill when they run /commit.", "question": "What should they do?", "options": [{"letter": "A", "text": "Create ~/.claude/skills/commit/SKILL.md as a personal copy with the same name to locally override the project skill for this developer only, leaving the shared project skill unchanged for teammates.", "correct": true, "explanation": "Personal skills in ~/.claude/skills/ are user-level and apply only to the local developer. By using the same skill name commit, the developer's environment can prefer the personal version when /commit is invoked, while teammates using the project's .claude/skills/commit/SKILL.md continue to see the original behavior. This meets both requirements: the developer keeps invoking /commit and teammates are unaffected."}, {"letter": "B", "text": "Add disable-model-invocation: true to the shared .claude/skills/commit/SKILL.md so that it no longer generates output and only the developer’s personal instructions apply.", "correct": false, "explanation": "disable-model-invocation is a valid skill frontmatter field, but editing the shared project skill changes behavior for all users, not just this developer. Teammates would see the skill disabled instead of the original commit-message style, and the developer would still not have a personal richer version. Isolation requires a local user-level change, not a modification to the shared file."}, {"letter": "C", "text": "Create a differently named skill, such as ~/.claude/skills/commit-verbose/SKILL.md, so it's invoked separately and the shared project skill remains untouched for teammates.", "correct": false, "explanation": "While creating a differently named personal skill does leave the shared project skill untouched, it also changes the slash command the developer must use to /commit-verbose. The requirement explicitly states the developer wants to continue using /commit, so this does not satisfy the scenario."}, {"letter": "D", "text": "Edit .claude/skills/commit/SKILL.md with the richer commit guidelines, and rely on the change remaining uncommitted so only the developer’s local experience uses it, while teammates’ copies are unaffected.", "correct": false, "explanation": "Project skills in .claude/skills/ are meant to be shared through version control. Leaving the edit uncommitted is fragile: any commit, sync, or checkout will propagate the change to teammates, violating the isolation requirement. The durable, local-only approach is to use a personal skill in ~/.claude/skills/."}], "correct": "A", "select": 1, "group": "D"}, {"id": "f3-024", "domain": 3, "task_id": "3.4", "objective": "Determine when to use plan mode vs direct execution", "situation": "An architect finishes a plan mode session where Claude proposed a design for restructuring a shared utilities package used by 12 services. The architect wants Claude to now make the changes exactly as proposed, without re-explaining the reasoning behind each file edit.", "question": "What should the architect do next?", "options": [{"letter": "A", "text": "Start an entirely new session from scratch and re-describe the utilities package restructuring in full before implementation begins", "correct": false, "explanation": "Starting over in a fresh session discards the plan and context Claude already built during the investigation, requiring the architect to redo work that plan mode already completed."}, {"letter": "B", "text": "Stay in plan mode indefinitely and have Claude describe each proposed file edit in the chat instead of ever actually making the edit", "correct": false, "explanation": "Staying in plan mode keeps edits blocked indefinitely, so the restructuring across the 12 services would never actually be applied to the files."}, {"letter": "C", "text": "Exit the session without approving anything and instead manually perform the twelve-service restructuring by hand rather than having Claude do it", "correct": false, "explanation": "Manually performing a 12-service restructuring by hand discards the value of having Claude execute the already-approved plan and reintroduces the risk of inconsistent manual changes."}, {"letter": "D", "text": "Approve the plan and choose an option that switches the session into an editing mode, so Claude proceeds directly with the approved design", "correct": true, "explanation": "Approving the plan and choosing an edit-enabling option exits plan mode and switches the session into a mode where Claude proceeds with implementing the design that was just investigated and approved."}], "correct": "D", "select": 1, "group": "C"}, {"id": "f3-025", "domain": 3, "task_id": "3.5", "objective": "Apply iterative refinement techniques for progressive improvement", "situation": "An architect asks Claude Code to write a function that converts free-form date strings like next Tuesday, 03/04/25, and the 1st of March into ISO 8601 dates. The architect has specified that ambiguous numeric inputs in MM/DD/YY format should be interpreted month-first, so 03/04/25 must mean March 4, 2025. Early attempts still inconsistently interpret ambiguous formats such as 03/04/25.", "question": "What is the most effective next step to fix this transformation ambiguity?", "options": [{"letter": "A", "text": "Rewrite the prompt using more forceful language demanding that Claude \"get the dates right this time\" without adding any new information, but emphasize that ambiguous formats must be resolved consistently.", "correct": false, "explanation": "Emphasis alone does not resolve ambiguity. Without adding new information or concrete examples, Claude still lacks a precise rule for interpreting 03/04/25, even though the stem defines the intended convention. Prompt engineering should add specificity, not just stronger wording."}, {"letter": "B", "text": "Ask Claude to throw an exception whenever an ambiguous format like 03/04/25 is encountered, preventing any ambiguous conversion and requiring the user to provide an explicit format specification.", "correct": false, "explanation": "This approach would be defensible if no interpretation rule had been specified, but the stem now establishes a month-first convention. With that business rule defined, rejecting 03/04/25 is unnecessarily restrictive and prevents valid conversions. The most effective next step is to encode the rule via concrete input-output examples, not to reject the input."}, {"letter": "C", "text": "Add a comment in the code explaining that date parsing is a hard problem and ask Claude to use its best judgment, for example by defaulting to month-first when formats like 03/04/25 are ambiguous.", "correct": false, "explanation": "Relying on best judgment and an implicit default leaves the parsing rule ambiguous and can still produce inconsistent results. Anthropic emphasizes using explicit examples and format conventions rather than asking the model to guess, especially when a concrete interpretation rule has already been defined."}, {"letter": "D", "text": "Provide 2-3 example input strings, including 03/04/25, each paired with the exact ISO 8601 output the architect expects (e.g., 03/04/25 → 2025-03-04), so the ambiguous format resolves to a concrete rule.", "correct": true, "explanation": "Anthropic's tool-use guidance recommends using concrete examples to establish format conventions. Because the stem now specifies month-first interpretation for MM/DD/YY, a paired example like 03/04/25 → 2025-03-04 explicitly encodes the required rule. This removes ambiguity and prevents inconsistent parsing, whereas implicit assumptions or stronger wording do not."}], "correct": "D", "select": 1, "group": "I"}, {"id": "f3-026", "domain": 3, "task_id": "3.6", "objective": "Integrate Claude Code into CI/CD pipelines", "situation": "A platform team wraps claude inside a GitHub Actions job that fires on every pull request to summarize the diff. The first run hangs until the job times out, with no output ever written to the log. The team confirms the API key secret is valid and the prompt text is correct.", "question": "What is the most likely cause of the hang?", "options": [{"letter": "A", "text": "The job checked out the repository with a shallow clone depth, so Claude Code could not resolve the diff and stalled while scanning git history", "correct": false, "explanation": "A shallow clone can limit git history available to Claude, but it does not cause a process to hang waiting on input; Claude would still print output or an error."}, {"letter": "B", "text": "The job invoked claude without the -p flag, so Claude Code started in interactive mode and sat waiting for terminal input that the runner never provides", "correct": true, "explanation": "Without -p (--print), Claude Code launches its interactive TUI, which blocks waiting on stdin. CI runners provide no interactive terminal, so the process hangs indefinitely until the job times out."}, {"letter": "C", "text": "The workflow forgot to set ANTHROPIC_API_KEY as a masked secret, so the CLI silently retried authentication in a loop until the runner timeout", "correct": false, "explanation": "An invalid or missing API key produces an explicit authentication error in the log almost immediately, not a silent indefinite hang with zero output."}, {"letter": "D", "text": "The runner's Node.js version is older than what the claude binary requires, so the process stayed alive while failing to parse its own CLI arguments", "correct": false, "explanation": "An incompatible Node/runtime version typically causes the CLI to fail fast with a startup error, not hang silently through the entire job timeout."}], "correct": "B", "select": 1, "group": "D"}, {"id": "f3-027", "domain": 3, "task_id": "3.6", "objective": "Integrate Claude Code into CI/CD pipelines", "situation": "A GitHub Actions PR-review workflow starts a review, then needs to ask a follow-up question in a second claude -p invocation later in the same job while explicitly targeting the same conversation (in case multiple reviews run concurrently on different PRs in parallel jobs).", "question": "What is the correct approach?", "options": [{"letter": "A", "text": "Capture the session_id from the first call's --output-format json response, then pass it to the second call with --resume \"$session_id\"", "correct": true, "explanation": "Capturing session_id from the first --output-format json response and passing it explicitly to --resume targets that specific conversation, which is the reliable approach when multiple concurrent sessions could otherwise be ambiguous."}, {"letter": "B", "text": "Pass --continue to the second call, which is guaranteed to reattach to the correct conversation even when several reviews run concurrently", "correct": false, "explanation": "--continue resumes the most recent conversation scoped to the current project directory, which becomes ambiguous or wrong when multiple reviews are running concurrently in parallel jobs rather than sequentially in one directory."}, {"letter": "C", "text": "Use --no-session-persistence on both calls so the two invocations share an in-memory session automatically", "correct": false, "explanation": "--no-session-persistence explicitly disables saving the session to disk so it cannot be resumed at all, which is the opposite of what's needed to link a follow-up call to a specific prior session."}, {"letter": "D", "text": "Re-send the entire original prompt text verbatim in the second call so Claude reconstructs the same context from scratch", "correct": false, "explanation": "Re-sending the original prompt does not give Claude access to intermediate reasoning, file reads, or tool results from the first call; it only restates the initial instructions, not the accumulated session context."}], "correct": "A", "select": 1, "group": "D"}, {"id": "f3-028", "domain": 3, "task_id": "3.5", "objective": "Apply iterative refinement techniques for progressive improvement", "situation": "A developer wants Claude Code to implement a rate limiter for an internal API and wants to reduce the number of correction cycles needed afterward.", "question": "Which opening approach best follows a test-driven iteration pattern?", "options": [{"letter": "A", "text": "Ask Claude to implement the rate limiter and manually verify it by running sample requests such as bursts, retries, and checking rate limit headers in the terminal, without automated tests.", "correct": false, "explanation": "Manual checks via terminal sampling of bursts, retries, and headers do not create a repeatable, shareable failure signal. Without automated tests, there is no structured basis for iterative refinement and verifying correctness after changes."}, {"letter": "B", "text": "Ask Claude to implement the rate limiter first, then write a matching test suite afterward that validates throttle responses, burst handling, and window resets to document the built behavior.", "correct": false, "explanation": "Writing tests after implementation only documents the behavior that was built, so it cannot catch misunderstandings of intended throttling or reset rules. This approach does not guide implementation and misses the opportunity to iterate based on test failures."}, {"letter": "C", "text": "Ask Claude to write a test suite covering the expected throttling behavior, burst limits, and reset timing first, then implement the rate limiter, iterating by sharing test failures.", "correct": true, "explanation": "Writing the test suite first specifying throttling behavior, burst limits, and reset timing establishes concrete expected behaviors. Iterating by sharing test failures then provides a focused signal that drives progressive refinement toward a correct implementation."}, {"letter": "D", "text": "Ask Claude to describe the rate limiting algorithm in a design document covering throttle rules, burst limits, and reset windows, then implement it from that document without writing tests.", "correct": false, "explanation": "A design document describing throttle rules, burst limits, and reset windows is only a prose specification; it does not provide executable checks. Without tests, Claude cannot verify its own implementation nor iteratively refine based on failures."}], "correct": "C", "select": 1, "group": "I"}, {"id": "f3-029", "domain": 3, "task_id": "3.5", "objective": "Apply iterative refinement techniques for progressive improvement", "situation": "During refinement of a caching layer, a code reviewer finds that a fix for a stale-read bug in the cache invalidation logic will change the locking behavior that a separate, still-open concurrency bug also touches. Both issues live in the same critical section.", "question": "How should these two issues be communicated to Claude for the next revision?", "options": [{"letter": "A", "text": "Describe both the stale-read bug and the concurrency bug together in one message, since a fix for the shared critical section will alter the locking constraints another bug relies on.", "correct": true, "explanation": "Both bugs share a critical section, so a fix for one will alter the locking constraints that the other relies on. Describing them together in a single message allows Claude to design a coherent solution that addresses the interaction and avoids creating a fix that must later be partially undone."}, {"letter": "B", "text": "Report the stale-read bug first in isolation, get its full fix merged with tests, then later file the concurrency bug as a separate issue without mentioning the shared critical section.", "correct": false, "explanation": "Fixing the stale-read bug in isolation may introduce locking changes that conflict with the unresolved concurrency bug, because both issues touch the same critical section. Handling them sequentially without cross-referencing risks wasting effort and producing a solution that is not optimal for the combined behavior."}, {"letter": "C", "text": "Ask Claude to randomly choose one of the two bugs, apply a complete fix including updated locking and cache invalidation logic, and only address the remaining bug if continuous integration tests later reveal a failure.", "correct": false, "explanation": "Randomly selecting one bug to fix and relying on CI failures to detect the interaction is unreliable and inefficient. The known interaction should be communicated directly so Claude can produce a comprehensive, deliberate fix rather than leaving the solution to chance and delayed test feedback."}, {"letter": "D", "text": "Report only the concurrency bug now, giving it full priority since it was discovered second, and defer the stale-read bug to a future sprint regardless of their shared critical section.", "correct": false, "explanation": "Deferring the stale-read bug while only addressing the concurrency bug ignores their shared critical section. A fix for the concurrency bug could alter locking in ways that complicate the stale-read issue later, making it harder to resolve; they should be tackled together to avoid regressions."}], "correct": "A", "select": 1, "group": "I"}, {"id": "f3-030", "domain": 3, "task_id": "3.3", "objective": "Apply path-specific rules for conditional convention loading", "situation": "An engineer works inside a symlinked checkout: the actual repository lives at /Users/eng/code/service and is symlinked to /workspace/service, and Claude Code is launched from the symlinked path. A path-scoped rule has paths: [\"src/handlers/**/*.go\"]. The engineer edits /workspace/service/src/handlers/middleware/auth.go.", "question": "Will the rule trigger given how Claude Code resolves paths through symlinks?", "options": [{"letter": "A", "text": "The rule triggers normally, because Claude Code matches path-scoped rules even when a file is reached through a symlinked path to the project directory, in addition to matching direct paths.", "correct": true, "explanation": "Path-scoped rules in Claude Code are evaluated against the logical path of a file—the path as observed through the symlinked working directory—not the canonical filesystem path. When editing /workspace/service/src/handlers/middleware/auth.go, the relative path from the project root (/workspace/service) is src/handlers/middleware/auth.go, which exactly matches the glob pattern src/handlers/**/*.go. Anthropic’s documentation and release notes confirm that path-scoped rules correctly trigger for files accessed via symlinked checkouts, ensuring consistent behavior regardless of whether the working directory is a symlink."}, {"letter": "B", "text": "The rule triggers only if the engineer manually adds a second paths entry pointing at /workspace/service/src/handlers/**/*.go, since symlinked roots require an explicit absolute pattern.", "correct": false, "explanation": "Absolute patterns are never required for path-scoped rules. The glob pattern src/handlers/**/*.go is relative and is matched against the file’s relative path from the project root. Even when the project root is accessed through a symlink, this relative matching works as designed, making an absolute entry redundant."}, {"letter": "C", "text": "The rule triggers only for read-only operations performed through the symlink, while edits made through the symlinked path are matched against a separate, unscoped rule set.", "correct": false, "explanation": "Path-scoped rules do not differentiate between read-only and write operations. A rule with a matching paths glob triggers for any file access (read, edit, or write) that involves a matching path, whether the path is direct or symlinked."}, {"letter": "D", "text": "The rule never triggers through a symlinked checkout, because path-scoped rules only evaluate against the canonical filesystem path returned by the operating system, bypassing any symlink entirely.", "correct": false, "explanation": "Contrary to this option, Claude Code intentionally uses the logical (symlink-aware) path for rule matching. While certain security-critical operations may resolve symlinks to canonical paths for validation, the loading of path-scoped rules is based on the path as seen from the working directory, including any symlinks in the path."}], "correct": "A", "select": 1, "group": "D"}, {"id": "f3-031", "domain": 3, "task_id": "3.1", "objective": "Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization", "situation": "A platform architect is deciding whether a new 'require code owner review before merging billing changes' policy should be enforced through a CLAUDE.md instruction or through a settings.json permission rule with a hook. The policy must hold even if Claude decides, based on its own reasoning, that skipping review would be fine in a particular case.", "question": "Which configuration choice is appropriate, and why?", "options": [{"letter": "A", "text": "Write it as a managed policy CLAUDE.md entry only, because managed policy is the sole CLAUDE.md tier that Claude cannot override regardless of its own reasoning", "correct": false, "explanation": "Even managed policy CLAUDE.md is still CLAUDE.md content — it shapes behavior but is not described as a hard enforcement mechanism the way hooks and permission settings are; its distinguishing feature is that individual settings can't exclude it, not that it forces compliance."}, {"letter": "B", "text": "Write it as a strongly worded instruction repeated in both the root CLAUDE.md and every subdirectory's CLAUDE.md, because repetition across the hierarchy is what converts guidance into enforcement", "correct": false, "explanation": "Repeating the same instruction across multiple CLAUDE.md files does not change its nature from guidance to enforcement; Claude can still deviate under its own reasoning."}, {"letter": "C", "text": "Implement it as a settings-based control such as a PreToolUse hook, because CLAUDE.md is context that shapes behavior but is not a hard enforcement layer Claude is guaranteed to obey", "correct": true, "explanation": "For instructions that must hold regardless of what Claude decides, a settings-based control like a PreToolUse hook enforces at the client level, whereas any CLAUDE.md tier remains behavioral guidance rather than a hard block."}, {"letter": "D", "text": "Write it as a CLAUDE.md instruction, because CLAUDE.md content carries the same enforcement guarantee as a technical control once it's committed to the project", "correct": false, "explanation": "CLAUDE.md instructions, including committed project-level ones, are context that Claude tries to follow but are not guaranteed enforcement; treating them as equivalent to a technical control is incorrect."}], "correct": "C", "select": 1, "group": "D"}, {"id": "f3-032", "domain": 3, "task_id": "3.2", "objective": "Create and configure custom slash commands and skills", "situation": "A platform team wants a /release-notes slash command that every engineer on the repository gets automatically after git pull, so nobody has to reinstall it manually.", "question": "Where should the team save the command file?", "options": [{"letter": "A", "text": "Save it to ~/.claude/skills/release-notes/SKILL.md so that Claude loads the command automatically in every project the team lead opens on their machine.", "correct": false, "explanation": "The ~/.claude/skills/ directory is personal to the team lead's machine and is not tracked by version control. This approach only loads the command in the lead's own projects and does not share it with the team."}, {"letter": "B", "text": "Save it to .claude/commands/release-notes.md inside the project and commit that file to the shared repository so it distributes through version control.", "correct": true, "explanation": "Project-scoped commands stored in .claude/commands/ inside the project root are normal repository files. Committing the file ensures that every engineer receives it automatically via git pull with no additional manual steps."}, {"letter": "C", "text": "Save it to ~/.claude/commands/release-notes.md on the team lead's machine and have each engineer copy the file to their own local ~/.claude/commands/ directory.", "correct": false, "explanation": "Saving the slash command file in the user-scoped ~/.claude/commands/ directory means it only works on the team lead's machine. Even if other engineers manually copy it, the process defeats the goal of automatic, git-based distribution after a pull."}, {"letter": "D", "text": "Save it to .claude/commands/release-notes.md in the project root, but add that file to .gitignore so each engineer's local changes stay conflict-free.", "correct": false, "explanation": "Adding the file to .gitignore prevents it from ever being committed to the shared repository. As a result, pulling the repository will not deliver the command to any other engineer."}], "correct": "B", "select": 1, "group": "D"}, {"id": "f3-033", "domain": 3, "task_id": "3.6", "objective": "Integrate Claude Code into CI/CD pipelines", "situation": "A security team wants a nightly Claude Code job to behave identically no matter which self-hosted runner picks it up, without being affected by a stray MCP server defined in one runner's shared .mcp.json or a hook left in a teammate's ~/.claude directory.", "question": "Which combination of choices best achieves this reproducibility goal?", "options": [{"letter": "A", "text": "Invoke claude with --output-format stream-json, and pipe the event stream through a filter that discards any hook or MCP-related events from .mcp.json or ~/.claude, so only the intended assistant content appears in the final output.", "correct": false, "explanation": "Redirecting the output with --output-format stream-json and filtering events only affects what appears in the final output; it does not prevent hooks or MCP servers from being auto-discovered and executed. The unwanted modifications from those components can still occur, compromising reproducibility."}, {"letter": "B", "text": "Invoke claude with --bare in print mode and pass only the explicit flags needed (such as --append-system-prompt or --settings) so nothing is auto-discovered from the working directory or home folder.", "correct": true, "explanation": "Using --bare disables automatic discovery of hooks, MCP servers, skills, plugins, and CLAUDE.md files, ensuring the runner's local or home-directory configuration is ignored. Only the explicitly provided flags (like --append-system-prompt or --settings) are applied, delivering consistent behavior across all runners."}, {"letter": "C", "text": "Invoke claude in ordinary -p mode, and set --max-turns to a low number such as 2, which limits any hook or MCP server from .mcp.json or ~/.claude to at most two interactions, preventing them from altering the final output.", "correct": false, "explanation": "Setting --max-turns to a low number only limits the number of agentic interactions; it does not stop hooks or MCP servers from being loaded and executed from .mcp.json or ~/.claude in the first place. Any discovered tool or hook can still run and alter the output before the turn limit is reached."}, {"letter": "D", "text": "Invoke claude with --continue and provide a pre-recorded session file that was created in a clean environment, so the job picks up that exact conversation state instead of auto-discovering any hooks or MCP servers from the runner.", "correct": false, "explanation": "Using --continue with a pre-recorded session file restores a previous conversation state, but it does not suppress the auto-discovery of hooks, MCP servers, or other configuration from the current runner's filesystem. Any local artifacts in .mcp.json or ~/.claude will still be loaded and may influence the job's behavior."}], "correct": "B", "select": 1, "group": "D"}, {"id": "f3-034", "domain": 3, "task_id": "3.2", "objective": "Create and configure custom slash commands and skills", "situation": "After a developer adds a /lint-report skill, it works when they run Claude Code themselves, but a teammate who pulled the latest commit reports the skill doesn't exist. The developer confirms the file was never staged in any commit.", "question": "What is the most likely cause?", "options": [{"letter": "A", "text": "The skill was saved under .claude/commands/lint-report.md, which requires a separate enterprise license to be shared across teammates.", "correct": false, "explanation": "There is no enterprise licensing requirement for sharing a project-scoped skill; committing the .claude/skills/ directory to version control is what makes it available to teammates."}, {"letter": "B", "text": "The skill was saved correctly, but skill availability only activates after the project's CI pipeline runs once successfully.", "correct": false, "explanation": "Skill availability is not gated behind a CI pipeline run."}, {"letter": "C", "text": "The skill was saved under ~/.claude/skills/lint-report/SKILL.md, which lives outside the repository and was never version-controlled.", "correct": true, "explanation": "The personal ~/.claude/skills/ directory lives outside the git repository, so a skill placed there is never staged, committed, or shared with anyone else who clones or pulls the project; only a skill committed under the project's .claude/skills/ directory reaches teammates."}, {"letter": "D", "text": "The skill was saved under .claude/skills/lint-report/SKILL.md, but the teammate needs to restart their machine before new skills are detected.", "correct": false, "explanation": "Claude Code does not require a restart to detect a newly added skill, and this doesn't explain why the file was never staged in a commit to begin with."}], "correct": "C", "select": 1, "group": "D"}, {"id": "f3-035", "domain": 3, "task_id": "3.5", "objective": "Apply iterative refinement techniques for progressive improvement", "situation": "A test run against a newly generated report-export module in Claude Code produces two failures: a currency-formatting mismatch in the summary totals and a completely unrelated failure where the export filename generator omits the file extension on Windows paths.", "question": "How should these two failures be presented to Claude for the next iteration?", "options": [{"letter": "A", "text": "Combine both failures into a single message that shows the full test run output, noting that the currency-formatting mismatch and the filename extension omission are unrelated defects, and ask Claude to fix both.", "correct": true, "explanation": "Claude Code's feedback-loop guidance is to run the tests and show the full output in one message so Claude can act on everything at once. Naming both failures as unrelated, rather than implying a shared cause, keeps Claude from wasting effort hunting for a connection between them that does not exist."}, {"letter": "B", "text": "Report each failure sequentially in its own message: first address the currency-formatting mismatch, and once confirmed fixed, the filename extension issue, because the two problems are independent.", "correct": false, "explanation": "Anthropic's own Claude Code guidance directs running the full test suite and reporting the complete output in one iteration, not withholding one failure until the other is confirmed fixed. Splitting two already-diagnosed, simple failures across separate turns adds review overhead with no documented benefit."}, {"letter": "C", "text": "Ask Claude to evaluate the test results and fix whichever failure appears more severe first, relying on its internal judgment of the impact on report correctness without stating the exact currency mismatch or missing extension details.", "correct": false, "explanation": "Asking Claude to guess which failure is more severe and omitting the actual currency-formatting and filename-extension details removes the specific information Claude needs to diagnose and fix each issue, which slows the loop down and risks a misdiagnosis instead of speeding it up."}, {"letter": "D", "text": "Withhold the filename extension failure and report only the currency-formatting mismatch as the primary defect found first, deferring the secondary extension issue to avoid distracting Claude from the more critical formatting fix.", "correct": false, "explanation": "Withholding a known failure removes information Claude needs to fix the report correctly and gains nothing: presenting both failures together in the same test-output message is no more distracting than presenting one, and it lets both get fixed in the same iteration instead of a later one."}], "correct": "A", "select": 1, "group": "I"}, {"id": "f3-036", "domain": 3, "task_id": "3.4", "objective": "Determine when to use plan mode vs direct execution", "situation": "A team wants Claude to evaluate three candidate caching strategies for a high-traffic API, each requiring different changes to the request-handling layer, before any strategy is implemented. Once a strategy is picked, the actual code changes are expected to be small and localized to two files.", "question": "How should the architect sequence this work?", "options": [{"letter": "A", "text": "Implement all three caching strategies with direct execution first and delete two of them once the comparison is complete", "correct": false, "explanation": "Fully implementing three caching strategies just to discard two wastes effort and risks committing side effects across the request-handling layer that are costly to undo, which is the rework plan mode is meant to prevent."}, {"letter": "B", "text": "Use plan mode to compare the three caching strategies and settle on one, then leave plan mode and apply the small change with direct execution", "correct": true, "explanation": "Comparing multiple valid approaches before committing to code is exactly what plan mode is for, and once a strategy is chosen, the small, localized two-file change is well suited to direct execution."}, {"letter": "C", "text": "Use direct execution for the strategy comparison itself, then switch into plan mode only for the small two-file implementation that follows afterward", "correct": false, "explanation": "This reverses the appropriate sequence: the multi-approach comparison needs the safe, no-edit exploration of plan mode, while the already-scoped small change doesn't need a plan mode review afterward."}, {"letter": "D", "text": "Skip the comparison entirely and implement whichever strategy is mentioned first, adjusting the choice later if it underperforms", "correct": false, "explanation": "Choosing arbitrarily and adjusting later after underperformance is discovered defeats the purpose of comparing strategies up front and risks costly rework in the request-handling layer."}], "correct": "B", "select": 1, "group": "C"}, {"id": "f3-037", "domain": 3, "task_id": "3.4", "objective": "Determine when to use plan mode vs direct execution", "situation": "While in plan mode, an architect notices that Claude keeps reading files and running read-only search commands without prompting for approval on each one, even though no edits have occurred.", "question": "What most plausibly explains this behavior?", "options": [{"letter": "A", "text": "The architect's local settings have disabled every single permission check for the remainder of this particular session entirely", "correct": false, "explanation": "Claude Code's permission model is granular; there is no single setting that disables all permission checks entirely. Approvals are controlled per action category and mode, and disabling all checks would violate the principle of least privilege. The observed behavior is better explained by the design of plan mode itself."}, {"letter": "B", "text": "Plan mode has silently switched over to acceptEdits mode partway through this session, which is why operations proceed without prompts", "correct": false, "explanation": "Plan mode is explicitly designed to be read-only and does not silently switch to a mode that allows edits. Any change to a mode that permits edits must be explicitly triggered by the user, and such a switch would not explain why no edits have occurred."}, {"letter": "C", "text": "Read-only commands are permanently exempt from every permission mode in Claude Code, no matter which mode is currently active", "correct": false, "explanation": "Read-only commands are not universally exempt from permission checks. Outside of plan mode, manual approval may be required depending on settings, and documented regressions have caused read operations to prompt for permission. The permanent exemption claim is too broad and inaccurate."}, {"letter": "D", "text": "Plan mode inherently blocks all edits and does not require approval for read operations, as the mode's read-only guarantee makes approval prompts unnecessary.", "correct": true, "explanation": "According to Anthropic's documentation, Plan mode is always read-only and does not edit source files. Because the mode itself prevents any destructive actions, read operations like file reads and search commands execute without prompting for user approval. This design eliminates the need for manual approval on safe operations while ensuring edits are blocked."}], "correct": "D", "select": 1, "group": "C"}, {"id": "f3-038", "domain": 3, "task_id": "3.3", "objective": "Apply path-specific rules for conditional convention loading", "situation": "A CLAUDE.md file has grown past 300 lines because it contains detailed conventions for GraphQL resolvers, database migrations, and CSS modules, each relevant only to a specific part of the codebase. Session-start context usage is climbing and adherence is dropping.", "question": "What is the most effective way to address this using path-specific rules?", "options": [{"letter": "A", "text": "Wrap the GraphQL, migration, and CSS sections in HTML maintainer comments within CLAUDE.md so they are stripped before injection, then re-insert each section manually into the conversation whenever Claude happens to work on that area.", "correct": false, "explanation": "HTML comments are stripped before injection for maintainer notes, not as a mechanism for conditional loading, and manually re-inserting content defeats the purpose of automated context management."}, {"letter": "B", "text": "Copy the same GraphQL, migration, and CSS sections into a new CLAUDE.local.md file so it loads alongside the existing CLAUDE.md, giving Claude two overlapping sources of the identical guidance to reinforce adherence at every session start.", "correct": false, "explanation": "Duplicating the content into CLAUDE.local.md does not reduce context usage; it doubles the amount of guidance loaded into every session instead of scoping it conditionally."}, {"letter": "C", "text": "Move the entire CLAUDE.md content into a single .claude/rules/ file without a paths field, which shrinks the file on disk but keeps every topic's guidance loading into every session exactly as it did before the move.", "correct": false, "explanation": "Moving content into a rules file without a paths field still loads it unconditionally at launch with the same priority as CLAUDE.md, so context usage is not reduced even though the file moved."}, {"letter": "D", "text": "Split each topic into its own .claude/rules/ file with a paths frontmatter scoped to the relevant file types, so each set of conventions only enters context when Claude works on matching files instead of loading all of them every session.", "correct": true, "explanation": "Splitting oversized CLAUDE.md content into topic-specific, path-scoped rules means each set of conventions only loads into context when relevant files are opened, cutting unnecessary context usage."}], "correct": "D", "select": 1, "group": "D"}, {"id": "f3-039", "domain": 3, "task_id": "3.3", "objective": "Apply path-specific rules for conditional convention loading", "situation": "In a Claude Code rules file, the same conventions must apply to .ts and .tsx files anywhere under src/, plus .ts files under lib/, using as few paths glob pattern entries as possible.", "question": "Which paths frontmatter accomplishes this most efficiently?", "options": [{"letter": "A", "text": "paths:\n  - \"src/**/*.{ts,tsx}\"\n  - \"lib/**/*.ts\"", "correct": true, "explanation": "Claude Code path-specific rules use YAML frontmatter with a paths field containing glob patterns. The brace syntax *.{ts,tsx} matches either extension in a single entry, while ** matches files recursively. This covers .ts and .tsx anywhere under src/ and .ts anywhere under lib/ in only two entries, making it the most compact and token-efficient option shown in Anthropic's documentation."}, {"letter": "B", "text": "paths:\n  - \"**/*.ts\"\n  - \"**/*.tsx\"", "correct": false, "explanation": "This applies the conventions to every .ts and .tsx file anywhere in the project, including .tsx files under lib/ and files outside src/ and lib/. That is broader than the requested scope and can load irrelevant rules into context, reducing token efficiency."}, {"letter": "C", "text": "paths:\n  - \"src/*.{ts,tsx}\"\n  - \"lib/*.ts\"", "correct": false, "explanation": "Without **, src/*.{ts,tsx} and lib/*.ts only match files directly in the top-level src/ or lib/ directories and do not recurse into nested subdirectories. The requirement states the rules must apply anywhere under src/ and under lib/, so these patterns do not satisfy the scenario."}, {"letter": "D", "text": "paths:\n  - \"src/**/*.ts\"\n  - \"src/**/*.tsx\"\n  - \"lib/**/*.ts\"", "correct": false, "explanation": "This pattern set is functionally correct and covers the requested files, but it uses three separate entries instead of two. The question asks for the most efficient approach, and Anthropic's documented brace syntax *.{ts,tsx} expresses the same src/ coverage in a single entry."}], "correct": "A", "select": 1, "group": "D"}, {"id": "f3-040", "domain": 3, "task_id": "3.6", "objective": "Integrate Claude Code into CI/CD pipelines", "situation": "A nightly job invokes claude -p \"run the migration script\" --allowedTools \"Bash(npm run migrate*)\" intending to allow only the exact migrate command, but the run unexpectedly also executes an unrelated command npm run migrate-cleanup without prompting.", "question": "What causes this broader match than intended?", "options": [{"letter": "A", "text": "The --bare flag was omitted, causing Claude Code to fall back to a permissive default that widens all --allowedTools prefix matches", "correct": false, "explanation": "No --bare flag exists in Claude Code that influences --allowedTools behavior. The matching rules are deterministic and based solely on the provided patterns; omitting a non‑existent flag does not change how Bash patterns are evaluated."}, {"letter": "B", "text": "acceptEdits was implicitly enabled alongside --allowedTools, which broadens matching to include any command containing the word migrate", "correct": false, "explanation": "There is no acceptEdits parameter that interacts with --allowedTools in this way. The --allowedTools flag only controls which tools the agent can use, and Bash matching is governed solely by prefix patterns, not by implicit broadening from other options."}, {"letter": "C", "text": "The trailing asterisk in Bash(npm run migrate*) triggers prefix matching, so the pattern matches any command that starts with 'npm run migrate', including npm run migrate-cleanup, not just the exact npm run migrate command.", "correct": true, "explanation": "Anthropic's official documentation states that Bash rules use prefix matching, not regex. The asterisk in Bash(npm run migrate*) causes the pattern to match any command that begins with \"npm run migrate\", such as npm run migrate-cleanup. To restrict the match to the exact command, drop the asterisk and specify Bash(npm run migrate) without any trailing wildcard."}, {"letter": "D", "text": "--allowedTools patterns are matched using regular expression alternation by default, so 'migrate*' matches 'migrate' or any suffix starting with a letter", "correct": false, "explanation": "--allowedTools does not use regular expressions. Bash tool patterns are simple prefix matches, as documented. The wildcard * is a prefix‑match operator, not a regex alternation or wildcard for arbitrary suffixes."}], "correct": "C", "select": 1, "group": "D"}, {"id": "f3-041", "domain": 3, "task_id": "3.5", "objective": "Apply iterative refinement techniques for progressive improvement", "situation": "", "question": "An architect is scoping a new internal search feature and gives Claude Code the single-line prompt \"I want to build a search feature for our support ticket system, interview me in detail using the AskUserQuestion tool.\" What is the most effective way for the architect to use the resulting interview before implementation starts?", "options": [{"letter": "A", "text": "Answer all of Claude's interview questions thoroughly in a single session, then instruct Claude to immediately start coding the search feature directly, relying on the conversation transcript as the sole record rather than first creating a written specification or design document.", "correct": false, "explanation": "While answering all interview questions thoroughly is valuable, proceeding directly to implementation without first creating a written specification discards the distilled clarity a spec provides. Relying solely on the conversation transcript forces the implementation to compete with the full interview context, increasing the risk of ambiguity or oversight."}, {"letter": "B", "text": "Skip the interactive interview and instead provide Claude with the existing support ticket schema, letting it infer the required search parameters, ranking rules, and filtering logic from the field names, data types, and existing relationships in that schema alone.", "correct": false, "explanation": "Providing only the ticket schema omits crucial design decisions about ranking strategy, UI behavior, and failure modes—none of which can be reliably inferred from field names and data types alone. The interview responses are essential to resolve these open questions before implementation."}, {"letter": "C", "text": "Continue answering Claude's questions until ranking, filtering, and failure-mode considerations are all covered. Then have Claude write a complete selfcontained spec naming files and interfaces involved, and start a fresh session focused only on implementing that spec.", "correct": true, "explanation": "By continuing until ranking, filtering, and failure-mode considerations are addressed, the architect ensures all critical design decisions are surfaced. Having Claude produce a self-contained spec and then starting a fresh session focused solely on that spec keeps implementation uncluttered and aligned with documented requirements."}, {"letter": "D", "text": "Answer only the first two or three questions Claude asks, then instruct Claude to immediately start implementing the search feature, filling in the missing ranking, filtering, and schema decisions with its own assumptions based on typical support ticket search patterns.", "correct": false, "explanation": "Answering only the first few questions and relying on Claude's assumptions skips the deeper, unanticipated questions that the interview is designed to uncover, leaving gaps in ranking, filtering, or failure handling. This defeats the purpose of the iterative refinement process and risks building on uninformed guesses."}], "correct": "C", "select": 1, "group": "I"}, {"id": "f3-042", "domain": 3, "task_id": "3.6", "objective": "Integrate Claude Code into CI/CD pipelines", "situation": "A CI pipeline needs to parse Claude's response programmatically to decide whether a build step should fail. The team wants the reply to always match a specific JSON shape, for example an object with a boolean passed field and an array of issues strings, regardless of how Claude phrases its reasoning.", "question": "Which invocation achieves this?", "options": [{"letter": "A", "text": "claude -p \"review this diff\" --output-format json --json-schema '{\"type\":\"object\",\"properties\":{\"passed\":{\"type\":\"boolean\"},\"issues\":{\"type\":\"array\",\"items\":{\"type\":\"string\"}}},\"required\":[\"passed\",\"issues\"]}'", "correct": true, "explanation": "Combining --output-format json with --json-schema validates the model's structured output against the supplied schema and returns it in the structured_output field, giving a guaranteed machine-parseable shape."}, {"letter": "B", "text": "claude -p \"review this diff and reply only with passed:true/false and a list of issues\" --output-format text", "correct": false, "explanation": "Asking nicely in the prompt with plain text output gives no schema enforcement; Claude can vary phrasing or omit fields, and the CI script would need fragile text parsing."}, {"letter": "C", "text": "claude -p \"review this diff\" --output-format stream-json --verbose --include-partial-messages", "correct": false, "explanation": "stream-json streams turn-by-turn events for real-time display, it does not validate or constrain the final answer to a specific object schema."}, {"letter": "D", "text": "claude -p \"review this diff\" --append-system-prompt \"Always respond with valid JSON matching {passed, issues}\"", "correct": false, "explanation": "Appending schema instructions to the system prompt is a soft hint with no validation; the model can still drift from the exact shape, unlike an enforced --json-schema."}], "correct": "A", "select": 1, "group": "D"}, {"id": "f3-043", "domain": 3, "task_id": "3.2", "objective": "Create and configure custom slash commands and skills", "situation": "A developer builds a /fix-issue skill that expects an issue number as input, such as /fix-issue 482. When teammates invoke it without typing a number, they get confused about what to supply.", "question": "Which frontmatter field should the developer add to SKILL.md so the autocomplete menu shows teammates what to type?", "options": [{"letter": "A", "text": "argument-hint: [issue-number], which displays a placeholder in the autocomplete menu without changing how the skill executes.", "correct": true, "explanation": "The argument-hint field is the standard frontmatter key for SKILL.md that provides a placeholder string in the slash‑command autocomplete interface. It helps users understand what input is expected without altering skill execution or validation logic."}, {"letter": "B", "text": "disable-model-invocation: true, which forces teammates to always type an issue number when invoking the command manually.", "correct": false, "explanation": "disable-model-invocation: true prevents the model from automatically calling the skill, requiring manual invocation. It does not add an input hint or guide the user on what to type."}, {"letter": "C", "text": "description: Requires an issue number argument, which Claude reads aloud to the user before executing the skill.", "correct": false, "explanation": "The description field is shown in the command menu as general help text, but it does not display a placeholder for the argument. It would not appear as a visual hint in the input area."}, {"letter": "D", "text": "arguments: issue-number, which validates that the typed value is numeric before the skill is allowed to run.", "correct": false, "explanation": "arguments is not a recognized frontmatter field for SKILL.md. The correct field to provide an argument hint is argument-hint. There is no built‑in numeric validation through frontmatter."}], "correct": "A", "select": 1, "group": "D"}, {"id": "f3-044", "domain": 3, "task_id": "3.4", "objective": "Determine when to use plan mode vs direct execution", "situation": "A CI pipeline runs Claude Code non-interactively to apply a pre-approved, narrowly scoped fix: updating one hardcoded timeout value in one configuration file, with the exact new value specified in advance. There is no ambiguity about approach and no exploration needed.", "question": "Which workflow choice best fits this automated scenario?", "options": [{"letter": "A", "text": "Run the task with direct execution, since it is a single, fully specified, well-scoped edit with no open design question to investigate", "correct": true, "explanation": "A single, fully specified value change to one known file with no ambiguity is a well-scoped, well-understood change, which fits direct execution even in a non-interactive CI context."}, {"letter": "B", "text": "Reject running this task in CI entirely, since plan mode is required for every single change to any configuration file", "correct": false, "explanation": "Plan mode is not required for every configuration file change; a fully specified, single-value edit is exactly the kind of well-scoped task direct execution is meant for."}, {"letter": "C", "text": "Route the task through plan mode so Claude proposes a design for the timeout value change before the pipeline applies it", "correct": false, "explanation": "Plan mode is meant for tasks with real design decisions or multi-file scope to investigate; there is nothing to explore or propose here since the exact value and file are already specified."}, {"letter": "D", "text": "Delegate the task to an Explore subagent to search the entire repository for every single place a timeout could plausibly still be configured", "correct": false, "explanation": "There is no discovery need, since the exact file and value are already known in advance; using the Explore subagent to search for other timeout locations goes beyond what the pre-approved fix asked for."}], "correct": "A", "select": 1, "group": "C"}, {"id": "f3-045", "domain": 3, "task_id": "3.3", "objective": "Apply path-specific rules for conditional convention loading", "situation": "A developer configured .claude/rules/db-migrations.md with paths: [\"db/migrations/**/*.sql\"], but during a session where Claude edited a migration file, the conventions in that rule do not seem to have been followed. The developer wants to confirm exactly which instruction files loaded and when, to determine whether the rule ever entered context.", "question": "What should they do?", "options": [{"letter": "A", "text": "Delete and recreate db-migrations.md without a paths field, so it loads unconditionally at every launch, then confirm its content shows up by reading through the full session transcript afterward.", "correct": false, "explanation": "Removing the paths field makes the rule load unconditionally, but this does not help verify when it loaded or whether the original paths filter caused the issue. Additionally, session transcripts may not clearly show rule loading context, making this method unreliable."}, {"letter": "B", "text": "Run the /init command against the project, which regenerates every rules file from scratch and reports any pattern that fails to match a corresponding file already tracked in the repository.", "correct": false, "explanation": "There is no documented /init command that regenerates rules files and reports pattern mismatches. This option is not a valid approach in Claude Code."}, {"letter": "C", "text": "Use the InstructionsLoaded hook to log which instruction files are loaded, when they load, and why, to verify whether db-migrations.md actually loaded when the migration file was read.", "correct": true, "explanation": "The InstructionsLoaded hook is an official Claude Code feature that fires when CLAUDE.md or .claude/rules/*.md files are loaded into context. It receives information about what changed and can be used to log details, making it the recommended way to verify instruction file loading."}, {"letter": "D", "text": "Add the rule's file path to claudeMdExcludes in settings.local.json, which prints a diagnostic message confirming whether the now-excluded rule would otherwise have loaded during the session.", "correct": false, "explanation": "claudeMdExcludes is used to exclude files from loading, not to print diagnostics about whether they would have loaded. There is no documented diagnostic message triggered by adding a rule to this exclusion list."}], "correct": "C", "select": 1, "group": "D"}, {"id": "f3-046", "domain": 3, "task_id": "3.2", "objective": "Create and configure custom slash commands and skills", "situation": "A developer keeps invoking a skill that walks the entire codebase and prints a long dependency analysis. Even after the skill finishes, this lengthy output keeps consuming space in the main conversation for the rest of the session.", "question": "What frontmatter change would keep this analysis out of the main conversation's context while still returning a summary?", "options": [{"letter": "A", "text": "Add allowed-tools: Read Grep to the skill's frontmatter so the analysis is limited to read-only tools that produce shorter output.", "correct": false, "explanation": "Restricting tool access changes which tools can run, not whether the resulting output pollutes the main conversation's context."}, {"letter": "B", "text": "Add disable-model-invocation: true to the skill's frontmatter so only manual invocation triggers the lengthy dependency analysis.", "correct": false, "explanation": "Preventing automatic invocation only changes who can trigger the skill; once invoked, the verbose output would still land directly in the main conversation."}, {"letter": "C", "text": "Add context: fork to the skill's frontmatter so the analysis runs in an isolated subagent and only a summarized result is returned to the main conversation.", "correct": true, "explanation": "context: fork runs the skill in an isolated subagent context; the verbose analysis stays in that forked context and only the returned result surfaces in the main conversation."}, {"letter": "D", "text": "Add argument-hint: [directory] to the skill's frontmatter so the analysis only scans one directory at a time instead of the whole codebase.", "correct": false, "explanation": "argument-hint only shows an autocomplete placeholder; it doesn't change scope of the scan or where the output lands."}], "correct": "C", "select": 1, "group": "D"}, {"id": "f3-047", "domain": 3, "task_id": "3.3", "objective": "Apply path-specific rules for conditional convention loading", "situation": "A large monorepo team wants to organize .claude/rules/ into frontend/ and backend/ subdirectories for maintainability, for example .claude/rules/frontend/react.md and .claude/rules/backend/api.md.", "question": "Will Claude Code still discover and load these rule files correctly?", "options": [{"letter": "A", "text": "Yes, because all .md files under .claude/rules/ are discovered recursively, so organizing them into subdirectories like frontend/ and backend/ does not prevent them from being found and evaluated.", "correct": true, "explanation": "Claude Code recursively scans all .md files under .claude/rules/, so rules in subdirectories are discovered and loaded just like those in the top level. This allows teams to organize rules by subsystem without any loss of functionality."}, {"letter": "B", "text": "No, because subdirectories under .claude/rules/ are reserved for skills and commands, so any .md files placed there are misinterpreted as skill or command artifacts and are simply skipped entirely.", "correct": false, "explanation": "Subdirectories under .claude/rules/ are not reserved for skills or commands; they are a supported way to organize rule files. Any .md files placed in them are properly recognized and processed as rules."}, {"letter": "C", "text": "No, because Claude Code only ever scans the top level of .claude/rules/ for .md files, so any rules placed in subdirectories such as frontend/ and backend/ are never discovered or loaded at all.", "correct": false, "explanation": "Claude Code performs recursive discovery of .md files under .claude/rules/, so rules in subdirectories like frontend/ and backend/ are still found and loaded. There is no limitation to the top-level directory."}, {"letter": "D", "text": "Only partially, because while the files are discovered, rules in subdirectories load unconditionally without respecting any paths frontmatter, so they are applied to all files, not just those matching the path constraints.", "correct": false, "explanation": "The behavior of rule files (whether they are unconditional or path‑scoped) is determined entirely by the paths field in their frontmatter, not by their subdirectory location. Rules in subdirectories still respect any path constraints defined in their metadata."}], "correct": "A", "select": 1, "group": "D"}, {"id": "f3-048", "domain": 3, "task_id": "3.4", "objective": "Determine when to use plan mode vs direct execution", "situation": "A team lead reviewing a completed library migration finds that early file conversions used one pattern, while later files in the same migration used a different, incompatible pattern, forcing a second pass to reconcile them.", "question": "Which earlier decision most likely caused this outcome?", "options": [{"letter": "A", "text": "The migration was executed directly from the start instead of first exploring the codebase in plan mode to settle on one pattern", "correct": true, "explanation": "Without first exploring the codebase and settling on one approach in plan mode, direct execution can lock in whatever pattern is discovered first, leading to inconsistent conversions once different patterns are found partway through the work."}, {"letter": "B", "text": "The migration was scoped to a single library instead of being combined into one larger project with an unrelated framework upgrade", "correct": false, "explanation": "Combining the migration with an unrelated framework upgrade would add scope, but the described problem is inconsistent conversion patterns, which stems from skipping upfront investigation, not from scoping choices about a second upgrade."}, {"letter": "C", "text": "The team approved a design in plan mode and then switched to direct execution to carry out the approved conversion steps", "correct": false, "explanation": "Planning first and then executing the approved steps is the recommended sequence for a migration like this and would tend to prevent the inconsistency, not cause it."}, {"letter": "D", "text": "The team used the Explore subagent to catalog usage patterns across the codebase before starting the conversion work itself", "correct": false, "explanation": "Using the Explore subagent to catalog patterns beforehand is a preventive step, not a cause of the described inconsistency; it would have helped surface the pattern variation earlier."}], "correct": "A", "select": 1, "group": "C"}, {"id": "f3-049", "domain": 3, "task_id": "3.3", "objective": "Apply path-specific rules for conditional convention loading", "situation": "An API team wants a rule requiring input validation and OpenAPI comments to apply only to TypeScript files inside src/api/, including deeply nested route handler subfolders, but not to TypeScript files elsewhere in the repo such as src/utils/ or src/components/.", "question": "Which paths frontmatter correctly scopes the rule to this requirement?", "options": [{"letter": "A", "text": "paths:\n  - \"src/api/**/*.ts\"", "correct": true, "explanation": "This pattern restricts matches to TypeScript files under src/api/, and the double-star ensures nested route handler subfolders are included."}, {"letter": "B", "text": "paths:\n  - \"src/api/*.ts\"", "correct": false, "explanation": "The single-star pattern only matches files directly inside src/api/, so it misses TypeScript files in nested route handler subfolders like src/api/users/handlers.ts."}, {"letter": "C", "text": "paths:\n  - \"src/**/*.ts\"", "correct": false, "explanation": "This pattern matches every TypeScript file anywhere under src/, including src/utils/ and src/components/, which the team explicitly wants excluded from the rule's scope."}, {"letter": "D", "text": "paths:\n  - \"api/**/*.ts\"", "correct": false, "explanation": "This pattern omits the src/ prefix, so it would not match files under the actual src/api/ path and would fail to scope the rule as intended."}], "correct": "A", "select": 1, "group": "D"}, {"id": "f3-050", "domain": 3, "task_id": "3.4", "objective": "Determine when to use plan mode vs direct execution", "situation": "A junior engineer asks why an architect insisted on plan mode for a task that adds a new authentication provider, when the engineer felt it could have started with direct edits right away. Investigation shows the task requires new session-handling logic, changes to token storage, updates to five different service entry points, and a decision between two competing library options.", "question": "Which justification best explains the architect's choice?", "options": [{"letter": "A", "text": "The task spans multiple files, involves a real choice between two competing libraries, and touches core session and token-handling behavior directly", "correct": true, "explanation": "Multi-file changes, a choice between competing valid approaches, and changes with architectural implications for session and token handling are exactly the conditions under which plan mode is designed to explore and settle a design before edits are made."}, {"letter": "B", "text": "Direct execution cannot make changes to more than one file at a time, so plan mode was the only technically available option here", "correct": false, "explanation": "Direct execution is capable of editing multiple files; the choice to use plan mode here is about managing an unresolved architectural decision and scope, not a technical limitation on file count."}, {"letter": "C", "text": "Plan mode should be used for any task that involves authentication, since security-related code is always exempt from direct execution by policy", "correct": false, "explanation": "There is no blanket rule tying plan mode usage to the word authentication; the actual reasons are the task's multi-file scope and open design decision, not the topic area alone."}, {"letter": "D", "text": "The engineer's instinct was correct, and the architect should have started with direct execution since every requirement was already fully described", "correct": false, "explanation": "The requirements were not fully settled, since a choice between two competing libraries remained open, which is exactly the kind of ambiguity plan mode is meant to resolve before committing to code."}], "correct": "A", "select": 1, "group": "C"}, {"id": "f3-051", "domain": 3, "task_id": "3.3", "objective": "Apply path-specific rules for conditional convention loading", "situation": "An engineering organization maintains ten separate repositories that should all follow the same security review checklist whenever Claude edits authentication-related code. The security team wants a single source of truth that automatically stays in sync across all ten repos without copy-pasting the rule file.", "question": "What should they do?", "options": [{"letter": "A", "text": "Paste the checklist into each repository's root CLAUDE.md file and rely on team members across ten repositories to manually update every copy whenever the checklist itself changes.", "correct": false, "explanation": "Manually pasting and updating ten separate copies is exactly the synchronization burden the team wants to avoid, and is prone to drift when someone forgets to update one repo."}, {"letter": "B", "text": "Store the security checklist as a rules file in a shared location, then symlink it into each repository's .claude/rules/ directory so Claude Code resolves and loads it normally from every project.", "correct": true, "explanation": "The .claude/rules/ directory supports symlinks, so a shared rules file can be linked into multiple projects and stays in sync automatically since every repo reads the same underlying file."}, {"letter": "C", "text": "Create a scheduled workflow that copies the checklist file into each repository's .claude/rules/ directory on a nightly schedule, producing a new commit in every repo each night.", "correct": false, "explanation": "A nightly copy job works but introduces unnecessary commit noise and a sync delay in every repository, when a symlink achieves the same single-source-of-truth result natively and instantly."}, {"letter": "D", "text": "Add the checklist text to a managed-settings.json claudeMd key deployed only to the security team's own machines, leaving every other engineer across the ten repositories without the guidance.", "correct": false, "explanation": "A managed CLAUDE.md deployed only to the security team's machines would not reach the other engineers across all ten repos who also need to follow the checklist."}], "correct": "B", "select": 1, "group": "D"}, {"id": "f3-052", "domain": 3, "task_id": "3.6", "objective": "Integrate Claude Code into CI/CD pipelines", "situation": "A release pipeline calls claude -p --output-format json \"summarize the changes in this release\" and pipes the result to a script that reads .result. Finance now also wants each invocation's per-model API spend recorded for cost tracking, without adding any new flags.", "question": "Where does that information already appear?", "options": [{"letter": "A", "text": "The cost data is only available by separately querying the Claude usage dashboard after the run completes", "correct": false, "explanation": "The usage dashboard is a separate aggregate view; per-invocation cost is already returned inline in the JSON response, so a dashboard lookup isn't required for this use case."}, {"letter": "B", "text": "Cost data only appears when --output-format stream-json and --verbose are both set, so the job must switch formats", "correct": false, "explanation": "Cost fields are part of the standard --output-format json response; switching to stream-json is unnecessary and changes the output shape the script already depends on."}, {"letter": "C", "text": "Per-invocation cost is written to a local .claude/usage.log file that the script must additionally parse", "correct": false, "explanation": "Claude Code does not write per-invocation cost to a local usage.log file; the cost fields are part of the JSON response returned directly by the CLI."}, {"letter": "D", "text": "The JSON response body already includes total_cost_usd along with a per-model cost breakdown alongside the result field", "correct": true, "explanation": "With --output-format json, the response payload includes total_cost_usd and a per-model cost breakdown, so scripted callers can track spend per invocation directly from the existing JSON output."}], "correct": "D", "select": 1, "group": "D"}, {"id": "f3-053", "domain": 3, "task_id": "3.3", "objective": "Apply path-specific rules for conditional convention loading", "situation": "An infrastructure team wants Claude Code to load Terraform-specific formatting and tagging conventions only when Claude edits files under the terraform/ directory, at any subfolder depth.", "question": "Which YAML frontmatter for a new .claude/rules/terraform.md file correctly scopes the rule to this requirement?", "options": [{"letter": "A", "text": "---\npaths:\n  - \"terraform/*\"\n---", "correct": false, "explanation": "A single star only matches files directly inside terraform/ and would miss files in nested subfolders such as terraform/modules/network/main.tf."}, {"letter": "B", "text": "---\npaths:\n  - \"terraform/**/*\"\n---", "correct": true, "explanation": "The paths field takes a YAML list, and the double-star glob terraform/**/* matches files under terraform/ at any depth, including nested subfolders."}, {"letter": "C", "text": "---\ninclude:\n  - \"terraform/**/*\"\n---", "correct": false, "explanation": "The frontmatter field recognized for path scoping is paths, not include; a rule using include would not be treated as path-scoped."}, {"letter": "D", "text": "---\npaths: \"terraform/**/*\"\n---", "correct": false, "explanation": "The paths field is documented as a YAML list of patterns, not a bare string; while a single pattern is common, it should be expressed as a list item under paths:."}], "correct": "B", "select": 1, "group": "D"}, {"id": "f3-054", "domain": 3, "task_id": "3.6", "objective": "Integrate Claude Code into CI/CD pipelines", "situation": "An engineering lead wants CI-invoked Claude Code reviews to consistently flag missing test coverage using the team's specific definition of a \"valuable test\" (asserts behavior, not implementation details) and to know which fixtures already exist in the test helpers directory.", "question": "Where should this project-specific guidance be encoded so every CI run picks it up automatically without being repeated in each workflow prompt?", "options": [{"letter": "A", "text": "In the GitHub Actions workflow YAML as a long inline prompt string duplicated across every review job", "correct": false, "explanation": "Duplicating standards inline in every workflow file means any update requires editing multiple YAML files and risks drift between jobs, unlike a single shared CLAUDE.md."}, {"letter": "B", "text": "In a CLAUDE.md file at the repository root describing testing standards, the valuable-test criteria, and available fixtures", "correct": true, "explanation": "CLAUDE.md is the mechanism Claude Code reads automatically for project context such as testing standards, fixture conventions, and review criteria, so it persists across every CI-invoked run without re-specifying it."}, {"letter": "C", "text": "In the pull request description template, so each contributor retypes the testing standards before requesting review", "correct": false, "explanation": "PR description templates depend on contributors manually retyping guidance, which is unreliable and does not guarantee Claude reads it during an automated CI run."}, {"letter": "D", "text": "In the --append-system-prompt flag value hardcoded into the CI runner's shell profile", "correct": false, "explanation": "Hardcoding standards into a shell profile ties them to a specific runner's environment rather than the repository, so they would not travel with the codebase or apply consistently across runners."}], "correct": "B", "select": 1, "group": "D"}, {"id": "f3-055", "domain": 3, "task_id": "3.5", "objective": "Apply iterative refinement techniques for progressive improvement", "situation": "While reviewing a diff from Claude Code that adds a background job scheduler, an architect notices two problems: the retry backoff logic uses the same shared timer instance that the job-locking logic also relies on to prevent double execution, so changing one affects the other; and separately, a log message uses the wrong log level and needs a one-word fix.", "question": "How should the architect sequence feedback on these two problems?", "options": [{"letter": "A", "text": "List all three issues as a numbered list, with each item named by its symptom (e.g., 'retry backoff', 'job-locking', 'log-level') while intentionally leaving out the fact that the retry backoff and job-locking logic are coupled through the shared timer.", "correct": false, "explanation": "Intentionally omitting the shared timer dependency removes the critical context that Claude needs to understand why these two behaviors cannot be changed independently. Without that explanation, Claude might propose a fix for one that inadvertently breaks the other."}, {"letter": "B", "text": "Describe the retry backoff and job-locking interaction in one detailed message because they depend on the same shared timer, and separately note the log-level fix as an unrelated change that should be addressed separately.", "correct": true, "explanation": "The retry backoff and job-locking logic are coupled via the shared timer, so describing them together gives Claude the context needed to avoid breaking one when fixing the other. The log-level fix is an unrelated, trivial change that should be addressed separately to avoid confusion."}, {"letter": "C", "text": "Combine all three items—the retry backoff logic, the job-locking logic, and the log-level fix—into a single message and instruct Claude to treat them as one inseparable change that must be implemented together due to their shared timer dependency.", "correct": false, "explanation": "The log-level fix is not related to the shared timer, so bundling it with the retry and locking logic would introduce noise and might cause Claude to waste effort searching for a nonexistent connection. The retry and locking logic are coupled, but the log fix is a separate cosmetic issue that should be handled independently."}, {"letter": "D", "text": "Report the log-level fix as a separate immediate change requiring confirmation, then after it is resolved describe the retry backoff and job-locking interaction that shares the same timer, ensuring they are addressed together and not conflated with the log-level fix.", "correct": false, "explanation": "Postponing the description of the timer interaction to first handle the trivial log-level fix wastes a feedback cycle and delays the more complex, coupled change. The retry and locking logic should be presented together early so Claude has full context for a coherent fix."}], "correct": "B", "select": 1, "group": "I"}, {"id": "f3-056", "domain": 3, "task_id": "3.2", "objective": "Create and configure custom slash commands and skills", "situation": "A developer creates a report-generator skill and wants to ensure it can only write and edit files, without the ability to run shell commands or delete files. They configure allowed-tools: Write Edit in the SKILL.md frontmatter without changing the permission mode.", "question": "Which statement about this approach is correct?", "options": [{"letter": "A", "text": "The allowed-tools field skips prompts for listed tools but does not restrict availability; other tools like Bash may still run.", "correct": true, "explanation": "Official documentation describes allowed-tools as tools Claude can use without asking permission when the skill is active. It pre-approves the listed Write and Edit operations, but in the normal permission mode it does not deny unlisted tools; those tools may still be available, usually after prompting. Therefore, this configuration alone does not enforce the developer's stated restriction."}, {"letter": "B", "text": "Setting allowed-tools to Write Edit automatically sandboxes the skill, disabling all tools except file operations.", "correct": false, "explanation": "allowed-tools does not sandbox the skill and does not automatically disable all unlisted tools. It only lists tools that can run without a permission prompt. To enforce a strict allowlist, combine it with permissionMode: \"dontAsk\", which denies unlisted tools outright."}, {"letter": "C", "text": "This configuration will effectively restrict the skill to only file write operations, preventing any destructive tool invocations.", "correct": false, "explanation": "allowed-tools by itself pre-approves the listed tools so Claude can use them without asking, but it is not a denial list. In the normal permission mode, unlisted tools such as Bash can still be invoked after a user permission prompt, and destructive operations may still be possible. To restrict the skill to only Write and Edit, pair allowed-tools with permissionMode: \"dontAsk\" or use explicit deny rules."}, {"letter": "D", "text": "Using allowed-tools is the recommended method to limit a skill's capabilities in compliance with least privilege principles.", "correct": false, "explanation": "allowed-tools is a useful component for declaring required tools, but it is not sufficient by itself for least privilege. Anthropic's recommended approach is to pair allowed-tools with permissionMode: \"dontAsk\" to create a deny-by-default whitelist; otherwise, unlisted tools can still be invoked after prompting."}], "correct": "A", "select": 1, "group": "D"}, {"id": "f3-057", "domain": 3, "task_id": "3.6", "objective": "Integrate Claude Code into CI/CD pipelines", "situation": "A DevOps engineer wants a locked-down CI runner where Claude Code can only use a small, explicitly-approved set of read-only commands and anything else causes the run to abort rather than prompt, since no human is present to answer a permission prompt.", "question": "Which configuration best matches this requirement?", "options": [{"letter": "A", "text": "Set --max-turns to 1, limiting Claude to one tool call so no command needing approval is ever reached and the run completes entirely without any prompts.", "correct": false, "explanation": "Restricting to a single tool call with --max-turns 1 does not define which commands are allowed; a single call could still be a disallowed action that triggers a permission prompt, causing the run to hang. This approach fails to ensure abort-on-denial and does not establish a locked-down, read-only command set."}, {"letter": "B", "text": "Set the permission mode to dontAsk, which denies anything not covered by permissions.allow rules or the built-in read-only command set without prompting.", "correct": true, "explanation": "dontAsk mode denies any command not explicitly allowed by permissions.allow rules or the built-in read-only set, without pausing for user input. This ensures the run aborts cleanly in unattended CI when an unapproved command is attempted, aligning with the locked-down requirement."}, {"letter": "C", "text": "Omit --allowedTools entirely, rely on the default interactive permission prompt to block unapproved commands, and set a short timeout so the run aborts when no one confirms.", "correct": false, "explanation": "The default interactive prompt will hang indefinitely when no human is present, even with a timeout, causing unpredictable delays and potentially failing to abort cleanly. Additionally, omitting --allowedTools does not restrict Claude Code to read-only commands, allowing arbitrary command attempts that could require approval."}, {"letter": "D", "text": "Set the permission mode to acceptEdits, so file writes and common filesystem commands get automatic approval, and the CI run continues without pausing.", "correct": false, "explanation": "acceptEdits mode automatically approves file writes and common filesystem operations, defeating the read-only restriction. While it reduces prompts for those operations, other unapproved commands may still trigger prompts, and it does not enforce a strict read-only, abort-on-denial policy."}], "correct": "B", "select": 1, "group": "D"}, {"id": "f3-058", "domain": 3, "task_id": "3.5", "objective": "Apply iterative refinement techniques for progressive improvement", "situation": "A team is using Claude Code to build a webhook ingestion service. During testing, three tests fail: one asserting that duplicate webhook deliveries are deduplicated using an idempotency key, one asserting that deduplication window expiry releases old keys correctly, and one asserting that a malformed JSON payload returns a 400 status. The first two failures both trace to the same deduplication store logic; the third is unrelated.", "question": "What is the best way to structure feedback across these three failures?", "options": [{"letter": "A", "text": "Wait until all three tests pass or fail consistently across several runs before reporting any of them, to confirm that the failures are reproducible and not transient, thereby avoiding premature reports on flaky test conditions.", "correct": false, "explanation": "Waiting through multiple runs to confirm reproducibility adds unneeded process overhead when the failure relationships are already clear from the test descriptions. Proactive, targeted feedback on the interacting issues enables faster, efficient iteration."}, {"letter": "B", "text": "Report the malformed-JSON failure first because it is the simplest, addressing the parsing error in an initial message, then combine both deduplication failures into a single follow-up message since they both involve the idempotency store.", "correct": false, "explanation": "Grouping the unrelated malformed-JSON failure into an initial message adds unnecessary complexity and delays focus on the more critical, interacting deduplication failures. There is no dependency that requires addressing the simple parsing error first."}, {"letter": "C", "text": "Report all three failures individually in separate messages, treating each test as a distinct issue so that each can be investigated in isolation, even when two failures share the same deduplication store logic, to keep the debugging process modular.", "correct": false, "explanation": "Reporting the two interacting deduplication failures separately risks a fix for one that does not resolve the other, since they share root-cause logic. Treating them as isolated issues would miss their relationship and likely lead to incomplete fixes."}, {"letter": "D", "text": "Report the two deduplication-store failures together in a single message since they interact through the same store logic and report the malformed-JSON failure separately since it does not interact with the other two.", "correct": true, "explanation": "Bundling the two deduplication failures into a single message is best because they stem from the same store logic, allowing Claude to design a coherent fix that addresses both. The malformed-JSON failure is independent and can be reported separately without complicating the deduplication solution."}], "correct": "D", "select": 1, "group": "I"}, {"id": "f3-059", "domain": 3, "task_id": "3.2", "objective": "Create and configure custom slash commands and skills", "situation": "An engineer keeps pasting the same eight-step deployment checklist into chat whenever they ask Claude to help ship a release, and the same steps have started to also live as a growing section in the project's CLAUDE.md. They want the procedure available on demand without it consuming context on every single turn of every session.", "question": "What should they do?", "options": [{"letter": "A", "text": "Move the checklist into a deploy-checklist skill under .claude/skills/, since a skill's body only loads into context when it's invoked, unlike CLAUDE.md content which loads every session.", "correct": true, "explanation": "A skill's body loads into context only when it is invoked, so moving the checklist into a deploy-checklist skill keeps it available on demand without adding to every session's baseline context, unlike CLAUDE.md content which loads every session."}, {"letter": "B", "text": "Move the checklist into a subagent definition under .claude/agents/ with the deployment steps, since subagents load conditionally only when invoked, keeping the main session focused on the current task.", "correct": false, "explanation": "Subagent definitions under .claude/agents/ describe an agent's own configuration and are not a mechanism for storing procedural checklists that load on demand. They do not conditionally inject content into the main session in the way skills do, and using them for this purpose would not keep the main session focused without overhead."}, {"letter": "C", "text": "Keep expanding the CLAUDE.md section with detailed steps and failover instructions for multiple environments, since always-loaded content lets Claude consistently follow the full procedure on every request without prompting.", "correct": false, "explanation": "Expanding the CLAUDE.md section would load the detailed steps into every session, consuming context continuously, which is exactly the recurring cost the engineer wants to avoid. The goal is to have the procedure available on demand without bloating every session's baseline context."}, {"letter": "D", "text": "Convert the checklist into a hooks entry in settings.json configured as a PreToolUse hook that runs the deployment steps before any tool invocation, since hooks execute at defined events without consuming context.", "correct": false, "explanation": "Hooks in settings.json are designed for automated event-triggered scripts, such as running a command before tool use, not for holding a multi-step interactive checklist that Claude follows conversationally. A PreToolUse hook would execute rigidly rather than providing the flexible, on-demand guidance the engineer needs."}], "correct": "A", "select": 1, "group": "D"}, {"id": "f3-060", "domain": 3, "task_id": "3.6", "objective": "Integrate Claude Code into CI/CD pipelines", "situation": "A development team is migrating their claude-code-action workflow from beta to v1.0. They previously relied on direct_prompt, custom_instructions, mode, and max_turns.", "question": "Which of the following best explains why the v1.0 action no longer uses these parameters in the same way?", "options": [{"letter": "A", "text": "The v1.0 action removed support for custom instructions and turn limits entirely because they are now managed automatically by the Anthropic API.", "correct": false, "explanation": "Custom instructions and turn limits are still supported, but they are no longer separate action inputs. Custom instruction content can be supplied with --append-system-prompt inside claude_args, while turn limits are set with --max-turns in claude_args rather than being removed."}, {"letter": "B", "text": "The v1.0 action takes a single prompt input, passes other settings as CLI arguments in claude_args, and detects the execution mode automatically.", "correct": true, "explanation": "The documented upgrade path from beta replaces direct_prompt with a single prompt input, removes the mode input because the action now detects interactive versus automation mode automatically from whether a prompt was supplied, and moves CLI options such as max_turns and model into claude_args, which the action's parameter table describes as CLI arguments passed to Claude Code. custom_instructions has no same-name flag and becomes --append-system-prompt inside claude_args."}, {"letter": "C", "text": "The parameters are still available but have been renamed to prompt_text, instructions, execution_mode, and max_steps for improved clarity.", "correct": false, "explanation": "The upgrade does not rename these inputs. direct_prompt is replaced by prompt and the remaining CLI options are passed through claude_args; no prompt_text, instructions, execution_mode, or max_steps input exists in the v1 action."}, {"letter": "D", "text": "The v1.0 action requires all configuration to be placed in a single config.yml file located in the repository root, which replaces environment variables.", "correct": false, "explanation": "Configuration for the GitHub Action is defined in the repository's workflow file, not a config.yml in the repository root. The v1.0 action uses workflow inputs such as prompt and claude_args, and does not require consolidating all settings into a single configuration file."}], "correct": "B", "select": 1, "group": "D"}, {"id": "f3-061", "domain": 3, "task_id": "3.4", "objective": "Determine when to use plan mode vs direct execution", "situation": "An architect is planning a task that involves both a broad investigation phase, reading dozens of files to understand a legacy authentication flow, and a narrow implementation phase, adding one new field to a single config file once the investigation clarifies where it belongs.", "question": "How should this task be structured for the best balance of thoroughness and context efficiency?", "options": [{"letter": "A", "text": "Run the dozens-of-files investigation directly in the main conversation without a subagent, then execute the config change directly as well", "correct": false, "explanation": "Running the dozens-of-files investigation directly in the main conversation would flood the context window with discovery output, the exact problem the Explore subagent exists to prevent."}, {"letter": "B", "text": "Run the entire task, including the small config field addition, inside plan mode so every step gets reviewed before any file changes happen", "correct": false, "explanation": "Routing the narrow, already-clarified config change through plan mode adds an unnecessary approval cycle to a step that is a simple, well-scoped edit once the investigation is done."}, {"letter": "C", "text": "Delegate the broad investigation to an Explore subagent to keep discovery output out of context, then add the field with direct execution", "correct": true, "explanation": "Isolating the verbose, dozens-of-files investigation in an Explore subagent preserves the main conversation's context, and once the target location is known, the one-file config change is a well-scoped edit suited to direct execution."}, {"letter": "D", "text": "Skip the investigation phase entirely and add the new config field to whichever file in the codebase looks most plausible at a glance", "correct": false, "explanation": "Skipping the investigation risks placing the field in the wrong location within the legacy authentication flow, since the task explicitly notes the correct location isn't yet clear."}], "correct": "C", "select": 1, "group": "C"}, {"id": "f3-062", "domain": 3, "task_id": "3.1", "objective": "Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization", "situation": "Two engineers on the same team report that Claude behaves inconsistently: one gets careful adherence to a 'always run the linter before committing' rule, while the other says Claude never mentions the linter at all, even in the same repository on the same branch.", "question": "Before assuming the rule text is unclear, what is the fastest way to confirm whether the inconsistency is actually a memory-loading problem?", "options": [{"letter": "A", "text": "Have each engineer run /memory and compare which CLAUDE.md, CLAUDE.local.md, and rules files are loaded to see if the linter rule is missing", "correct": true, "explanation": "/memory lists exactly which CLAUDE.md, CLAUDE.local.md, and rules files are loaded in the current session; if the linter instruction is missing from one engineer's list, that immediately isolates the problem to a loading/scoping issue rather than instruction wording."}, {"letter": "B", "text": "Compare the two engineers' Claude Code version numbers, since instruction-following is versioned per release and needs a version match", "correct": false, "explanation": "Version mismatches could cause differences, but comparing versions doesn't reveal what's actually loaded, and instruction-following isn't gated by a required version match for diagnosis."}, {"letter": "C", "text": "Ask both engineers to restart their machines, since CLAUDE.md changes only take effect after a full operating system reboot occurs", "correct": false, "explanation": "CLAUDE.md files are read fresh at the start of each session; no OS reboot is required or relevant to whether an instruction is loaded."}, {"letter": "D", "text": "Have each engineer delete their local git history and re-clone the repository, since stale git objects usually cause instruction drift between machines", "correct": false, "explanation": "Stale git objects do not affect which memory files Claude Code discovers and loads, and re-cloning is a disruptive step with no diagnostic value here."}], "correct": "A", "select": 1, "group": "D"}, {"id": "f3-063", "domain": 3, "task_id": "3.1", "objective": "Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization", "situation": "A platform team maintains a monorepo with packages/api, packages/web, and packages/shared. The api package's maintainer wants api-specific database and testing conventions to load only when someone is actively working in packages/api, without duplicating those conventions into the root CLAUDE.md.", "question": "Which approach best satisfies this requirement?", "options": [{"letter": "A", "text": "Paste the api-specific conventions directly into every developer's ~/.claude/CLAUDE.md so they apply automatically whenever anyone opens the api package", "correct": false, "explanation": "User-level CLAUDE.md is personal, not shared through version control, and would need to be manually duplicated by every developer rather than maintained once by the package owner."}, {"letter": "B", "text": "Reference the api-specific standards file from the root CLAUDE.md using @packages/api/CLAUDE.md so it always loads at launch regardless of working directory", "correct": false, "explanation": "Importing the api file from root would load it into every session regardless of which package is being worked on, defeating the goal of scoping it to api-only work."}, {"letter": "C", "text": "Add @packages/api/CLAUDE.md as an import inside packages/api/CLAUDE.md itself, since imports referencing their own containing file take precedence over root-level rules", "correct": false, "explanation": "A file cannot meaningfully import itself, and this describes a non-existent precedence rule; imports don't create self-referential precedence."}, {"letter": "D", "text": "Create packages/api/CLAUDE.md with the api-specific conventions, so it loads at launch when Claude starts from packages/api or on demand when Claude reads files there", "correct": true, "explanation": "A per-directory CLAUDE.md in packages/api loads at launch when Claude starts there, or lazily when Claude reads files in that subdirectory, keeping it scoped without touching the root file."}], "correct": "D", "select": 1, "group": "D"}, {"id": "f3-064", "domain": 3, "task_id": "3.2", "objective": "Create and configure custom slash commands and skills", "situation": "A developer configured allowed-tools: Write Edit on a skill, expecting this to prevent Claude from ever calling Bash while the skill runs. During a session, Claude still calls Bash after asking for the user's approval.", "question": "Why did this happen, and what should the developer configure instead to fully remove Bash from the available pool while the skill is active?", "options": [{"letter": "A", "text": "allowed-tools requires trailing wildcards, such as Write* and Edit*, to restrict tools; without them, all tools including Bash are implicitly allowed. To block Bash, append * to each allowed tool so only those tools can be called.", "correct": false, "explanation": "allowed-tools does not require trailing wildcards to restrict tools; it already restricts Claude to only those specifically listed tools. Wildcards are used to match tool names by pattern, but omitting wildcards does not implicitly allow all tools. To block Bash, disallowed-tools: Bash must be explicitly set."}, {"letter": "B", "text": "allowed-tools is evaluated only after the skill finishes running, so Bash calls made during the skill are unaffected by its configuration. To remove Bash, set context: fork on the skill to sandbox execution and block unlisted tools.", "correct": false, "explanation": "allowed-tools constrains the tool pool for the entire duration the skill is active, not just after it finishes. context: fork changes where the skill executes but does not block tools; it's not a sandbox that automatically blocks unlisted tools. To fully remove Bash, disallowed-tools should be used."}, {"letter": "C", "text": "allowed-tools only pre-approves the listed tools without prompting; it does not remove other tools from availability. Adding disallowed-tools: Bash removes Bash from Claude's pool while the skill is active.", "correct": true, "explanation": "allowed-tools pre-approves the listed tools so they run without prompting, but it does not remove unlisted tools; they remain callable subject to normal permission settings. To actually remove Bash from Claude's available tools during skill execution, disallowed-tools: Bash must be configured. This is why Bash was still callable with approval."}, {"letter": "D", "text": "allowed-tools only constrains tools invoked directly by the user, and does not limit tools that Claude chooses during skill execution. To remove Bash, set model: inherit on the skill to override Claude's autonomous tool selection.", "correct": false, "explanation": "allowed-tools restricts all tools that Claude can call during skill execution, not just those invoked directly by the user. model: inherit only determines which model handles the request and has no impact on tool availability. To remove Bash, you need disallowed-tools in the skill configuration."}], "correct": "C", "select": 1, "group": "D"}, {"id": "f3-065", "domain": 3, "task_id": "3.6", "objective": "Integrate Claude Code into CI/CD pipelines", "situation": "A build script wants to pipe a large build-error log into Claude for a root-cause explanation: cat build-error.txt | claude -p 'explain the root cause' > output.txt. On one particularly verbose failure, the job exits immediately with an error and no explanation is produced.", "question": "What is the most likely cause given how Claude Code handles piped stdin?", "options": [{"letter": "A", "text": "The redirect operator > is not supported when combined with -p, so the shell discarded the response before Claude could write it", "correct": false, "explanation": "Standard shell output redirection works normally with -p mode; there is no special restriction preventing '>' from capturing Claude Code's output."}, {"letter": "B", "text": "The build log exceeded the 10MB cap on piped stdin, so Claude Code exited with an error instead of processing the oversized input", "correct": true, "explanation": "Piped stdin in Claude Code is capped at 10MB; exceeding it causes the CLI to exit with a clear error and non-zero status rather than processing the content. The documented workaround is to write the content to a file and reference the path in the prompt instead."}, {"letter": "C", "text": "The build log contained non-UTF8 bytes, which -p mode rejects outright regardless of file size", "correct": false, "explanation": "The described stdin limitation in Claude Code's documentation is a size cap, not a character-encoding restriction; an encoding issue is not the documented failure mode here."}, {"letter": "D", "text": "Piped stdin is only accepted when --input-format stream-json is explicitly set, so plain text piping silently failed", "correct": false, "explanation": "Plain text stdin piping is the default and documented pattern for -p mode; --input-format stream-json is only needed for the stream-json input/output pairing, not for ordinary piped text."}], "correct": "B", "select": 1, "group": "D"}, {"id": "f3-066", "domain": 3, "task_id": "3.2", "objective": "Create and configure custom slash commands and skills", "situation": "A contractor clones a repository containing a project skill at .claude/skills/publish/SKILL.md, whose frontmatter sets allowed-tools: Bash(npm publish *). The skill's instructions tell Claude to run npm run build && npm publish to publish the package. On first opening the project in Claude Code, a workspace trust dialog appears, but the contractor dismisses it without accepting. Invoking the skill still prompts for approval before the build-and-publish command finishes running.", "question": "What is the most likely explanation?", "options": [{"letter": "A", "text": "The Bash(npm publish *) pattern matches only the npm publish subcommand; Claude Code checks each subcommand of a compound command independently, so the unmatched npm run build step still requires approval.", "correct": true, "explanation": "Claude Code splits a compound command at its shell operators and evaluates each subcommand against the permission rules independently, so a pattern scoped to one subcommand does not extend approval to the rest of the line. Bash(npm publish *) covers only the publish step, leaving the build step with no matching rule and therefore subject to the ordinary approval prompt."}, {"letter": "B", "text": "The skill's frontmatter is missing a context: fork declaration, and Claude Code only honors allowed-tools grants for skill content it runs inside a forked subagent rather than inline in the main conversation.", "correct": false, "explanation": "The context: fork field controls whether a skill's content becomes the prompt for a new forked subagent instead of running inline; it has no bearing on which Bash patterns allowed-tools pre-approves. The prompt persists because one subcommand of the compound command is not covered by the listed pattern, not because of how the skill executes."}, {"letter": "C", "text": "The allowed-tools field only applies to skills stored in the personal ~/.claude/skills/ directory, so any project-scoped skill checked into a shared repository always requires the contractor's manual approval regardless of trust status.", "correct": false, "explanation": "Claude Code's skills documentation states that a project skill's allowed-tools grant applies whenever the skill is invoked, including in a directory that has never been trusted, so neither the personal-versus-project distinction nor the dismissed trust dialog explains the prompt here. The actual cause is that the compound command's build step falls outside the one Bash pattern the frontmatter pre-approves."}, {"letter": "D", "text": "The npm publish command needs to be listed under a separate arguments frontmatter field instead of allowed-tools before Claude Code treats it as pre-approved for invocations of this particular skill.", "correct": false, "explanation": "allowed-tools is the documented frontmatter field for pre-approving Bash patterns, and no arguments field substitutes for it. Even a correctly written pattern would only ever cover the npm publish subcommand, leaving the npm run build step that precedes it unapproved."}], "correct": "A", "select": 1, "group": "D"}, {"id": "f3-067", "domain": 3, "task_id": "3.1", "objective": "Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization", "situation": "A new engineer joins a team that has used Claude Code for six months. All other teammates report that Claude consistently follows the project's commit message format and test-running conventions, but for the new engineer Claude ignores these conventions entirely, even though they cloned the repository fresh and confirmed that a CLAUDE.md file is present in the repo on GitHub. Upon inspection, however, the file is empty and contains none of the expected conventions.", "question": "What is the most likely root cause?", "options": [{"letter": "A", "text": "The repository's CLAUDE.md exceeds the token limit for brand-new sessions, so it is dropped only for engineers with no prior history.", "correct": false, "explanation": "This does not align with documented behavior. While CLAUDE.md files are recommended to be concise (ideally under 200 lines) to preserve context window space, exceeding a token limit would not cause the file to be silently dropped only for new engineers. If the file is present and non-empty, it is loaded for every session regardless of user history. The scenario explicitly states the file was empty, so token limits are irrelevant."}, {"letter": "B", "text": "The conventions are stored in the existing team members' personal CLAUDE.md files (e.g., ~/.claude/CLAUDE.md), not in the project's committed CLAUDE.md, so the new engineer never loads them.", "correct": true, "explanation": "This is the most likely cause. Per Anthropic's documentation, team-wide conventions must be placed in a project-level CLAUDE.md file (e.g., ./CLAUDE.md or ./.claude/CLAUDE.md) that is committed to the repository. This ensures every developer, including new hires, automatically receives and loads these instructions. User-level ~/.claude/CLAUDE.md files are for personal preferences only and are not shared via Git, so conventions stored there remain invisible to others. Since the committed project file was empty, the new engineer's Claude Code session had no team conventions to follow."}, {"letter": "C", "text": "The new engineer's local git client is silently skipping markdown files during checkout, so CLAUDE.md never lands on disk at all.", "correct": false, "explanation": "This scenario is implausible. Git does not silently skip markdown files by default, and the engineer confirmed the file was present in the repository. No known Git configuration or behavior would cause selective omission of .md files during clone or checkout. The issue is clearly with the content of the file, not its absence."}, {"letter": "D", "text": "Claude Code caches CLAUDE.md content per machine on first run, so the new engineer must trigger a manual cache rebuild locally.", "correct": false, "explanation": "Claude Code does not cache CLAUDE.md file contents in a way that would cause a new user to ignore conventions. The loading of CLAUDE.md files happens dynamically at the start of each session, reading the current file system state. There is no persistent cache that would require a manual rebuild. Official documentation describes the loading as automatic and hierarchical, not dependent on prior runs."}], "correct": "B", "select": 1, "group": "D"}, {"id": "f3-068", "domain": 3, "task_id": "3.2", "objective": "Create and configure custom slash commands and skills", "situation": "A project's .claude/skills/review/SKILL.md is checked into the repo for team-wide use. A senior engineer also keeps a ~/.claude/skills/review/SKILL.md on their laptop for personal projects, unaware it uses the same name.", "question": "When this engineer works in the shared repository and runs /review, which skill actually executes?", "options": [{"letter": "A", "text": "The project skill at .claude/skills/review/SKILL.md, because project-scoped skills always take priority over personal ones with the same name.", "correct": false, "explanation": "This is incorrect. The official skill precedence places personal skills above project-level skills. Therefore, a project skill will not override a personal one; it is the reverse. For team-wide standards, it is recommended to commit project skills to the repository, but they will be superseded if a user has a personal skill with the identical name."}, {"letter": "B", "text": "Neither skill executes, and Claude Code reports a naming conflict error until one of the two files is renamed.", "correct": false, "explanation": "Claude Code silently resolves naming conflicts using its priority order without raising an error. The higher-precedence skill (personal, in this hierarchy) simply overrides the lower one. While using unique, descriptive skill names is recommended to avoid accidental overrides, no error is thrown if a conflict exists."}, {"letter": "C", "text": "The personal skill at ~/.claude/skills/review/SKILL.md, because a skill with the same name at the personal level overrides one from the project level.", "correct": true, "explanation": "According to official Claude Code documentation, skill priority follows a strict hierarchy: Enterprise > Personal (~/.claude/skills/) > Project (.claude/skills/) > Plugins. When a personal and a project skill share the same name, the personal skill takes precedence and is the one that executes. This behavior allows individual users to override team-level skills with their own customizations."}, {"letter": "D", "text": "Both skills execute in sequence, first the personal skill and then the project skill, since Claude Code merges same-named skills rather than choosing one.", "correct": false, "explanation": "Claude Code does not merge skills with identical names or chain their execution. The priority hierarchy is used to select exactly one skill definition, and only that skill runs when invoked. The conflict resolution is based on precedence, not concatenation."}], "correct": "C", "select": 1, "group": "D"}, {"id": "f3-069", "domain": 3, "task_id": "3.1", "objective": "Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization", "situation": "A project's single CLAUDE.md has grown to over 500 lines covering code style, testing, security requirements, and deployment steps, and the team notices Claude adheres to instructions less reliably than before. The architect wants to reorganize this into focused, maintainable files while keeping everything loaded for every session.", "question": "Which restructuring best achieves this?", "options": [{"letter": "A", "text": "Move the entire 500 lines into one .claude/rules/all-conventions.md file so CLAUDE.md itself can be deleted, since rules files replace CLAUDE.md", "correct": false, "explanation": "Consolidating everything into one large rules file just relocates the monolith and does not solve the adherence or maintainability problem the team is trying to fix."}, {"letter": "B", "text": "Convert each topic into a paths-scoped rule matching **/* so every rule loads on every single file read across the whole session", "correct": false, "explanation": "Using a **/* paths glob on every rule makes them conditional on file reads rather than loaded at launch, and offers no organizational benefit over unconditional rules for content meant to apply everywhere."}, {"letter": "C", "text": "Split the content into code-style.md, testing.md, security.md, and deployment.md under .claude/rules/, leaving CLAUDE.md as a short project pointer", "correct": true, "explanation": "Splitting a monolithic CLAUDE.md into topic-specific files under .claude/rules/ (testing.md, security.md, etc.) is the documented pattern; rules without paths frontmatter load at launch with the same priority as CLAUDE.md, keeping everything available while improving maintainability."}, {"letter": "D", "text": "Wrap the 500 lines in @ import syntax pointing at four external gist URLs so the content lives outside the repository and downloads at launch", "correct": false, "explanation": "The @ import syntax references files by path, not arbitrary external URLs, and Claude Code does not fetch remote gists as CLAUDE.md imports."}], "correct": "C", "select": 1, "group": "D"}, {"id": "f3-070", "domain": 3, "task_id": "3.1", "objective": "Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization", "situation": "An architect is designing a configuration strategy for a company rolling out Claude Code to all engineering teams. They need security and compliance instructions that every developer receives on every machine, in every repository, and that individual developers or teams cannot disable through their own settings.", "question": "Which approach satisfies this requirement?", "options": [{"letter": "A", "text": "Populate each developer's personal ~/.claude/CLAUDE.md via a one-time onboarding script, which developers can then edit as needed.", "correct": false, "explanation": "Global user settings in ~/.claude/CLAUDE.md apply to all projects for that user, but developers can freely edit or delete this file at any time. There is no mechanism to lock these settings or prevent a developer from altering them. As a result, this approach does not enforce compliance and cannot ensure that the instructions remain active on every machine."}, {"letter": "B", "text": "Commit a CLAUDE.md file to the root of each repository, as source control ensures every team member receives the same project-level instructions.", "correct": false, "explanation": "While a CLAUDE.md file at the repository root provides project‑level instructions that can be shared via source control, it is not enforceable. In Claude Code’s configuration hierarchy, global user settings (e.g., ~/.claude/CLAUDE.md) take precedence over project settings. An individual developer could simply set conflicting instructions in their global file, effectively disabling the project‑level instructions. Project files can also be deleted or modified, so they cannot guarantee that every developer receives and follows the required security and compliance instructions."}, {"letter": "C", "text": "Place a CLAUDE.local.md at the root of each repository, distributed by an onboarding script and added to .gitignore to prevent accidental commits.", "correct": false, "explanation": "A CLAUDE.local.md file is a local override meant for developer‑specific, temporary, or machine‑specific customizations. It has the lowest precedence in the configuration hierarchy; both project‑level (CLAUDE.md) and global user settings (~/.claude/CLAUDE.md) override it. Even if distributed by a script, developers can easily delete or modify it, and because it is added to .gitignore, there is no source control to maintain consistency. This does not provide the required enforcement or universal application."}, {"letter": "D", "text": "Deploy a managed-settings.json file to the system directory (e.g., /etc/claude-code/ on Linux) using the organization's configuration management system, as it enforces policies that override developer and project settings.", "correct": true, "explanation": "Managed settings in Claude Code are designed to be non‑overridable by any user, project, or local settings. Deploying a managed-settings.json file to the system directory via configuration management ensures these policies are present on every machine. According to Anthropic documentation, this file‑based deployment is a recommended method for organizations that manage system configuration through automation. The managed policy has the highest precedence in the permission hierarchy, meaning it cannot be overridden by developer global settings (~/.claude/CLAUDE.md), project settings (CLAUDE.md), or local overrides (CLAUDE.local.md), thus satisfying the requirement."}], "correct": "D", "select": 1, "group": "D"}, {"id": "f3-071", "domain": 3, "task_id": "3.1", "objective": "Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization", "situation": "An architect wants each developer to have their personal editor and formatting preferences automatically apply across all projects on their local machine, without committing those preferences to any project repository.", "question": "Where should these preferences be configured?", "options": [{"letter": "A", "text": "In ~/.claude/rules/ on the developer's machine, since user-level rules apply to every project and load before project rules do", "correct": true, "explanation": "According to Anthropic's official documentation, ~/.claude/rules/ is the designated directory for user-level rules. Files placed here are personal to the developer, automatically apply to all projects on that machine, and are not committed to any repository. They load before project rules, establishing a baseline that can be overridden by project-specific settings if needed. This is the recommended approach for setting personal defaults like editor and formatting preferences."}, {"letter": "B", "text": "In .claude/rules/ inside every project's repository, duplicated identically across each repo so the preferences travel with the code", "correct": false, "explanation": "Placing preferences in .claude/rules/ inside a project's repository would commit them to version control, violating the requirement not to commit preferences to any project. These project-level rules would also only apply to that specific project, not across all projects a developer works on. User-level rules in ~/.claude/rules/ are the correct mechanism for cross-project personal preferences."}, {"letter": "C", "text": "In CLAUDE.local.md at the root of every project, since local files are the only mechanism applying across a developer's projects", "correct": false, "explanation": "CLAUDE.local.md is a real, documented location: Anthropic's Claude Code documentation lists it under \"Local instructions\" for personal, project-specific preferences that you add to your project's .gitignore so they never reach version control. But its scope is one project at a time - it applies only to the project it lives in, and a separate copy is needed in every other repository. That fails the requirement for preferences to automatically apply across all of a developer's projects without per-project setup. The mechanism documented to apply automatically everywhere is ~/.claude/rules/, which is personal to the developer and loads for every project on the machine."}, {"letter": "D", "text": "In the organization's managed policy CLAUDE.md, since only managed policy files can hold personal, non-project-specific preferences", "correct": false, "explanation": "Managed policy files (e.g., server-managed settings) are intended for organizational governance and enforce company-wide rules that apply uniformly to all developers. They cannot hold personal preferences tailored to individual developers. Personal preferences are meant to be configured at the user level via ~/.claude/rules/, not through managed policies, which override user settings."}], "correct": "A", "select": 1, "group": "D"}, {"id": "f3-072", "domain": 3, "task_id": "3.6", "objective": "Integrate Claude Code into CI/CD pipelines", "situation": "A CI job runs claude -p \"apply the lint fixes\" --permission-mode acceptEdits, but the run aborts partway through when Claude attempts to invoke a network request to fetch an updated dependency list.", "question": "Why did this specific action fail to auto-approve under acceptEdits, even though file edits earlier in the same run went through without prompting?", "options": [{"letter": "A", "text": "Network requests are always blocked outright in print mode, meaning the attempt to fetch an updated dependency list would have been stopped regardless of acceptEdits or any other permission configuration.", "correct": false, "explanation": "Network requests are not universally blocked in print mode; their permissibility depends on the configured permission mode and allow rules. The abort occurred because acceptEdits did not cover this network action, not because print mode blocks all network requests."}, {"letter": "B", "text": "The job needed the --bare flag alongside acceptEdits, since acceptEdits has no effect at all unless bare mode is also enabled; without --bare, the network request to fetch dependency lists required interactive approval and caused the abort.", "correct": false, "explanation": "acceptEdits functions independently of --bare; the --bare flag controls auto-discovery of local configuration files like CLAUDE.md and hooks, not the effect of permission modes. The abort was due to the network request not being covered by acceptEdits, not the absence of --bare."}, {"letter": "C", "text": "acceptEdits only applies to the first tool call in a run, so while it auto-approved earlier file edits, the later network request to fetch dependency lists required interactive approval and caused the abort.", "correct": false, "explanation": "acceptEdits applies to the entire session for the categories of actions it covers, not just the first tool call. The network request failed not because of call order, but because network requests are not included in the actions that acceptEdits auto-approves."}, {"letter": "D", "text": "acceptEdits auto-approves file writes and common filesystem commands like mkdir, mv, and cp, but other shell commands and network requests require an --allowedTools entry or a permissions.allow rule.", "correct": true, "explanation": "acceptEdits auto-approves file writes and common filesystem commands such as mkdir, mv, and cp, but other shell commands and network requests are not covered. This action failed because fetching a dependency list requires network access, which falls outside the scope of acceptEdits and must be explicitly allowed via --allowedTools or a permissions.allow rule."}], "correct": "D", "select": 1, "group": "D"}, {"id": "f3-073", "domain": 3, "task_id": "3.4", "objective": "Determine when to use plan mode vs direct execution", "situation": "A request asks Claude to rename a single internal helper function and update its handful of call sites within one module, where every call site is already visible in the file the developer has open.", "question": "Which characteristic of this task most justifies skipping plan mode?", "options": [{"letter": "A", "text": "Plan mode is technically incapable of ever being used for a rename operation, regardless of the circumstances involved", "correct": false, "explanation": "Plan mode has no restriction against being used for rename operations; it is simply unnecessary here because the scope is already fully known."}, {"letter": "B", "text": "Renaming operations never require any verification of call sites at all, so the risk of this kind of change is always assumed zero", "correct": false, "explanation": "Renames can still miss call sites or affect external consumers; the reason plan mode is unnecessary here is the visible, confined scope, not an inherent zero risk in all renames."}, {"letter": "C", "text": "The helper function was written recently, so its current behavior is assumed to already be correct without any review", "correct": false, "explanation": "The recency of the helper function's authorship is unrelated to whether plan mode is needed for this rename; scope and ambiguity are the relevant factors."}, {"letter": "D", "text": "The full scope of the change is already known and confined to one module, leaving no exploration or design decision to make", "correct": true, "explanation": "When the scope is already fully known and confined to one module with no ambiguity about approach, there is no exploration or architectural decision for plan mode to add value on, so direct execution is appropriate."}], "correct": "D", "select": 1, "group": "C"}, {"id": "f3-074", "domain": 3, "task_id": "3.2", "objective": "Create and configure custom slash commands and skills", "situation": "A developer creates a migration-plan skill with the body: 'Follow the approach we discussed for moving the billing service.' When invoked as a subagent, it returns a generic, unhelpful plan.", "question": "What is the most likely reason?", "options": [{"letter": "A", "text": "The skill is missing an arguments field, so the system ignores the entire skill body and substitutes a placeholder prompt.", "correct": false, "explanation": "Anthropic's subagent implementation does not silently discard the skill body when arguments are omitted. The body serves as the instruction prompt; missing arguments might lead to unparameterized behavior but not a wholesale replacement with a placeholder."}, {"letter": "B", "text": "The model cannot write multi-step plans, so it defaults to a short generic response regardless of the prompt.", "correct": false, "explanation": "There is no evidence that language models used in subagents are inherently unable to produce multi-step plans. With proper context, Anthropic's models can generate detailed, structured plans. The failure is due to missing context, not an inability to plan."}, {"letter": "C", "text": "Subagents cannot return text results to the main conversation, so the developer only sees an error message instead of a plan.", "correct": false, "explanation": "Subagents are designed to return condensed results, typically as text summaries, back to the main agent. Receiving an error message is not standard behavior and would indicate a different issue, such as a misconfiguration or runtime error."}, {"letter": "D", "text": "The subagent does not have access to the main conversation's history, so the reference to an earlier discussion provides no actionable detail.", "correct": true, "explanation": "Standard subagents operate with isolated context windows and do not inherit history from the parent conversation. A directive like 'Follow the approach we discussed' is meaningless without that context. Official guidance recommends explicitly summarizing relevant context in the subagent’s prompt (see Anthropic’s sub-agent architecture documentation)."}], "correct": "D", "select": 1, "group": "D"}, {"id": "f3-075", "domain": 3, "task_id": "3.3", "objective": "Apply path-specific rules for conditional convention loading", "situation": "A design-system team scoped a rule to paths: [\"src/components/*.tsx\"] to enforce prop-naming conventions. After reorganizing, components now live in nested subfolders like src/components/forms/Input.tsx and src/components/layout/Grid.tsx. The team notices the rule no longer applies to these files.", "question": "What is the cause, and how should they fix it?", "options": [{"letter": "A", "text": "The pattern is unaffected by folder depth, so the rule should still apply; the real cause is that .tsx files require a separate paths entry using brace expansion syntax to be recognized at all.", "correct": false, "explanation": "Brace expansion is for combining multiple extensions in one pattern and is unrelated to whether nested folders are matched; the single-star pattern is the actual cause of the mismatch."}, {"letter": "B", "text": "The rule only fails because the frontmatter is missing a leading forward slash before src/components, and adding one would restore matching for both top-level and nested files.", "correct": false, "explanation": "A leading slash is not required or meaningful for these relative glob patterns, and adding one would not change whether nested subfolders are matched by a single-star pattern."}, {"letter": "C", "text": "Path-scoped rules stop matching automatically once a project exceeds a certain number of files, so the team needs to split components into multiple smaller rules files regardless of the glob pattern used.", "correct": false, "explanation": "There is no documented file-count threshold that disables path-scoped rule matching; the mismatch is explained by the glob pattern's depth, not the project's size."}, {"letter": "D", "text": "The single-level pattern src/components/*.tsx only matches files directly inside src/components/, not nested subfolders; changing it to src/components/**/*.tsx would match files at any depth underneath that directory.", "correct": true, "explanation": "A single star matches only one path segment, so files nested further inside src/components/ are not matched; a double-star wildcard is needed to match at any depth."}], "correct": "D", "select": 1, "group": "D"}, {"id": "f3-076", "domain": 3, "task_id": "3.3", "objective": "Apply path-specific rules for conditional convention loading", "situation": "A team places .claude/rules/general-style.md in the repo without adding a paths field to its YAML frontmatter, alongside a separate .claude/rules/api.md that does declare paths: [\"src/api/**/*.ts\"].", "question": "How will Claude Code treat general-style.md compared to api.md?", "options": [{"letter": "A", "text": "general-style.md loads only when Claude opens a file matching a default wildcard of \"**/*\", functioning identically to api.md but with a broader glob pattern applied automatically.", "correct": false, "explanation": "There is no implicit default wildcard pattern applied to rules missing a paths field; omitting paths simply makes the rule load unconditionally at launch."}, {"letter": "B", "text": "general-style.md and api.md both load only when Claude edits a file located inside the .claude/rules/ directory itself, since rules files are scoped to their own containing directory by default.", "correct": false, "explanation": "Rules apply to files across the project based on their paths patterns (or unconditionally if none is set); they are not scoped to files inside the .claude/rules/ directory itself."}, {"letter": "C", "text": "general-style.md is ignored entirely because every file placed in .claude/rules/ requires a paths field to be recognized before Claude Code will load it, while api.md loads correctly at launch as configured.", "correct": false, "explanation": "A paths field is optional; rules without one are simply treated as unconditional and load at launch rather than being ignored."}, {"letter": "D", "text": "general-style.md loads at launch with the same priority as .claude/CLAUDE.md, since omitting the paths field makes the rule unconditional, while api.md only loads when Claude reads a matching file.", "correct": true, "explanation": "Rules without a paths field are loaded unconditionally, with the same priority as .claude/CLAUDE.md, while rules with a paths field only load when a matching file is read."}], "correct": "D", "select": 1, "group": "D"}, {"id": "f3-077", "domain": 3, "task_id": "3.4", "objective": "Determine when to use plan mode vs direct execution", "situation": "An architect is evaluating whether a requested change to a billing module needs plan mode. The change touches one file, has a single obvious implementation matching an existing pattern used elsewhere in the same file, and does not affect any other module.", "question": "Which factor most strongly indicates that direct execution, rather than plan mode, is the right choice here?", "options": [{"letter": "A", "text": "The existing pattern being followed in the file was written by a different engineer than the one requesting this change", "correct": false, "explanation": "Who originally authored the existing pattern has no bearing on whether the current change is architecturally complex or spans multiple files."}, {"letter": "B", "text": "The billing module is business-critical, so every change to it should always go through the exact same fixed review process", "correct": false, "explanation": "Business criticality alone does not determine whether plan mode is needed; the deciding factors are scope, ambiguity of approach, and architectural impact, not the module's importance."}, {"letter": "C", "text": "The change can be described to Claude in a single short sentence, regardless of how many modules it actually ends up touching", "correct": false, "explanation": "Sentence length describing a change says nothing about its actual scope; a one-sentence request can still imply a large, multi-file, architecturally significant change."}, {"letter": "D", "text": "The change is confined to a single file, has one clear implementation path, and touches no other module in the codebase", "correct": true, "explanation": "Plan mode is reserved for multi-file changes, multiple valid approaches, or architectural decisions; a single-file change with one clear path and no cross-module impact lacks all three, so direct execution fits."}], "correct": "D", "select": 1, "group": "C"}, {"id": "f3-078", "domain": 3, "task_id": "3.5", "objective": "Apply iterative refinement techniques for progressive improvement", "situation": "A team lead wants Claude Code to implement a CSV-to-JSON conversion utility whose column-to-field mapping rules are difficult to describe precisely in words, since the exact handling of empty cells, quoted commas, and duplicate headers matters.", "question": "Which approach best sets up an effective iterative refinement loop before implementation begins?", "options": [{"letter": "A", "text": "Give Claude a few sample input rows paired with the exact expected JSON output including one row with an empty cell and one with a quoted comma, before asking it to write the converter.", "correct": true, "explanation": "Providing sample rows with exact expected JSON outputs, including edge cases like empty cells and quoted commas, gives Claude a precise, unambiguous specification to implement from the start. This avoids the ambiguity that purely verbal descriptions would introduce, setting up a reliable iterative refinement loop."}, {"letter": "B", "text": "Ask Claude to write the full converter first, without providing any example rows, then describe the handling of empty cells and quoted commas verbally once the initial output on a small CSV file shows unexpected results.", "correct": false, "explanation": "Writing the full converter first without examples and then verbally describing edge-case handling after seeing unexpected results leads to unnecessary rework and confusion. Verbal descriptions of CSV parsing nuances are inherently ambiguous, making this approach less effective for iterative refinement."}, {"letter": "C", "text": "Instruct Claude to select the most popular CSV parsing library from GitHub, implement the converter using that library, and rely on its default behaviors for handling empty cells and quoted commas.", "correct": false, "explanation": "Selecting a popular library and relying on its default behaviors does not guarantee that the handling of empty cells and quoted commas will match the team's specific requirements. This approach assumes the library's defaults are correct without verification, undermining the refinement process."}, {"letter": "D", "text": "Tell Claude to implement the converter, adding TODO comments for any edge cases like empty cells or quoted commas it cannot resolve from the prompt, then proceed to the next task without further iteration.", "correct": false, "explanation": "Adding TODO comments for unresolved edge cases and moving on without further iteration defers ambiguity rather than resolving it. This leaves the implementation incomplete and risks inconsistent behavior when those edge cases are eventually addressed."}], "correct": "A", "select": 1, "group": "I"}, {"id": "f3-079", "domain": 3, "task_id": "3.1", "objective": "Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization", "situation": "An architect is deciding whether a new set of database-migration conventions should live in a per-directory CLAUDE.md inside src/db/ or as a path-scoped rule under the repository root's .claude/rules/ with paths: [\"**/migrations/**\"]. The migration conventions need to apply to migration files scattered across several unrelated subsystems, not just one directory.", "question": "Which structure fits this requirement better, and why?", "options": [{"letter": "A", "text": "A per-directory CLAUDE.md in src/db/, because path-scoped rules can only match a single exact directory, never recursive glob wildcards", "correct": false, "explanation": "Path-scoped rules do support recursive glob wildcards such as **/migrations/**, so this claim about their limitation is incorrect."}, {"letter": "B", "text": "A path-scoped rule under .claude/rules/ with the migrations glob, since it targets files by pattern tree-wide rather than one directory", "correct": true, "explanation": "Path-scoped rules match by glob pattern across the whole tree, so a single rule with a **/migrations/** pattern applies wherever migration files exist, regardless of which subsystem they're under, which is exactly suited to conventions tied to a file pattern rather than a single directory."}, {"letter": "C", "text": "A per-directory CLAUDE.md in src/db/, because directory files always take precedence over rules no matter how scattered the matches are", "correct": false, "explanation": "A CLAUDE.md placed in src/db/ only loads for that directory and its own subtree; it would not reach migration files located in other, unrelated subsystems, and there's no precedence rule that makes directory files universally win over path-scoped rules."}, {"letter": "D", "text": "Neither approach works for scattered files; the only option is duplicating the same CLAUDE.md manually into every directory with migrations", "correct": false, "explanation": "Manual duplication across every directory is unnecessary and is precisely the maintenance burden that path-scoped rules with glob patterns are designed to avoid."}], "correct": "B", "select": 1, "group": "D"}, {"id": "f3-080", "domain": 3, "task_id": "3.5", "objective": "Apply iterative refinement techniques for progressive improvement", "situation": "An architect wants Claude Code to build a function that validates and reformats postal addresses, but plain-English descriptions of the expected reformatting keep producing inconsistent results across a handful of edge cases (missing unit numbers, PO boxes, and rural route addresses).", "question": "Which action best resolves this without adding unnecessary process overhead?", "options": [{"letter": "A", "text": "Ask Claude to implement three separate reformatting functions, one for each edge case (missing unit numbers, PO boxes, and rural routes), and let the caller decide which function to invoke.", "correct": false, "explanation": "Creating separate functions for each edge case introduces unnecessary code complexity; the core issue is clarifying the expected output for ambiguous inputs, not requiring distinct code paths."}, {"letter": "B", "text": "Supply a small set of concrete input/output examples, one per edge case (missing unit number, PO box, rural route) so each ambiguous case has an unambiguous target output.", "correct": true, "explanation": "Providing concrete input/output examples for each edge case directly clarifies the expected reformatting behavior, eliminating ambiguity without adding overhead like meetings or extra code."}, {"letter": "C", "text": "Schedule a full team design review meeting to formally document postal address reformatting standards, including edge cases such as missing unit numbers and PO boxes, before continuing with Claude Code.", "correct": false, "explanation": "A full team design review meeting is a heavy process for resolving a few specific edge cases; a handful of examples can directly communicate the desired output without formal documentation."}, {"letter": "D", "text": "Increase the temperature or creativity setting used by Claude Code so that it explores a wider variety of reformatting outputs for edge cases such as missing unit numbers, PO boxes, and rural routes.", "correct": false, "explanation": "Raising the temperature would increase output randomness, making reformatting even less consistent across edge cases, contrary to the goal of achieving consistent results."}], "correct": "B", "select": 1, "group": "I"}, {"id": "f3-081", "domain": 3, "task_id": "3.1", "objective": "Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization", "situation": "An engineer wants to keep their personal sandbox database URL and preferred test fixtures available to Claude Code whenever they work in a specific repository, without ever committing these personal details to the shared repository.", "question": "Which configuration best achieves this?", "options": [{"letter": "A", "text": "Add the sandbox URL and fixtures to a CLAUDE.local.md file at the project root and add CLAUDE.local.md to the repository's .gitignore (or your global Git excludes) so it remains uncommitted.", "correct": true, "explanation": "According to Anthropic's Claude Code memory documentation, CLAUDE.local.md is the documented file for private per-project preferences that should not be checked into version control. The docs explicitly list use cases such as \"Your sandbox URLs, preferred test data\", making this the recommended approach. Adding CLAUDE.local.md to .gitignore or using global Git excludes ensures the file is not committed while remaining available to Claude Code in that repository."}, {"letter": "B", "text": "Store the sandbox URL and fixtures in a .claude/settings.local.json file at the project root. Claude Code automatically adds this file to the global Git excludes, so it is never committed.", "correct": false, "explanation": ".claude/settings.local.json is designed for structured, machine-local project settings (such as permissions, environment variables, MCP servers, and sandbox filesystem/network access) and has a fixed schema. It is not the documented place to store arbitrary personal context like sandbox database URLs and preferred test fixture names. Although Claude Code automatically excludes this file from Git, the appropriate file for free-form personal notes is CLAUDE.local.md."}, {"letter": "C", "text": "Append the sandbox URL and fixtures to the bottom of the shared project CLAUDE.md, since Claude Code automatically strips personal values before committing.", "correct": false, "explanation": "The shared project CLAUDE.md is intended to be committed to source control as persistent, team-wide context for Claude Code. Claude Code does not automatically strip personal values from committed files, so adding a sandbox URL and fixtures there would likely expose them to the repository. Personal secrets should instead live in an uncommitted local file like CLAUDE.local.md."}, {"letter": "D", "text": "Store the sandbox URL and fixtures in ~/.claude/CLAUDE.md, because user-level configuration files are automatically scoped to the current repository.", "correct": false, "explanation": "~/.claude/CLAUDE.md is a user-level context file that applies globally across all projects, not automatically scoped to the current repository. Storing repository-specific personal details there could apply them in the wrong context or expose them unintentionally. The documented mechanism for per-project, uncommitted personal preferences is CLAUDE.local.md at the project root."}], "correct": "A", "select": 1, "group": "D"}, {"id": "f3-082", "domain": 3, "task_id": "3.4", "objective": "Determine when to use plan mode vs direct execution", "situation": "During a multi-phase refactor, Claude needs to grep across hundreds of files to catalog every usage of a deprecated helper function before deciding what to do next. The architect is worried this discovery output will consume most of the main conversation's context window.", "question": "What should the architect do to prevent this?", "options": [{"letter": "A", "text": "Delegate the cataloging work to an Explore subagent, which runs the searches in its own context and returns a condensed summary to the conversation", "correct": true, "explanation": "The Explore subagent isolates verbose discovery work like broad greps in its own context window and returns only the summary, preserving the main conversation's context for the rest of the multi-phase task."}, {"letter": "B", "text": "Have Claude paste the full grep output directly into the conversation and then manually delete the least relevant lines from the transcript afterward", "correct": false, "explanation": "Pasting the full output into the main conversation first consumes the context budget the architect is trying to protect, even if some lines are removed afterward."}, {"letter": "C", "text": "Raise the configured context window limit for the session so the raw search results no longer need to be trimmed at all before use", "correct": false, "explanation": "There is no setting that expands a model's context window on demand; the discovery output still needs to be isolated rather than resized."}, {"letter": "D", "text": "Switch the main session into plan mode, which is specifically designed to automatically summarize all grep output before it reaches the conversation", "correct": false, "explanation": "Plan mode restricts Claude to read-only actions before edits are approved; it does not isolate or summarize search output the way a dedicated subagent does."}], "correct": "A", "select": 1, "group": "C"}, {"id": "f3-083", "domain": 3, "task_id": "3.3", "objective": "Apply path-specific rules for conditional convention loading", "situation": "A developer has both a personal rule at ~/.claude/rules/formatting.md (no paths field) and their team's project rule at ./.claude/rules/formatting.md (no paths field) with conflicting formatting guidance. Both load unconditionally.", "question": "According to Claude Code memory documentation, how should the developer understand load order and conflict resolution?", "options": [{"letter": "A", "text": "User-level rules load before project rules, so when the two conflict, the project rule wins. Claude gives it priority for loading later, not for being more specific.", "correct": true, "explanation": "Claude Code's memory documentation states, in the User-level rules section: 'User-level rules are loaded before project rules, giving project rules higher priority.' The personal ~/.claude/rules/formatting.md loads first, and when it conflicts with the project's ./.claude/rules/formatting.md, the project rule is treated as higher priority simply because it entered context after the personal one -- not because Claude judged it more specific or more important."}, {"letter": "B", "text": "User-level rules under ~/.claude/rules/ are ignored whenever a project defines a rules file with the same name, so only the project's formatting.md file ever loads.", "correct": false, "explanation": "Neither file is ignored. Both rules load unconditionally because neither carries a paths field restricting it, so the personal ~/.claude/rules/formatting.md and the project's ./.claude/rules/formatting.md both enter context; a same-named project file does not suppress the personal one."}, {"letter": "C", "text": "User-level rules load before project rules, and project rules win any conflict, but only because project-level guidance is inherently more specific than personal guidance.", "correct": false, "explanation": "The load order and the outcome are both right here -- project rules do win -- but the stated reason is not documented anywhere. The memory documentation attributes the project rule's priority to load order, since it loads after the user-level rule, not to project-level guidance being inherently more specific. Assuming a specificity rule would incorrectly suggest a narrower personal rule could ever outrank a broader project one."}, {"letter": "D", "text": "Project rules load before user-level rules, so the personal formatting.md file takes higher priority over the team's shared project guidance.", "correct": false, "explanation": "The documented load order is the reverse of this: user-level rules load before project rules. Reversing the order also reverses the outcome -- it is the project's shared file, loaded second, that ends up with the higher-priority position, not the personal one."}], "correct": "A", "select": 1, "group": "D"}, {"id": "f3-084", "domain": 3, "task_id": "3.5", "objective": "Apply iterative refinement techniques for progressive improvement", "situation": "A backend engineer runs a test suite against a new Claude Code implementation of an inventory reconciliation function and gets three failing tests: a rounding error on fractional quantities, an off-by-one error in a loop boundary, and a missing null check for a discontinued-item field. All three appear to stem from unrelated parts of the function.", "question": "What is the most effective way to communicate these results to Claude for the next iteration?", "options": [{"letter": "A", "text": "Report the null check issue first because it is a runtime crash risk, and postpone mentioning the other two failures until a later session.", "correct": false, "explanation": "Anthropic's guidance is to give Claude the symptom, the likely location and what fixed looks like, not to ration failures by perceived severity. Postponing the rounding and loop-boundary failures to a later session also discards the context of this one, and the documented advice is the opposite: paste the error, fix it, and verify the build or suite succeeds. Two known defects would ship unfixed for no documented benefit."}, {"letter": "B", "text": "Ask Claude to guess which of the three failures is most likely to be the underlying root cause before revealing any of the specific test output.", "correct": false, "explanation": "Asking Claude to guess withholds exactly the evidence the documentation says to supply - the best-practices guide replaces \"fix the login bug\" with a prompt naming the symptom, the likely location and what fixed looks like, and suggests piping data in with \"cat error.log | claude\". Claude can read the three failures directly from the test output, so a guess made before seeing them is unverifiable and wastes a turn."}, {"letter": "C", "text": "Report each test failure sequentially in separate messages, addressing one issue at a time and verifying the fix before moving to the next.", "correct": false, "explanation": "Splitting one suite's failures into separate messages is not a documented Claude Code practice. The guidance about unrelated work is about session context - \"/clear: reset context between unrelated tasks\" - which concerns separate tasks, not separate assertions failing in the same function, and the same guide shows a single prompt covering every failure at once, as in \"fix all lint errors\". Feeding the failures one at a time withholds context Claude already needs to read the function correctly and adds a round trip per fix."}, {"letter": "D", "text": "Share the full test output for all three failures in one message, let Claude propose fixes, and verify by rerunning the entire test suite.", "correct": true, "explanation": "This matches the workflow Anthropic documents for Claude Code: give Claude the failing output and a check it can run. The best-practices guide's implementation prompt is \"write tests for the callback handler, run the test suite and fix any failures\", and its root-cause example pastes the error itself - \"the build fails with this error: [paste error]. fix it and verify the build succeeds. address the root cause, don't suppress the error\". Handing over all three failures at once lets Claude see the whole reconciliation function while reasoning about the rounding, boundary and null-check defects, and rerunning the full suite proves that no fix broke another test."}], "correct": "D", "select": 1, "group": "I"}, {"id": "f3-085", "domain": 3, "task_id": "3.4", "objective": "Determine when to use plan mode vs direct execution", "situation": "A stack trace points to a null check missing in a single function inside one file, and the fix is a one-line conditional.", "question": "Which workflow best matches the scope of this change?", "options": [{"letter": "A", "text": "Split the one-line fix into a multi-phase plan with a separate research session and a separate implementation session afterward", "correct": false, "explanation": "Breaking a one-line, already-diagnosed fix into multiple phases adds unnecessary process overhead for a change this small."}, {"letter": "B", "text": "Delegate the fix to an Explore subagent so the discovery work stays fully isolated from the rest of the main conversation entirely", "correct": false, "explanation": "The Explore subagent isolates verbose discovery output for large investigations; there is no substantial discovery phase here since the cause is already known from the trace."}, {"letter": "C", "text": "Direct execution, since the change is a well-scoped, single-file fix with a clear cause already identified from the trace", "correct": true, "explanation": "A single-file fix with a clear, already-identified cause is a well-understood, well-scoped change, which is exactly what direct execution is suited for."}, {"letter": "D", "text": "Plan mode, since any production bug fix always warrants a written proposal and a full approval cycle no matter how small the change is", "correct": false, "explanation": "Plan mode is for tasks with multiple valid approaches or architectural implications; forcing every bug fix through it adds friction without benefit when the fix and cause are already clear."}], "correct": "C", "select": 1, "group": "C"}, {"id": "f3-086", "domain": 3, "task_id": "3.3", "objective": "Apply path-specific rules for conditional convention loading", "situation": "A platform engineer maintains a monorepo where test files (matching *.test.ts) are scattered across src/, lib/, packages/*/tests/, and e2e/ directories. The engineer wants testing conventions to load automatically whenever Claude works on any test file, without duplicating the same instructions across multiple subdirectory CLAUDE.md files.", "question": "Which approach best achieves this?", "options": [{"letter": "A", "text": "Configure claudeMdExcludes in project settings to point at each test directory, which causes Claude to load the testing conventions only when those specific paths are excluded from the default context.", "correct": false, "explanation": "claudeMdExcludes is used to skip loading specific CLAUDE.md or rules files by path; it does not cause conditional loading of guidance and does not scope a rule to a file-type pattern."}, {"letter": "B", "text": "Add a CLAUDE.md file to each subdirectory that contains test files, duplicating the testing conventions text in every location so Claude picks up the guidance regardless of which directory it is currently exploring.", "correct": false, "explanation": "Duplicating a CLAUDE.md per directory means every new test directory needs its own copy, and edits must be kept in sync manually across all copies, which is exactly the maintenance burden path-specific rules avoid."}, {"letter": "C", "text": "Create a single .claude/rules/testing.md file with paths: [\"**/*.test.ts\"] in its YAML frontmatter, so the rule loads whenever Claude reads a file matching that pattern anywhere in the repo.", "correct": true, "explanation": "A single rules file scoped with a recursive glob pattern in the paths frontmatter applies by file type regardless of directory, and loads only when a matching file is opened, avoiding duplication and reducing irrelevant context."}, {"letter": "D", "text": "Move the testing conventions into the root-level CLAUDE.md file so they are loaded into context at the start of every session, ensuring Claude always has the guidance available before touching any file.", "correct": false, "explanation": "Placing the conventions in the root CLAUDE.md loads them into every session regardless of whether Claude is touching test files, consuming context on unrelated work instead of loading conditionally."}], "correct": "C", "select": 1, "group": "D"}, {"id": "f3-087", "domain": 3, "task_id": "3.4", "objective": "Determine when to use plan mode vs direct execution", "situation": "An architect is migrating a project from one ORM library to another across 60 files, where call sites use inconsistent query patterns and the correct replacement idiom differs by pattern. The architect wants Claude to first survey how the old library is used before writing any replacement code.", "question": "What is the most effective way to structure this work?", "options": [{"letter": "A", "text": "Skip the pattern survey and rewrite every file in one direct-execution pass, relying on the test suite to catch mismatched replacements later", "correct": false, "explanation": "Skipping the survey step for a migration with inconsistent patterns across many files is likely to produce mismatched replacements that require costly correction later."}, {"letter": "B", "text": "Apply the replacement idiom seen in the first converted file to every remaining file directly, without ever checking whether other query patterns exist", "correct": false, "explanation": "Assuming one idiom applies everywhere ignores the stated inconsistency in call site patterns and will misapply the wrong replacement to many files."}, {"letter": "C", "text": "Use plan mode to survey the inconsistent call site patterns and draft a migration approach, then exit plan mode and execute the approved plan directly", "correct": true, "explanation": "Plan mode fits the investigation phase across many files with varying patterns, and once the migration approach is approved, switching to direct execution carries it out efficiently, matching the combined plan-then-execute pattern for library migrations."}, {"letter": "D", "text": "Enable acceptEdits mode so Claude can begin swapping calls file by file, reconciling newly discovered query patterns as extraction proceeds", "correct": false, "explanation": "Starting in acceptEdits lets edits happen before the varying call-site patterns across 60 files are fully understood, risking inconsistent replacements that need rework."}], "correct": "C", "select": 1, "group": "C"}, {"id": "f3-088", "domain": 3, "task_id": "3.1", "objective": "Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization", "situation": "A team wants a rule requiring input validation and a standard error response format to apply automatically whenever Claude works on any file under src/api/, but they don't want this rule to consume context space during frontend or infrastructure work.", "question": "Which .claude/rules/ configuration satisfies this?", "options": [{"letter": "A", "text": "A rules file with no frontmatter, placed inside src/api/.claude/rules/ so its location alone scopes it to that one directory", "correct": false, "explanation": "Rules are discovered under .claude/rules/ at the project level; placing a rules directory inside src/api isn't the documented mechanism, and a rule without paths frontmatter loads unconditionally rather than being scoped by its folder location."}, {"letter": "B", "text": "A rules file with YAML frontmatter paths: [\"src/api/**/*\"], so it only enters context when Claude reads a matching file", "correct": true, "explanation": "The paths frontmatter field takes glob patterns and makes the rule conditional, so it only loads into context when Claude reads a file matching src/api/**/*, keeping it out of context for unrelated work."}, {"letter": "C", "text": "A rules file named src-api-only.md, since Claude Code infers scoping from filename prefixes matching directory names", "correct": false, "explanation": "Claude Code does not infer path scoping from filename conventions; scoping is controlled explicitly through the paths frontmatter field."}, {"letter": "D", "text": "A rules file with paths set to the literal string \"src/api\" with no glob wildcard, treated as matching every file underneath recursively", "correct": false, "explanation": "A bare literal path string without glob wildcards matches only that exact literal path, not files recursively beneath it, so it would not match individual files under src/api."}], "correct": "B", "select": 1, "group": "D"}, {"id": "f3-089", "domain": 3, "task_id": "3.3", "objective": "Apply path-specific rules for conditional convention loading", "situation": "A developer configured .claude/rules/api-security.md scoped with paths: [\"src/api/**/*.ts\"]. During a session, Claude runs git status and lists the repository tree with Glob, but has not yet opened any file under src/api/.", "question": "Based on how path-scoped rules are triggered, what should the developer expect?", "options": [{"letter": "A", "text": "The api-security.md rule will never load during this session unless the developer explicitly runs the /memory command to force path-scoped rules to activate.", "correct": false, "explanation": "/memory is used to view and browse loaded instruction files, not to force path-scoped rules to activate; the rule loads automatically once a matching file is read."}, {"letter": "B", "text": "The api-security.md rule has not been loaded into context yet, because path-scoped rules load when Claude reads a file matching the pattern, not merely when it uses other tools like Glob or Bash.", "correct": true, "explanation": "Path-scoped rules trigger when Claude reads a file matching the pattern, not on every tool use, so running git status or listing files with Glob does not load the rule."}, {"letter": "C", "text": "The api-security.md rule loaded as soon as Glob returned a listing that included files under src/api/, because any tool invocation surfacing a matching path counts as a trigger.", "correct": false, "explanation": "Seeing a matching filename in a directory listing is not the same as reading the file's contents; the trigger is Claude reading a file matching the pattern, not any tool surfacing the path."}, {"letter": "D", "text": "The api-security.md rule loaded automatically the moment the session started, because all rules under .claude/rules/ are loaded unconditionally regardless of their paths frontmatter.", "correct": false, "explanation": "Rules with a paths field are conditional, not unconditional; only rules omitting the paths field load automatically at launch."}], "correct": "B", "select": 1, "group": "D"}, {"id": "f3-090", "domain": 3, "task_id": "3.2", "objective": "Create and configure custom slash commands and skills", "situation": "A team's explore-alternatives skill asks Claude to brainstorm several competing architecture approaches before settling on a recommendation. The team wants that exploratory back-and-forth kept separate from the main conversation, with only the final recommendation surfacing back to the user.", "question": "Which configuration accomplishes this?", "options": [{"letter": "A", "text": "Set context: fork on the skill so the brainstorming happens in a subagent, and only the returned result reaches the main conversation.", "correct": true, "explanation": "context: fork isolates the exploratory reasoning in a subagent context, so the brainstorming itself never enters the main conversation, and only the resulting recommendation is surfaced back."}, {"letter": "B", "text": "Set paths: **/* on the skill so brainstorming activates automatically no matter which file is being edited.", "correct": false, "explanation": "paths only controls when a skill auto-activates based on file patterns; it has no effect on whether the skill's output is isolated from the main conversation."}, {"letter": "C", "text": "Set disable-model-invocation: true on the skill so brainstorming only happens when a teammate manually types the command.", "correct": false, "explanation": "This only restricts who can trigger the skill; it does nothing to keep the brainstorming content out of the main conversation once invoked."}, {"letter": "D", "text": "Set user-invocable: false on the skill so Claude silently runs the brainstorming in the background without a command name.", "correct": false, "explanation": "This only hides the skill from the / menu; the skill would still run inline in the main conversation, exposing the full brainstorming."}], "correct": "A", "select": 1, "group": "D"}, {"id": "f3-091", "domain": 3, "task_id": "3.1", "objective": "Configure CLAUDE.md files with appropriate hierarchy, scoping, and modular organization", "situation": "A maintainer writes a project CLAUDE.md that includes the line See @README for project overview and @package.json for available npm commands. They also want to mention the file @docs/legacy-notes.md in a sentence purely as a pointer for humans, without Claude actually importing and loading its contents into context.", "question": "How should they write that second reference?", "options": [{"letter": "A", "text": "Wrap the path in backticks so it becomes a Markdown code span, since import parsing skips code spans and treats the reference as literal text", "correct": true, "explanation": "Claude Code's memory documentation states that import parsing skips Markdown code spans and fenced code blocks, and that to mention a path without importing it you wrap it in backticks, which keeps the text literal while the same path outside backticks imports the file. A code span therefore leaves the reference visible to human readers of the file and out of the import expansion."}, {"letter": "B", "text": "Put a space between the @ and the path like @ docs/legacy-notes.md, since only paths immediately adjacent to @ with no space are treated as imports", "correct": false, "explanation": "The memory documentation defines no whitespace rule for suppressing an import; the escape it specifies is a Markdown code span. Separating the at-sign from the path also stops the text being a file reference at all, so the sentence no longer points a human reader at the file under discussion."}, {"letter": "C", "text": "Prefix it with a double @@docs/legacy-notes.md, since a doubled @ symbol tells the parser to render the path without importing it", "correct": false, "explanation": "The memory documentation describes exactly one way to keep a path literal in a CLAUDE.md file — wrapping it in backticks so it sits inside a Markdown code span — and never mentions a doubled at-sign. Nothing in the documented import syntax gives @@ an escaping meaning, so this form is not a supported way to suppress the import."}, {"letter": "D", "text": "Place it inside an HTML comment <!-- @docs/legacy-notes.md -->, since HTML comments are rendered as plain visible text but excluded from imports", "correct": false, "explanation": "HTML comments are not visible in rendered Markdown, so they cannot serve as the visible pointer the maintainer wants. The memory documentation also states that block-level HTML comments in CLAUDE.md files are stripped before the content is injected into Claude's context, which hides the note entirely rather than displaying it while skipping the import."}], "correct": "A", "select": 1, "group": "D"}, {"id": "f3-092", "domain": 3, "task_id": "3.4", "objective": "Determine when to use plan mode vs direct execution", "situation": "A team is choosing between two integration approaches for a new payment provider: one requires a new message queue and asynchronous worker fleet, the other requires synchronous calls from existing services with added retry logic. Each has different infrastructure and operational tradeoffs, and the team has not yet decided which to pursue.", "question": "What is the most appropriate way to proceed with Claude Code?", "options": [{"letter": "A", "text": "Have Claude implement both approaches in parallel branches with direct execution and compare the resulting pull requests once both are finished", "correct": false, "explanation": "Building out two full infrastructure approaches directly is far more costly than evaluating tradeoffs first in plan mode, and defeats the purpose of preventing costly rework."}, {"letter": "B", "text": "Use plan mode to have Claude explore both approaches and surface the infrastructure tradeoffs before any implementation starts", "correct": true, "explanation": "Choosing between integration approaches with different infrastructure requirements is an architectural decision, and plan mode lets Claude investigate and propose a design safely before either approach is committed to code."}, {"letter": "C", "text": "Skip codebase exploration entirely and let Claude choose an approach based solely on which one requires fewer new files to be created", "correct": false, "explanation": "File count is not a meaningful proxy for infrastructure or operational tradeoffs, and this ignores the exploration step needed to compare the approaches properly."}, {"letter": "D", "text": "Have Claude implement the asynchronous queue approach directly and immediately, since new infrastructure work always outranks synchronous changes", "correct": false, "explanation": "Committing to the asynchronous approach without evaluating the tradeoffs skips the architectural decision the team explicitly has not made yet, and could commit them to unwanted new infrastructure."}], "correct": "B", "select": 1, "group": "C"}, {"id": "f3-093", "domain": 3, "task_id": "3.2", "objective": "Create and configure custom slash commands and skills", "situation": "A large, rarely-needed compliance reference document is currently pasted into CLAUDE.md, and the team notices every session now starts with a noticeably larger context footprint even on days nobody needs the compliance information. Once loaded, a skill's body stays in context for the rest of that session too.", "question": "Given this, which change reduces the typical per-session cost most, while still making the reference available when it is actually needed?", "options": [{"letter": "A", "text": "Move the document into a .claude/skills/compliance-reference/SKILL.md skill, since its content is loaded only in the sessions where it's actually invoked, rather than in every session by default.", "correct": true, "explanation": "Moving the document into a .claude/skills/compliance-reference/SKILL.md skill means its content is only loaded when the skill is actually invoked, not in every session by default. This avoids the baseline context cost of having the full compliance text in CLAUDE.md and keeps it available on demand."}, {"letter": "B", "text": "Add effort: low to the YAML frontmatter of CLAUDE.md, which instructs Claude Code to load a reduced token-count representation of the file's instructions, leaving out the rarely-needed compliance details until explicitly requested.", "correct": false, "explanation": "The effort key in YAML frontmatter controls the model's reasoning depth (e.g., thinking budget), not how much of a file's content is loaded. It does not instruct Claude Code to load a reduced token-count representation or omit specific sections; the entire file would still enter context, so per-session cost remains unchanged."}, {"letter": "C", "text": "Split CLAUDE.md into CLAUDE_core.md and CLAUDE_compliance.md, so that Claude Code loads only the alphabetically first file's content per session, keeping the compliance content unloaded unless referenced.", "correct": false, "explanation": "Claude Code does not load files based on alphabetical order; splitting CLAUDE.md into CLAUDE_core.md and CLAUDE_compliance.md would not cause the system to automatically select only the first alphabetically. Without a single CLAUDE.md, neither file might be loaded by default, so this fails to reliably reduce context while maintaining availability."}, {"letter": "D", "text": "Move the document into a .claude/commands/compliance.md command file, since commands undergo transparent compression that leads to a smaller context footprint than raw CLAUDE.md instructions, without any manual intervention.", "correct": false, "explanation": "Command files, like skills, are only loaded when invoked, not automatically compressed. There is no transparent compression that reduces their context footprint; the claim is false, and while moving the content to a command would make it invocation‑based, the stated reasoning is inaccurate and not the primary recommended pattern for keeping reusable knowledge available on demand."}], "correct": "A", "select": 1, "group": "D"}, {"id": "f3-094", "domain": 3, "task_id": "3.4", "objective": "Determine when to use plan mode vs direct execution", "situation": "A senior engineer says a proposed change to the checkout service is small enough that it doesn't need plan mode, but a teammate insists it does because it touches a database schema. Investigation shows the change adds one nullable column to one table and updates one model file to read it, with no other files affected and no schema decisions still open.", "question": "Who has correctly assessed the situation?", "options": [{"letter": "A", "text": "Neither, since schema changes should bypass both plan mode and direct execution entirely and go straight to a migration tool", "correct": false, "explanation": "The question is about which Claude Code workflow to use for the change, not about which database tooling applies; this answer sidesteps the plan-mode-versus-direct-execution decision entirely."}, {"letter": "B", "text": "Neither, since plan mode should be the default for every change touching the checkout service given its business importance", "correct": false, "explanation": "Defaulting to plan mode purely because of the service's business importance ignores the actual scope of the change, which is narrow and unambiguous."}, {"letter": "C", "text": "The senior engineer, since the change is confined to one table and one file with a single clear implementation, matching direct execution", "correct": true, "explanation": "Plan mode is warranted for multi-file changes, multiple valid approaches, or architectural decisions; a single nullable column added to one table with one model file updated and no open design questions is a well-scoped change suited to direct execution."}, {"letter": "D", "text": "The teammate, since any schema change is inherently architectural by nature and must always go through plan mode no matter how small it is", "correct": false, "explanation": "Not every schema change is architectural by definition; the deciding factor is scope and ambiguity, and this change has neither multiple files nor open design decisions."}], "correct": "C", "select": 1, "group": "C"}, {"id": "f3-095", "domain": 3, "task_id": "3.6", "objective": "Integrate Claude Code into CI/CD pipelines", "situation": "A pipeline step runs git diff main | claude -p \"you are a typo linter...\" as part of a lint script in package.json. The maintainer wants Claude to have zero ability to run arbitrary Bash commands during this step, while still keeping the command portable across Windows and Linux runners.", "question": "Which approach achieves the no-Bash-permission goal?", "options": [{"letter": "A", "text": "The double-quoted prompt string itself tells Claude not to use Bash, so the invocation never makes any tool calls regardless of input, and no permission is needed.", "correct": false, "explanation": "Prompt instructions do **not** restrict Claude’s tool usage. Claude may still attempt to call Bash even if told not to, so a natural‑language directive is not a reliable security control."}, {"letter": "B", "text": "Adding --disallowedTools Bash to the command line is the explicit way to block any Bash tool use, because piping alone does not remove the default tool access.", "correct": true, "explanation": "Official documentation confirms that --disallowedTools Bash (or Bash in a deny rule) explicitly removes the Bash tool from Claude’s context. Piping input or using -p does not disable tool access, so the flag is the portable, intended mechanism to guarantee zero Bash execution."}, {"letter": "C", "text": "Running the script through npm instead of a direct shell invocation strips Bash tool access, as npm’s encapsulation blocks tool invocations by the claude process.", "correct": false, "explanation": "npm executes the script as a shell command; it does **not** intercept or modify Claude’s internal tool permissions. The Claude process itself still has Bash tool access unless explicitly denied."}, {"letter": "D", "text": "Piping the diff via stdin sends the changes directly to Claude without a tool invocation, so no Bash permission is needed to access the content itself.", "correct": false, "explanation": "Piping input via stdin does **not** remove Claude’s default access to the Bash tool. Unless --disallowedTools Bash or a similar explicit deny rule is used, Claude can still attempt to invoke Bash during the session."}], "correct": "B", "select": 1, "group": "D"}, {"id": "f3-096", "domain": 3, "task_id": "3.2", "objective": "Create and configure custom slash commands and skills", "situation": "A team using Claude Code wants every request in the project, regardless of topic, to follow a fixed rule: never commit directly to main and always open a pull request instead. The rule must be enforced automatically on relevant actions and must not depend on Claude remembering or choosing to follow an instruction.", "question": "Where should this rule live?", "options": [{"letter": "A", "text": "In a .claude/skills/git-workflow/SKILL.md skill with disable-model-invocation: true, since skills only load when manually typed.", "correct": false, "explanation": "Skills are reusable procedures that load on demand only when relevant to the task; they are not always-on project rules. disable-model-invocation: true prevents Claude from auto-triggering the skill, so it will not shape every relevant action."}, {"letter": "B", "text": "In a PreToolUse hook in the project's .claude/settings.json that inspects Bash tool inputs, blocks any direct commit to main, and returns a message instructing the user to open a pull request.", "correct": true, "explanation": "Claude Code hooks are the documented enforcement mechanism for critical project rules. A PreToolUse hook runs before the Bash tool executes, can inspect the command, and return a blocking decision. This cannot be skipped by the model because the hook is executed by the Claude Code platform, not by following text instructions. Claude Code security guidance explicitly recommends keeping enforcement in hooks rather than relying solely on CLAUDE.md."}, {"letter": "C", "text": "In the project's CLAUDE.md, since that file is always loaded and applies as a universal standard across every conversation in the project.", "correct": false, "explanation": "CLAUDE.md is loaded as project context, but it is guidance, not enforcement. Models can sometimes skip or ignore instructions, especially when files are long or when competing objectives exist. For a fixed rule that must not be violated, use hooks. Project memory files like CLAUDE.md provide context but do not guarantee compliance."}, {"letter": "D", "text": "In a .claude/commands/git-rule.md command, since commands run automatically before every user message without being typed.", "correct": false, "explanation": "Slash commands are user-invoked actions in Claude Code. They do not run automatically before every user message; the user must type or select the command. Thus they cannot enforce a rule on every relevant turn."}], "correct": "B", "select": 1, "group": "D"}, {"id": "g17", "domain": 4, "scenario": "Claude Code for Continuous Integration", "situation": "Your team uses Claude Code for generating code suggestions, but you notice a pattern: non-obvious issues—performance optimizations that break edge cases, cleanups that unexpectedly change behavior—are only caught when another team member reviews the PR. Claude’s reasoning during generation shows it considered these cases but concluded its approach was correct. Which approach directly addresses the root cause of this self-check limitation?", "question": "Which approach directly addresses the root cause?", "options": [{"letter": "A", "text": "Run a second independent instance of Claude Code to review the changes without access to the generator’s reasoning.", "correct": true, "explanation": "A second independent Claude Code instance without access to the generator’s reasoning directly addresses the root cause by avoiding confirmation bias. This “fresh eyes” perspective mirrors human peer review, where another reviewer catches issues the author rationalized."}, {"letter": "B", "text": "Enable extended thinking mode for the generation stage to allow more thorough deliberation before producing suggestions.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Add explicit self-review instructions to the generation prompt asking Claude to critique its own suggestions before finalizing output.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Include full test files and documentation in prompt context so Claude better understands expected behavior during generation.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "4.6", "objective": "Design multi-instance and multi-pass review architectures", "group": "F"}, {"id": "g19", "domain": 4, "scenario": "Claude Code for Continuous Integration", "situation": "Your CI/CD system runs three Claude-based analyses: (1) fast style checks on every PR that block merging until completion, (2) comprehensive weekly security audits of the entire codebase, and (3) nightly test-case generation for recently changed modules. The Message Batches API offers 50% savings but processing can take up to 24 hours. You want to optimize API cost while maintaining an acceptable developer experience. Which combination correctly matches each task to an API approach?", "question": "Which combination is correct?", "options": [{"letter": "A", "text": "Use the Message Batches API for all three tasks to maximize 50% savings, configuring the pipeline to poll for batch completion.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Use synchronous calls for PR style checks; use the Message Batches API for weekly security audits and nightly test generation.", "correct": true, "explanation": "PR style checks block developers and require immediate responses via synchronous calls, while weekly security audits and nightly test generation are scheduled tasks with flexible deadlines that can tolerate up to a 24-hour batch window—capturing 50% savings for both."}, {"letter": "C", "text": "Use synchronous calls for all three tasks for consistent response times, relying on prompt caching to reduce costs across workloads.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Use synchronous calls for PR style checks and nightly test generation; use the Message Batches API only for weekly security audits.", "correct": false, "explanation": ""}], "correct": "B", "task_id": "4.5", "objective": "Design efficient batch processing strategies", "group": "H"}, {"id": "g20", "domain": 4, "scenario": "Claude Code for Continuous Integration", "situation": "Your automated reviews find real issues, but developers report the feedback is not actionable. Findings include phrases like “complex ticket routing logic” or “potential null pointer” without specifying what exactly to change. When you add detailed instructions like “always include concrete fix suggestions,” the model still produces inconsistent output—sometimes detailed, sometimes vague. Which prompting technique most reliably produces consistently actionable feedback?", "question": "Which prompting technique is most reliable?", "options": [{"letter": "A", "text": "Further refine instructions with more explicit requirements for each part of the feedback format (location, issue, severity, proposed fix).", "correct": false, "explanation": ""}, {"letter": "B", "text": "Expand the context window to include more surrounding codebase so the model has enough information to propose concrete fixes.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Implement a two-pass approach where one prompt identifies issues and a second generates fixes, allowing specialization.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Add 3–4 few-shot examples showing the exact required format: identified issue, location in code, concrete fix suggestion.", "correct": true, "explanation": "Few-shot examples are the most effective technique for achieving consistent output format when instructions alone produce variable results. Providing 3–4 examples that show the exact desired structure (issue, location, concrete fix) gives the model a concrete pattern to follow, which is more reliable than abstract instructions."}], "correct": "D", "task_id": "3.5", "objective": "Apply iterative refinement techniques for progressive improvement", "group": "I"}, {"id": "g21", "domain": 4, "scenario": "Claude Code for Continuous Integration", "situation": "Your CI pipeline includes two Claude-based code review modes: a pre-merge-commit hook that blocks PR merge until completion, and a “deep analysis” that runs overnight, polls for batch completion, and posts detailed suggestions to the PR. You want to reduce API cost using the Message Batches API, which offers 50% savings but requires polling and can take up to 24 hours. Which mode should use batch processing?", "question": "Which mode should use batch processing?", "options": [{"letter": "A", "text": "Only the pre-merge-commit hook.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Only the deep analysis.", "correct": true, "explanation": "Deep analysis is an ideal candidate for batch processing because it already runs overnight, tolerates delay, and uses a polling model before publishing results—matching the asynchronous, polling-based architecture of the Message Batches API while capturing 50% savings."}, {"letter": "C", "text": "Both modes.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Neither mode.", "correct": false, "explanation": ""}], "correct": "B", "task_id": "4.5", "objective": "Design efficient batch processing strategies", "group": "H"}, {"id": "g22", "domain": 4, "scenario": "Claude Code for Continuous Integration", "situation": "Your automated review analyzes comments and docstrings. The current prompt instructs Claude to “check that comments are accurate and up to date.” Findings often flag acceptable patterns (TODO markers, simple descriptions) while missing comments describing behavior the code no longer implements. What change addresses the root cause of this inconsistent analysis?", "question": "What change addresses the root cause?", "options": [{"letter": "A", "text": "Include `git blame` data so Claude can identify comments that predate recent code changes.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Add few-shot examples of misleading comments to help the model recognize similar patterns in the codebase.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Filter TODO, FIXME, and descriptive comment patterns before analysis to reduce noise.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Specify explicit criteria: flag comments only when the behavior they claim contradicts the code’s actual behavior.", "correct": true, "explanation": "Explicit criteria—flagging comments only when claimed behavior contradicts actual code behavior—directly addresses the root cause by replacing a vague instruction with a precise definition of what constitutes a problem. This reduces false positives on acceptable patterns and misses of truly misleading comments."}], "correct": "D", "task_id": "4.1", "objective": "Design prompts with explicit criteria to improve precision and reduce false positives", "group": "I"}, {"id": "g23", "domain": 4, "scenario": "Claude Code for Continuous Integration", "situation": "Your automated code review system shows inconsistent severity ratings—similar issues like null pointer risks are rated “critical” in some PRs but only “medium” in others. Developer surveys show growing distrust—many start dismissing findings without reading because “half are wrong.” High-false-positive categories erode trust in accurate categories. Which approach best restores developer trust while improving the system?", "question": "Which approach best restores developer trust?", "options": [{"letter": "A", "text": "Temporarily disable high-false-positive categories (style, naming, documentation) and keep only high-precision categories while improving prompts.", "correct": true, "explanation": "Temporarily disabling high-false-positive categories immediately stops trust erosion by removing noisy findings that cause developers to dismiss everything, while preserving value from high-precision categories like security and correctness. It also creates space to improve prompts for problematic categories before re-enabling them."}, {"letter": "B", "text": "Keep all categories enabled but display confidence scores with each finding so developers can decide what to investigate.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Keep all categories enabled and add few-shot examples to improve accuracy for each category over the next few weeks.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Apply a uniform strictness reduction across all categories to bring the overall false-positive rate down.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "4.1", "objective": "Design prompts with explicit criteria to improve precision and reduce false positives", "group": "I"}, {"id": "g24", "domain": 4, "scenario": "Claude Code for Continuous Integration", "situation": "Your automated review generates test-case suggestions for each PR. Reviewing a PR that adds course completion tracking, Claude suggests 10 test cases, but developer feedback shows that 6 duplicate scenarios already covered by the existing test suite. What change most effectively reduces duplicate suggestions?", "question": "What change is most effective?", "options": [{"letter": "A", "text": "Include the existing test file in context so Claude can determine what scenarios are already covered.", "correct": true, "explanation": "Including the existing test file fixes the root cause of duplication: Claude can only avoid suggesting already-covered scenarios if it knows what tests already exist. This gives Claude the information needed to propose genuinely new, valuable tests."}, {"letter": "B", "text": "Reduce the requested number of suggestions from 10 to 5, assuming Claude prioritizes the most valuable cases first.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Add instructions directing Claude to focus exclusively on edge cases and error conditions rather than success paths.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Implement post-processing that filters suggestions whose descriptions match existing test names via keyword overlap.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "3.5", "objective": "Apply iterative refinement techniques for progressive improvement", "group": "I"}, {"id": "g25", "domain": 4, "scenario": "Claude Code for Continuous Integration", "situation": "After an initial automated review identifies 12 findings, a developer pushes new commits to address issues. Re-running review produces 8 findings, but developers report that 5 duplicate previous comments on code that was already fixed in the new commits. What is the most effective way to eliminate this redundant feedback while maintaining thoroughness?", "question": "What is the most effective way to eliminate redundant feedback?", "options": [{"letter": "A", "text": "Run review only when the PR is created and in the final pre-merge state, skipping intermediate commits.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Add a post-processing filter that removes findings that match previous ones by file paths and issue descriptions before posting comments.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Restrict review scope to files changed in the most recent push, excluding files from earlier commits.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Include previous review findings in context and instruct Claude to report only new or still-unresolved issues.", "correct": true, "explanation": "Including prior review findings in context lets Claude distinguish new problems from those already addressed in recent commits. This preserves review thoroughness while using Claude’s reasoning to avoid redundant feedback on fixed code."}], "correct": "D", "task_id": "4.1", "objective": "Design prompts with explicit criteria to improve precision and reduce false positives", "group": "I"}, {"id": "g27", "domain": 4, "scenario": "Claude Code for Continuous Integration", "situation": "A pull request changes 14 files in an inventory tracking module. A single-pass review that analyzes all files together produces inconsistent results: detailed feedback on some files but shallow comments on others, missed obvious bugs, and contradictory feedback (a pattern is flagged in one file but identical code is approved in another file in the same PR). How should you restructure the review?", "question": "How should you restructure the review?", "options": [{"letter": "A", "text": "Run three independent full-PR review passes and flag only issues that appear in at least two of the three runs.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Split into focused passes: review each file individually for local issues, then run a separate integration-oriented pass to examine cross-file data flows.", "correct": true, "explanation": "Focused per-file passes address the root cause—attention dilution—by ensuring consistent depth and reliable local issue detection. A separate integration-oriented pass then covers cross-file concerns such as dependency and data-flow interactions."}, {"letter": "C", "text": "Require developers to split large PRs into smaller submissions of 3–4 files before running automated review.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Switch to a larger model with a bigger context window so it can pay sufficient attention to all 14 files in one pass.", "correct": false, "explanation": ""}], "correct": "B", "task_id": "4.6", "objective": "Design multi-instance and multi-pass review architectures", "group": "F"}, {"id": "g28", "domain": 4, "scenario": "Claude Code for Continuous Integration", "situation": "Your automated code review averages 15 findings per pull request, and developers report a 40% false-positive rate. The bottleneck is investigation time: developers must click into each finding to read Claude’s rationale before deciding whether to fix or dismiss it. Your CLAUDE.md already contains comprehensive rules for acceptable patterns, and stakeholders rejected any approach that filters findings before developers see them. What change best addresses investigation time?", "question": "What change best addresses investigation time?", "options": [{"letter": "A", "text": "Require Claude to include its rationale and confidence estimate directly in each finding.", "correct": true, "explanation": "Including rationale and confidence directly in each finding reduces investigation time by letting developers quickly triage without opening each finding. It satisfies the “no filtering” constraint because all findings remain visible while accelerating developer decision-making."}, {"letter": "B", "text": "Add a post-processor that analyzes finding patterns and automatically suppresses those that match historical false-positive signatures.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Categorize findings as “blocking issues” vs “suggestions,” with different review requirements by level.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Configure Claude to show only high-confidence findings, filtering uncertain flags before developers see them.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "4.6", "objective": "Design multi-instance and multi-pass review architectures", "group": "F"}, {"id": "g29", "domain": 4, "scenario": "Claude Code for Continuous Integration", "situation": "Analysis of your automated code review shows large differences in false-positive rates by finding category: security/correctness findings have 8% false positives, performance findings 18%, style/naming findings 52%, and documentation findings 48%. Developer surveys show growing distrust—many start dismissing findings without reading because “half are wrong.” High-false-positive categories erode trust in accurate categories. Which approach best restores developer trust while improving the system?", "question": "Which approach best restores developer trust?", "options": [{"letter": "A", "text": "Temporarily disable high-false-positive categories (style, naming, documentation) and keep only high-precision categories while improving prompts.", "correct": true, "explanation": "Temporarily disabling high-false-positive categories immediately stops trust erosion by removing noisy findings that cause developers to dismiss everything, while preserving value from high-precision categories like security and correctness. It also creates space to improve prompts for problematic categories before re-enabling them."}, {"letter": "B", "text": "Keep all categories enabled but display confidence scores with each finding so developers can decide what to investigate.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Keep all categories enabled and add few-shot examples to improve accuracy for each category over the next few weeks.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Apply a uniform strictness reduction across all categories to bring the overall false-positive rate down.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "4.1", "objective": "Design prompts with explicit criteria to improve precision and reduce false positives", "group": "I"}, {"id": "g30", "domain": 4, "scenario": "Claude Code for Continuous Integration", "situation": "Your team wants to reduce API costs for automated analysis. Currently, synchronous Claude calls support two workflows: (1) a blocking pre-merge check that must complete before developers can merge, and (2) a technical debt report generated overnight for review the next morning. Your manager proposes moving both to the Message Batches API to save 50%. How should you evaluate this proposal?", "question": "How should you evaluate this proposal?", "options": [{"letter": "A", "text": "Move both to batch processing with fallback to synchronous calls if batches take too long.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Move both workflows to batch processing with status polling to verify completion.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Use batch processing only for technical debt reports; keep synchronous calls for pre-merge checks.", "correct": true, "explanation": "Message Batches API processing can take up to 24 hours with no latency SLA, which is acceptable for overnight technical debt reports but unacceptable for blocking pre-merge checks where developers wait. This matches each workflow to the right API based on latency requirements."}, {"letter": "D", "text": "Keep synchronous calls for both workflows to avoid issues with batch result ordering.", "correct": false, "explanation": ""}], "correct": "C", "task_id": "4.5", "objective": "Design efficient batch processing strategies", "group": "H"}, {"id": "g31", "domain": 4, "scenario": "Code Generation with Claude Code", "situation": "You asked Claude Code to implement a function that transforms API responses into an internal normalized format. After two iterations, the output structure still doesn’t match expectations—some fields are nested differently and timestamps are formatted incorrectly. You described requirements in prose, but Claude interprets them differently each time.", "question": "Which approach is most effective for the next iteration?", "options": [{"letter": "A", "text": "Write a JSON schema describing the expected output structure and validate Claude’s output against it after each iteration.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Provide 2–3 concrete input-output examples showing the expected transformation for representative API responses.", "correct": true, "explanation": "Concrete input-output examples remove ambiguity inherent in prose descriptions by showing Claude the exact expected transformation results. This directly addresses the root cause—misinterpretation of textual requirements—by providing unambiguous patterns for field nesting and timestamp formatting."}, {"letter": "C", "text": "Rewrite requirements with more technical precision, specifying exact field mappings, nesting rules, and timestamp format strings.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Ask Claude to explain its current understanding of the requirements to identify where interpretations diverge.", "correct": false, "explanation": ""}], "correct": "B", "task_id": "3.5", "objective": "Apply iterative refinement techniques for progressive improvement", "group": "I"}, {"id": "g49", "domain": 4, "scenario": "Customer Support Agent", "situation": "Your agent achieves 55% first-contact resolution, well below the 80% target. Logs show it escalates simple cases (standard replacements for damaged goods with photo proof) while trying to handle complex situations requiring policy exceptions autonomously. What is the most effective way to improve escalation calibration?", "question": "What is the most effective way to improve escalation calibration?", "options": [{"letter": "A", "text": "Require the agent to self-rate confidence on a 1–10 scale before each response and automatically route to humans when confidence drops below a threshold.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Deploy a separate classifier model trained on historical tickets to predict which requests need escalation before the main agent starts processing.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Add explicit escalation criteria to the system prompt with few-shot examples showing when to escalate versus resolve autonomously.", "correct": true, "explanation": "Explicit escalation criteria with few-shot examples directly address the root cause—unclear decision boundaries between simple and complex cases. It’s the most proportional, effective first intervention that teaches the agent when to escalate and when to resolve autonomously without extra infrastructure."}, {"letter": "D", "text": "Implement sentiment analysis to determine customer frustration level and automatically escalate past a negative sentiment threshold.", "correct": false, "explanation": ""}], "correct": "C", "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "group": "C"}, {"id": "g52", "domain": 4, "scenario": "Customer Support Agent", "situation": "Production metrics show that when resolving complex billing disputes or multi-order returns, customer satisfaction scores are 15% lower than for simple cases—even when the resolution is technically correct. Root-cause analysis shows the agent provides accurate solutions but inconsistently explains rationale: sometimes omitting relevant policy details, sometimes missing timeline info or next steps. The specific context gaps vary case by case. You want to improve solution quality without adding human oversight. What approach is most effective?", "question": "What approach is most effective?", "options": [{"letter": "A", "text": "Add a self-critique stage where the agent evaluates a draft response for completeness—ensuring it resolves the customer’s issue, includes relevant context, and anticipates follow-up questions.", "correct": true, "explanation": "A self-critique stage (the evaluator-optimizer pattern) directly addresses inconsistent explanation completeness by forcing the agent to assess its own draft against concrete criteria—such as policy context, timelines, and next steps—before presenting it. This catches case-specific gaps without human oversight."}, {"letter": "B", "text": "Add a confirmation stage where the agent asks “Does this fully resolve your issue?” before closing, allowing customers to request additional information if needed.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Upgrade the model from Haiku to Sonnet for complex cases, routing based on a defined complexity metric.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Implement few-shot examples in the system prompt showing complete explanations for five common complex case types, demonstrating how to include policy context, timelines, and next steps.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "3.5", "objective": "Apply iterative refinement techniques for progressive improvement", "group": "I"}, {"id": "g56", "domain": 4, "scenario": "Customer Support Agent", "situation": "Production logs show a consistent pattern: when customers include the word “account” in their message (e.g., “I want to check my account for an order I made yesterday”), the agent calls `get_customer` first 78% of the time. When customers phrase similar requests without “account” (e.g., “I want to check an order I made yesterday”), it calls `lookup_order` first 93% of the time. Tool descriptions are clear and unambiguous. What is the most likely root cause of this discrepancy?", "question": "What is the most likely root cause?", "options": [{"letter": "A", "text": "The system prompt contains keyword-sensitive instructions that steer behavior based on terms like “account,” creating unintended tool-selection patterns.", "correct": true, "explanation": "The systematic keyword-driven pattern (78% vs 93%) strongly indicates explicit routing logic in the system prompt reacting to the word “account” and steering the agent toward customer-related tools. Since tool descriptions are already clear, the discrepancy points to prompt-level instructions creating unintended behavioral steering."}, {"letter": "B", "text": "The model’s base training creates associations between “account” terminology and customer-related operations that override tool descriptions.", "correct": false, "explanation": ""}, {"letter": "C", "text": "The model needs more training data on multi-concept messages and should be fine-tuned on examples containing both account and order terminology.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Tool descriptions need additional negative examples specifying when NOT to use each tool to prevent this keyword-induced confusion.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "4.1", "objective": "Design prompts with explicit criteria to improve precision and reduce false positives", "group": "I"}, {"id": "g60", "domain": 4, "scenario": "Customer Support Agent", "situation": "Production logs show the agent sometimes chooses `get_customer` when `lookup_order` would be more appropriate, especially for ambiguous queries like “I need help with my recent purchase.” You decide to add few-shot examples to the system prompt to improve tool selection. Which approach most effectively addresses the problem?", "question": "Which approach is most effective?", "options": [{"letter": "A", "text": "Add explicit “use when” and “don’t use when” guidance in each tool description covering ambiguous cases.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Add examples grouped by tool—all `get_customer` scenarios together, then all `lookup_order` scenarios.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Add 4–6 examples targeted at ambiguous scenarios, each with rationale for why one tool was chosen over plausible alternatives.", "correct": true, "explanation": "Targeting few-shot examples at the specific ambiguous scenarios where errors occur, with explicit rationale for why one tool is preferable to alternatives, teaches the model the comparative decision process needed for edge cases. This is more effective than generic examples or declarative rules."}, {"letter": "D", "text": "Add 10–15 examples of clear, unambiguous requests demonstrating correct tool choice for typical scenarios for each tool.", "correct": false, "explanation": ""}], "correct": "C", "task_id": "2.1", "objective": "Design effective tool interfaces with clear descriptions and boundaries", "group": "J"}, {"id": "g69", "domain": 4, "scenario": "Conversational AI Architecture Patterns", "situation": "During QA testing, Claude follows system prompt guidelines for the first 10–15 turns, but later responses deviate. The conversation is still within token limits.", "question": "What is the best solution?", "options": [{"letter": "A", "text": "Move behavioral guidelines into the first user message.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Start a new conversation after 20 turns.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Insert user-role messages reinforcing guidelines at conversation breakpoints.", "correct": true, "explanation": "Periodic injection of behavioral reminders directly combats instruction drift by re-establishing constraints at regular intervals as conversation history accumulates. Moving guidelines to the first user message (A) reduces their authority. Starting a new conversation (B) destroys context. Post-response validation (D) is corrective rather than preventive and adds significant latency."}, {"letter": "D", "text": "Use post-response validation to regenerate non-compliant responses.", "correct": false, "explanation": ""}], "correct": "C", "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "group": "G"}, {"id": "g71", "domain": 4, "scenario": "Conversational AI Architecture Patterns", "situation": "Your assistant must maintain an enthusiastic tone, explain its reasoning, and ask clarifying questions. Where should these behavioral guidelines be defined?", "question": "Where should these behavioral guidelines be defined?", "options": [{"letter": "A", "text": "Prepended to each user message.", "correct": false, "explanation": ""}, {"letter": "B", "text": "In the system prompt.", "correct": true, "explanation": "The system prompt is specifically designed for persistent behavioral constraints and guidelines that apply throughout the entire conversation. Prepending to each user message (A) is redundant overhead. The first assistant message (C) is unreliable because the model can deviate from its own prior statements. Environment variables (D) have no effect on model behavior."}, {"letter": "C", "text": "In the first assistant message.", "correct": false, "explanation": ""}, {"letter": "D", "text": "In environment variables.", "correct": false, "explanation": ""}], "correct": "B", "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "group": "G"}, {"id": "g72", "domain": 4, "scenario": "Conversational AI Architecture Patterns", "situation": "Users report repetitive response openings like \"Certainly!\" and \"I'd be happy to help!\"", "question": "What is the most effective approach?", "options": [{"letter": "A", "text": "Append a partial assistant message with a direct response opening.", "correct": true, "explanation": "Prefilling the assistant's response with the beginning of a direct answer prevents greeting patterns at the generation level—the model continues from the prefill rather than generating new opening phrases. System prompt instructions (D) can help but are less reliable since the model may still produce variants. Post-processing (C) is a fragile workaround. Temperature (B) controls randomness, not specific phrase patterns."}, {"letter": "B", "text": "Lower the temperature setting.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Post-process responses to remove greetings.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Add system prompt instructions to avoid those phrases.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "4.3", "objective": "Enforce structured output using tool use and JSON schemas", "group": "H"}, {"id": "g73", "domain": 4, "scenario": "Conversational AI Architecture Patterns", "situation": "A webhook notifies your system that a user's package has shipped while the user is actively chatting. You want the assistant to incorporate this naturally into the next response.", "question": "What is the best approach?", "options": [{"letter": "A", "text": "Add shipping status to the system prompt.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Send an immediate synthetic user message.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Force the assistant to call a status tool on each turn.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Append the status update as a prefix to the next user message.", "correct": true, "explanation": "Prefixing the status update to the next user message injects real-time context at a natural conversation boundary without disrupting the flow. Modifying the system prompt (A) requires rebuilding the session or is architecturally cumbersome. A synthetic user message (B) can break the natural dialogue flow and confuse attribution. Forcing a tool call each turn (C) is wasteful when events are rare."}], "correct": "D", "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "group": "G"}, {"id": "g74", "domain": 4, "scenario": "Conversational AI Architecture Patterns", "situation": "Users frequently send requests like \"Book a venue for the party.\" The assistant asks 4+ clarifying questions, causing 35% abandonment.", "question": "What approach best improves the trade-off?", "options": [{"letter": "A", "text": "Proceed with hidden defaults.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Ask all clarifying questions in one compound message.", "correct": false, "explanation": ""}, {"letter": "C", "text": "State assumptions explicitly and proceed while inviting corrections.", "correct": true, "explanation": "Stating assumptions explicitly and proceeding gives the user an immediate, useful response while preserving their ability to correct wrong assumptions. Hidden defaults (A) leave the user unaware of what was assumed. A compound question list (B) still demands upfront effort from the user. A structured form (D) adds more friction, not less—contradicting the goal of reducing abandonment."}, {"letter": "D", "text": "Use a structured intake form.", "correct": false, "explanation": ""}], "correct": "C", "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "group": "C"}, {"id": "g76", "domain": 4, "scenario": "Conversational AI Architecture Patterns", "situation": "Users ask vague requests like \"Can you help with the report?\" The assistant responds by asking multiple questions (which report? what help? deadline?), causing 40% abandonment.", "question": "What is the best solution?", "options": [{"letter": "A", "text": "Make reasonable assumptions, state them explicitly, and offer to adjust.", "correct": true, "explanation": "Proceeding with reasonable stated assumptions eliminates the back-and-forth entirely while keeping the user informed and in control. Predefined silent interpretations (C) leave users confused when the response doesn't match their intent. A single-question limit (D) still requires turns of back-and-forth. A smaller classification model (B) adds latency and infrastructure complexity without solving the core UX problem."}, {"letter": "B", "text": "Classify ambiguity with a smaller model before responding.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Use predefined interpretations without stating assumptions.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Limit the assistant to one clarifying question per turn.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "group": "C"}, {"id": "m8", "domain": 4, "scenario": null, "situation": "Production reviews reveal inconsistent handling of uncertainty in final reports. Sometimes conflicting subagent findings are synthesized into a single confident statement (losing nuance), while other times reports over-hedge with excessive qualifications (becoming unhelpful). When the web search agent returns \"industry analysts estimate $50B market size (methodology varies)\" and the document analysis agent returns \"peer-reviewed study estimates 35B(±7B, 95% CI),\" the coordinator either picks one arbitrarily or produces vague statements like \"the market may be 35B−50B depending on factors.\"", "question": "What systematic approach best addresses this?", "options": [{"letter": "A", "text": "Configure subagents to only report findings meeting a high-confidence threshold, filtering uncertain information before it reaches the coordinator.", "correct": false, "explanation": "Filtering throws away useful-but-uncertain evidence. It doesn't help the synthesis agent reason about conflicts — it just hides them."}, {"letter": "B", "text": "Implement a confidence calibration layer that normalizes subagent uncertainty expressions to standardized probability scores (0.0-1.0), then weight-average findings by their calibrated confidence.", "correct": false, "explanation": "Collapsing peer-reviewed CIs and analyst estimates into a single averaged number destroys the methodological difference that's the whole point."}, {"letter": "C", "text": "Instruct the synthesis agent to structure reports with explicit sections distinguishing well-established findings from contested ones, preserving original source characterizations and methodological context.", "correct": true, "explanation": "Correct. Report structure that keeps methodological context and separates settled vs. contested claims is how you get nuance without over-hedging."}, {"letter": "D", "text": "Add a verification subagent that cross-references findings across sources, only passing claims to synthesis that are corroborated by at least two independent sources.", "correct": false, "explanation": "Two-source minimums drop legitimate single-source findings and don't resolve genuinely conflicting estimates with different methodologies."}], "correct": "C", "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "group": "B"}, {"id": "m11", "domain": 4, "scenario": null, "situation": "A user is expanding the research system beyond its single web search agent by adding specialized data sources. They add a financial API agent that returns structured JSON with revenue, margins, and growth rates; a news monitoring agent that returns prose summaries of recent developments; and a patent analysis agent that returns structured lists of technology areas. The synthesis agent combines these into executive briefings. Currently, it converts everything to bullet points, causing financial comparisons to lose tabular clarity and news summaries to lose narrative flow.", "question": "What change would most improve briefing quality?", "options": [{"letter": "A", "text": "Standardize all subagent outputs to prose summaries with inline citations.", "correct": false, "explanation": "Flattening structured financial data into prose loses the comparability you want in an executive briefing."}, {"letter": "B", "text": "Add a format conversion layer between subagents and synthesis that transforms all outputs to a common intermediate representation.", "correct": false, "explanation": "A single IR usually ends up too lossy — either it keeps structure (news suffers) or it keeps prose (tables suffer). The fix is type-aware rendering, not uniformity."}, {"letter": "C", "text": "Update the synthesis agent to render each content type appropriately—financial data as tables, news as prose.", "correct": true, "explanation": "Correct. Executive briefings need mixed rendering: tables for numbers, prose for narrative. Asking synthesis to preserve the native format of each input is the right abstraction."}, {"letter": "D", "text": "Standardize all subagent outputs to JSON with fields for claim, evidence, source, and confidence.", "correct": false, "explanation": "Forcing prose-heavy news into a claim/evidence schema strips the narrative flow executives actually read."}], "correct": "C", "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "group": "B"}, {"id": "m46", "domain": 4, "scenario": null, "situation": "Your extraction system processes two document types: standard monthly reports (archived after processing) and urgent exception reports (must trigger business alerts within 30 minutes of receipt). Both use the same JSON schema. You want to minimize API costs while meeting latency requirements.", "question": "How should you architect the processing pipeline?", "options": [{"letter": "A", "text": "Submit all documents to the real-time Messages API to ensure consistent processing latency across document types.", "correct": false, "explanation": "Consistent but expensive. Standard monthly reports don't need real-time latency and shouldn't pay for it."}, {"letter": "B", "text": "Submit all documents to the `Batch API` with `custom_ids` for tracking. When results arrive, immediately process urgent documents and trigger delayed alerts for exceptions.", "correct": false, "explanation": "Batch has up to a 24-hour window — 30-minute alerting SLAs on exception reports can't be met by routing them through batch."}, {"letter": "C", "text": "Queue all documents and submit hourly batches, flagging urgent documents for expedited handling when batch results return.", "correct": false, "explanation": "Hourly batches still ride the batch SLO and don't guarantee the 30-minute alert window for exceptions."}, {"letter": "D", "text": "Route standard reports to the `Batch API` for 50% cost savings, and route urgent exception reports to the real-time Messages API.", "correct": true, "explanation": "Correct. Match latency profile to document urgency: batch for the bulk (cheap), real-time for the latency-sensitive exceptions (fast). Minimizes cost while meeting SLA."}], "correct": "D", "task_id": "4.5", "objective": "Design efficient batch processing strategies", "group": "H"}, {"id": "m47", "domain": 4, "scenario": null, "situation": "Your schema includes a skills: string[] field. Production monitoring reveals three consistency issues: (1) compound phrases like \"Python and SQL\" are sometimes kept as one entry, sometimes split; (2) implied but unstated skills occasionally appear in extractions; (3) similar documents produce wildly different array lengths (5-10 vs 40+ entries). Your prompt currently says \"Extract all skills mentioned.\"", "question": "What's the most effective improvement?", "options": [{"letter": "A", "text": "Add few-shot examples demonstrating compound phrase handling, explicit mention criteria, and appropriate entry granularity.", "correct": true, "explanation": "Correct. All three issues are about the model's interpretation of what counts as 'a skill.' Few-shot examples teach the pattern concretely — split vs. not-split, mentioned vs. inferred, appropriate granularity."}, {"letter": "B", "text": "Add constraints: \"Extract 10-20 skills maximum, one skill per entry, only explicitly named skills.\"", "correct": false, "explanation": "Hard count caps are arbitrary and can force the model to drop real skills or invent filler to hit the range."}, {"letter": "C", "text": "Add post-extraction normalization that maps skills to a canonical taxonomy and deduplicates similar entries.", "correct": false, "explanation": "Post-processing can't fix inferred-but-not-mentioned skills or choose the right split for 'Python and SQL' — that decision has to be made at extraction time."}, {"letter": "D", "text": "Enrich the schema to {skill: string, confidence: float, `source_quote`: string}[] to capture extraction metadata.", "correct": false, "explanation": "Useful signal, but doesn't address the underlying inconsistency in how skills are parsed in the first place."}], "correct": "A", "task_id": "4.2", "objective": "Apply few-shot prompting to improve output consistency and quality", "group": "I"}, {"id": "m48", "domain": 4, "scenario": null, "situation": null, "question": "Your system has been operating with 100% human review for 3 months. Analysis shows that extractions with model confidence >90% have 97% accuracy overall. To reduce reviewer workload, you plan to automate high-confidence extractions. Before deploying, what validation step is most critical?", "options": [{"letter": "A", "text": "Analyze accuracy by document type and field to verify high-confidence extractions perform consistently across all segments, not just in aggregate.", "correct": true, "explanation": "Correct. Aggregate accuracy hides segment failures — one document type could be 70% while others are 99%. Auto-routing by overall number alone risks systemic errors in the weak segments."}, {"letter": "B", "text": "Compare accuracy at different confidence thresholds (85%, 90%, 95%) to find the optimal cutoff that maximizes automation while minimizing errors.", "correct": false, "explanation": "Threshold tuning is worth doing, but only after you know the accuracy holds uniformly across segments — otherwise you're optimizing a number that lies."}, {"letter": "C", "text": "Run a two-week pilot routing 25% of high-confidence extractions directly to downstream systems and monitor error reports.", "correct": false, "explanation": "A pilot is good practice, but it's downstream of the segment analysis — if a segment is systematically wrong, you'll learn it by harming users."}, {"letter": "D", "text": "Verify that 97% accuracy meets requirements for all downstream systems that consume the extracted data.", "correct": false, "explanation": "An important product question, but it assumes the 97% holds everywhere — again, that's what segment analysis verifies first."}], "correct": "A", "task_id": "5.5", "objective": "Design human review workflows and confidence calibration", "group": "C"}, {"id": "m49", "domain": 4, "scenario": null, "situation": "Your extraction pipeline processes contracts that frequently include amendments. When a contract contains both original terms and later amendments (e.g., original clause specifies \"30-day payment terms\" while Amendment 1 changes this to \"45 days\"), the model inconsistently extracts one value or the other with no indication of which applies.", "question": "What's the most effective approach to improve extraction accuracy for documents with amendments?", "options": [{"letter": "A", "text": "Redesign the schema so amended fields capture multiple values, each with source location and effective date.", "correct": true, "explanation": "Correct. Amendments are structurally about versioned values. A schema with value + source location + effective date models the domain correctly and stops forcing the model to pick one."}, {"letter": "B", "text": "Add prompt instructions to always extract the most recent amendment value and ignore superseded original terms.", "correct": false, "explanation": "Downstream consumers sometimes need the original too (e.g., for dispute resolution). Hard-coding 'most recent wins' throws that away."}, {"letter": "C", "text": "Preprocess documents with a classifier that identifies and removes superseded sections before the main extraction step.", "correct": false, "explanation": "Deleting sections with a classifier can accidentally strip content that's still legally effective — amendments often modify subsets of clauses."}, {"letter": "D", "text": "Implement post-extraction validation using pattern matching to detect amendments and flag those extractions for manual review.", "correct": false, "explanation": "Routes every amended contract to a human — expensive, and still doesn't give downstream systems a structured answer."}], "correct": "A", "task_id": "4.3", "objective": "Enforce structured output using tool use and JSON schemas", "group": "H"}, {"id": "m50", "domain": 4, "scenario": null, "situation": "Your extraction system implements automatic retries when validation fails. On each retry, the specific validation error is appended to the prompt. This retry-with-error-feedback approach resolves most failures within 2-3 attempts.", "question": "For which failure pattern would additional retries be LEAST effective?", "options": [{"letter": "A", "text": "The model extracts keywords as a nested object organized by category when the schema requires a flat array of strings", "correct": false, "explanation": "A structural mismatch the model can definitely correct given feedback about the expected shape."}, {"letter": "B", "text": "The model extracts citation counts as locale-formatted strings (\"1,234\") when the schema requires integers", "correct": false, "explanation": "The required number is present in the document — the model just needs to stop formatting it. Easy for feedback-retry to fix."}, {"letter": "C", "text": "The model extracts dates as ISO 8601 datetime strings (\"2023-03-15T00:00:00Z\") when the schema requires only the date portion (YYYY-MM-DD)", "correct": false, "explanation": "Pure formatting fix — the underlying value is correct and the model can trim the time portion on retry."}, {"letter": "D", "text": "The model extracts \"et al.\" for co-authors when the full list exists only in an external document not in the input", "correct": true, "explanation": "Correct. No amount of retrying teaches the model information that isn't in the input. Retry-with-error-feedback only fixes mistakes the model could have gotten right from the source."}], "correct": "D", "task_id": "3.5", "objective": "Apply iterative refinement techniques for progressive improvement", "group": "I"}, {"id": "m51", "domain": 4, "scenario": null, "situation": "Your extraction pipeline processes restaurant menus and must output structured JSON with fields for item names, descriptions, prices, and dietary tags. Some menus use inconsistent formatting—prices as \"$12\" vs \"12.00\", dietary info as icons vs text.", "question": "What's the most reliable approach?", "options": [{"letter": "A", "text": "Use separate extraction calls for each field to ensure consistent handling of each type.", "correct": false, "explanation": "Multiple calls per menu multiplies cost and latency and loses cross-field context (e.g., which dietary icons sit next to which item)."}, {"letter": "B", "text": "Extract data as-is and normalize formats in post-processing code after Claude returns.", "correct": false, "explanation": "Post-processing a raw blob is brittle — deterministic code has to handle every format permutation. Better to let the model normalize at extraction time."}, {"letter": "C", "text": "Request multiple extraction attempts per document and select the most common format.", "correct": false, "explanation": "Ensemble voting is expensive and can still pick a wrong-format majority."}, {"letter": "D", "text": "Define a strict output schema and include format normalization rules in your prompt.", "correct": true, "explanation": "Correct. Strict schema + explicit normalization rules ('prices as decimal with two places', 'dietary as enumerated tags') lets the model both extract and normalize in one pass."}], "correct": "D", "task_id": "4.2", "objective": "Apply few-shot prompting to improve output consistency and quality", "group": "I"}, {"id": "m52", "domain": 4, "scenario": null, "situation": "Your system extracts event metadata (date, location, organizer, `attendee_count`) from news articles using a JSON schema with all nullable fields. During evaluation, you observe the model frequently generates plausible but incorrect values for fields not mentioned in the article—for example, outputting \"500\" for `attendee_count` when the source contains no attendance information.", "question": "What's the most effective way to reduce these false extractions?", "options": [{"letter": "A", "text": "Add a post-processing step using a second LLM call to verify each extracted value exists in the source document.", "correct": false, "explanation": "Expensive second pass for a behavior you can correct at the source — tell the model to return null when the field isn't stated."}, {"letter": "B", "text": "Add prompt instructions to return null for any field where information is not directly stated in the source.", "correct": true, "explanation": "Correct. The fields are already nullable; the model just needs an explicit instruction to prefer null over a plausible guess. This is the standard fix for schema-aware hallucination."}, {"letter": "C", "text": "Make all schema fields required (non-nullable) with strict validation rules to ensure the model only outputs verifiable data.", "correct": false, "explanation": "Required fields force the model to invent values when information is missing — the opposite of what you want."}, {"letter": "D", "text": "Upgrade to a more capable model tier with improved instruction-following to reduce hallucination tendencies.", "correct": false, "explanation": "Adds cost and still doesn't tell the model your policy — returning null instead of inventing."}], "correct": "B", "task_id": "4.3", "objective": "Enforce structured output using tool use and JSON schemas", "group": "H"}, {"id": "m53", "domain": 4, "scenario": null, "situation": "After implementing tool use with strict schema definitions, JSON syntax errors are eliminated, but 5% of extractions still have valid JSON with empty arrays or null values for required fields like citations and methodology. Spot-checking reveals that source documents contain this information, but in varied formats—inline citations vs. bibliographies, methodology sections vs. details embedded in introductions.", "question": "What's the most effective way to address these failures?", "options": [{"letter": "A", "text": "Implement retry logic that re-sends requests when validation detects empty required fields.", "correct": false, "explanation": "Retry on the same prompt won't suddenly teach the model to recognize methodology embedded in an intro. The extraction capability is the gap."}, {"letter": "B", "text": "Build a regex-based post-processing layer that scans source documents for citation patterns and methodology keywords, populating empty fields when the model fails to extract.", "correct": false, "explanation": "Regex over prose is brittle, misses the nuanced cases (intro-embedded methodology), and ships silently wrong data."}, {"letter": "C", "text": "Modify your schema to make citations and methodology optional, and flag incomplete records for manual review rather than failing validation.", "correct": false, "explanation": "Papers over the symptom. The fields really are required downstream — routing them to humans trades one problem for reviewer workload."}, {"letter": "D", "text": "Add few-shot examples demonstrating extractions from documents with varied structures—showing how to identify citations in different formats and locate methodology details across section types.", "correct": true, "explanation": "Correct. The failure mode is the model not recognizing varied formats. Concrete examples across the format distribution directly raise recall on the 5%."}], "correct": "D", "task_id": "4.2", "objective": "Apply few-shot prompting to improve output consistency and quality", "group": "I"}, {"id": "m54", "domain": 4, "scenario": null, "situation": "Your extraction pipeline processes invoices and extracts line items, subtotals, tax amounts, and grand totals. During evaluation, you discover that in 18% of extractions, the sum of extracted line item amounts doesn't match the extracted grand total—sometimes due to OCR errors in the source document, sometimes due to extraction mistakes by the model. Downstream accounting systems reject records with mismatched totals.", "question": "What's the most effective approach to improve extraction reliability?", "options": [{"letter": "A", "text": "Add a \"`calculated_total`\" field where the model sums extracted line items alongside a \"`stated_total`\" field. Flag records for human review when values differ.", "correct": true, "explanation": "Correct. Capturing both values makes the discrepancy a first-class signal — you catch OCR errors and extraction mistakes uniformly, and you can route only the mismatched 18% to humans."}, {"letter": "B", "text": "Extract line items and totals independently, then use a separate validation model to reconcile discrepancies by determining which extracted values are most likely correct.", "correct": false, "explanation": "A reconciliation model can mask OCR errors by silently picking one number over another — the accounting system still gets wrong data without a flag."}, {"letter": "C", "text": "Add few-shot examples demonstrating invoices where extracted line items sum correctly to the stated total, encouraging the model to produce mathematically consistent extractions.", "correct": false, "explanation": "Can nudge the model toward consistency, but when the source itself is internally inconsistent (OCR errors), the model has to choose one 'wrong'."}, {"letter": "D", "text": "Implement post-processing that automatically adjusts line item amounts proportionally when their sum doesn't match the stated total.", "correct": false, "explanation": "Silent financial data rewrites are dangerous — you'd be fabricating line items for a downstream accounting system."}], "correct": "A", "task_id": "4.2", "objective": "Apply few-shot prompting to improve output consistency and quality", "group": "I"}, {"id": "m56", "domain": 4, "scenario": null, "situation": "Your extraction uses tool use with a JSON schema where `property_type` is defined as an enum: ['house', 'apartment', 'condo', 'townhouse']. After deployment, 8% of extractions fail schema validation. Investigation reveals listings mention many uncommon property types—\"studio\", \"loft\", \"duplex\", \"mobile home\", \"tiny house\", \"converted warehouse\"—and new types continue appearing regularly.", "question": "What's the most effective long-term solution?", "options": [{"letter": "A", "text": "Continuously expand the enum to include newly observed property types and add monitoring for additional edge cases.", "correct": false, "explanation": "Whack-a-mole enum maintenance — every new listing style risks another validation failure."}, {"letter": "B", "text": "Add an \"other\" value to your enum with a separate `property_type_detail` string field for specifics when \"other\" is selected.", "correct": true, "explanation": "Correct. Keeps the strong enum for the common cases (clean downstream joins) while giving a well-typed escape hatch that preserves detail. Stable long-term."}, {"letter": "C", "text": "Change `property_type` from an enum to a free-form string and implement a normalization step in post-processing.", "correct": false, "explanation": "Loses the validation guarantee entirely and pushes the normalization problem downstream with no canonical vocabulary."}, {"letter": "D", "text": "Add few-shot examples to your prompt demonstrating how to map unexpected property types to the closest existing enum value.", "correct": false, "explanation": "Forcing 'tiny house' into 'house' drops real distinctions downstream — you're silently lossy."}], "correct": "B", "task_id": "4.3", "objective": "Enforce structured output using tool use and JSON schemas", "group": "H"}, {"id": "m57", "domain": 4, "scenario": null, "situation": "Your extraction system parses e-commerce product descriptions to extract specifications like dimensions, weight, and materials into JSON. Despite having a well-defined schema, the model inconsistently extracts the \"materials\" field—sometimes returning \"cotton blend\", other times \"Cotton/Polyester mix\", and occasionally omitting the field when material information is clearly present in the source.", "question": "What's the most effective way to improve extraction consistency?", "options": [{"letter": "A", "text": "Make the \"materials\" field required instead of optional in the schema to force the model to always extract a value", "correct": false, "explanation": "Required-ness doesn't solve format inconsistency, and it can force invented values when material info is genuinely missing."}, {"letter": "B", "text": "Switch to a more capable model tier since inconsistent extraction indicates insufficient model capability", "correct": false, "explanation": "Costly, and doesn't address the actual cause — the model doesn't know the canonical output format you want."}, {"letter": "C", "text": "Set temperature to 0 to eliminate randomness and ensure deterministic outputs", "correct": false, "explanation": "Determinism doesn't define the right format; it just makes the wrong-format answers consistent."}, {"letter": "D", "text": "Add few-shot examples showing 2-3 complete input-output pairs with standardized material description formats", "correct": true, "explanation": "Correct. Few-shot examples demonstrate the exact canonical format you want ('cotton/polyester' as a normalized list), and they also raise recall on material info that was being skipped."}], "correct": "D", "task_id": "4.2", "objective": "Apply few-shot prompting to improve output consistency and quality", "group": "I"}, {"id": "m58", "domain": 4, "scenario": null, "situation": "Documents arrive continuously throughout business hours and need structured data extracted. To reduce costs, you want to use the `Message Batches API` (50% discount, up-to-24-hour processing window). Your SLA specifies that extraction results must be available within 30 hours of document arrival with 99.9% reliability.", "question": "Which batching strategy is most appropriate?", "options": [{"letter": "A", "text": "Submit batches every 6 hours containing documents from that window", "correct": false, "explanation": "Worst-case a document waits 6 hours to be batched and then up to 24 hours to process = 30 hours exactly. No safety margin for a 99.9% reliability target."}, {"letter": "B", "text": "Submit a single batch at end of day containing all documents from that day", "correct": false, "explanation": "A doc arriving just after morning startup waits all day to be batched plus up to 24 hours to process — well over 30 hours."}, {"letter": "C", "text": "Submit batches every 4 hours containing documents from that window", "correct": true, "explanation": "Correct. Max 4-hour wait + up to 24-hour batch SLO = 28-hour worst case, leaving a 2-hour cushion under the 30-hour SLA to absorb batch variance and hit 99.9%."}, {"letter": "D", "text": "Use the real-time API for all documents instead of batch processing", "correct": false, "explanation": "Meets SLA trivially but throws away the 50% cost saving the question asks you to capture."}], "correct": "C", "task_id": "4.5", "objective": "Design efficient batch processing strategies", "group": "H"}, {"id": "m59", "domain": 4, "scenario": null, "situation": "After deployment, you find that 12% of extractions contain semantic errors that pass JSON schema validation (e.g., a duration like \"30 minutes\" incorrectly placed in an ingredient quantity field). Human reviewers have capacity to check only 20% of extractions.", "question": "Which approach most effectively allocates reviewer attention?", "options": [{"letter": "A", "text": "Have the model output field-level confidence scores, then calibrate review thresholds using a labeled validation set.", "correct": true, "explanation": "Correct. Field-level confidence lets you route the low-confidence 20% — which is where the 12% semantic errors concentrate — to humans. Calibration makes the threshold choice data-driven."}, {"letter": "B", "text": "Randomly sample 20% of extractions for review, using corrections to track accuracy and identify error patterns.", "correct": false, "explanation": "Great for measurement, poor for coverage — random sampling only catches 20% of the 12% errors (~2.4% total). Confidence-based routing finds far more."}, {"letter": "C", "text": "Prioritize review of all extractions where required fields are empty or explicitly marked as not found.", "correct": false, "explanation": "Empty fields are a narrow error mode. The 12% problem is wrong values in populated fields, which this filter misses."}, {"letter": "D", "text": "Review all extractions from documents with formatting anomalies such as unusual layouts or mixed content types.", "correct": false, "explanation": "A useful heuristic, but formatting anomaly doesn't reliably predict where fields got confused with each other."}], "correct": "A", "task_id": "5.5", "objective": "Design human review workflows and confidence calibration", "group": "C"}, {"id": "f4-001", "domain": 4, "task_id": "4.3", "objective": "Enforce structured output using tool use and JSON schemas", "situation": "A support-automation platform routes incoming tickets to one of several extraction tools depending on ticket type: bug_report, feature_request, or billing_issue. The routing logic currently inspects keywords in the ticket text with regex before choosing which single tool to force via tool_choice, but the regex misclassifies many tickets, leading to the wrong tool being forced.", "question": "What is a better architecture using tool_choice?", "options": [{"letter": "A", "text": "Register all three tools and set tool_choice to {\"type\": \"auto\"} so Claude can choose to skip calling a tool for tickets that seem too ambiguous to classify", "correct": false, "explanation": "\"auto\" would allow Claude to respond with plain text instead of calling any extraction tool, undermining the goal of guaranteeing a structured extraction result for every ticket."}, {"letter": "B", "text": "Keep the regex-based routing exactly as-is, but improve the accuracy of the regex patterns since tool_choice cannot influence which tool Claude ultimately calls", "correct": false, "explanation": "tool_choice does influence which tool can be called; forcing a single tool via regex-based pre-selection is precisely the fragile pattern causing misrouting, and improving regex patterns doesn't address the architectural issue."}, {"letter": "C", "text": "Force each of the three tools in three separate parallel requests per ticket, then keep whichever single response happens to return without an error", "correct": false, "explanation": "Running three forced calls per ticket triples cost and latency unnecessarily and doesn't reliably identify the correct tool, since forcing the wrong tool would still return a valid-looking but incorrect result without an error."}, {"letter": "D", "text": "Register all three extraction tools in the same request and set tool_choice to {\"type\": \"any\"}, letting Claude read the ticket and select the correct tool itself", "correct": true, "explanation": "Setting tool_choice to \"any\" with all three tools registered lets Claude use its understanding of the ticket content to pick the correct extraction tool, removing the fragile regex-based pre-classification step entirely."}], "correct": "D", "select": 1, "group": "H"}, {"id": "f4-002", "domain": 4, "task_id": "4.6", "objective": "Design multi-instance and multi-pass review architectures", "situation": "An architect implements verification passes where the reviewer instance labels each finding 'high confidence' or 'low confidence.' A colleague suggests skipping the independent reviewer and instead having the generator self-report confidence on its own perceived flaws.", "question": "What is the flaw in the colleague's proposal?", "options": [{"letter": "A", "text": "Confidence scoring only works when a single instance reviews a single file, so it cannot be applied across a multi-file change regardless of which instance produces it", "correct": false, "explanation": "Confidence self-reporting is not limited to single-file, single-instance reviews; the flaw here is about which instance reports it, not the scope of files."}, {"letter": "B", "text": "The generator cannot report confidence levels unless it is also given write access to the codebase, which defeats the purpose of a review-only pass", "correct": false, "explanation": "Reporting confidence about a finding has no dependency on write access to the codebase."}, {"letter": "C", "text": "Self-reported confidence from the generator is produced in the same session that generated the code, so it inherits the same blind spots as self-review", "correct": true, "explanation": "Confidence labels generated within the same session as the code still carry the generator's retained reasoning bias, so they don't provide the independence that makes calibrated routing meaningful."}, {"letter": "D", "text": "Confidence scores are not a supported output format for any Claude instance, so the generator could not attach a confidence label to its own findings", "correct": false, "explanation": "Confidence labeling is a prompting pattern, not a format restriction, and any instance can be asked to produce it."}], "correct": "C", "select": 1, "group": "F"}, {"id": "f4-003", "domain": 4, "task_id": "4.5", "objective": "Design efficient batch processing strategies", "situation": "A support team is building a live chat assistant that must answer customers while they are actively typing in a conversation. An engineer proposes routing every chat turn through the Message Batches API to cut token costs.", "question": "What is the most accurate assessment of this proposal?", "options": [{"letter": "A", "text": "It is unsuitable, because live chat needs an immediate response and the Batches API offers no guaranteed turnaround, so answers could arrive far too late", "correct": true, "explanation": "Live chat is a blocking, latency-sensitive workflow; the Batches API can take up to 24 hours with no SLA, making it incompatible with real-time conversation."}, {"letter": "B", "text": "It is suitable, because customers already expect a short delay in chat interfaces, which matches the pacing the Batches API provides", "correct": false, "explanation": "Chat users tolerate seconds of delay, not the unbounded, unpredictable wait the Batches API allows for."}, {"letter": "C", "text": "It is suitable, because the Batches API returns results the moment a request finishes rather than waiting for a fixed processing window to elapse", "correct": false, "explanation": "While individual batch requests can finish quickly under light load, there is no guarantee of this, and demand spikes can push completion out for hours."}, {"letter": "D", "text": "It is unsuitable, because the Batches API charges a per-token premium above standard synchronous pricing that live chat's request volume cannot justify", "correct": false, "explanation": "The Batches API is priced at a 50% discount versus standard rates, not a premium; the mismatch here is about latency, not cost."}], "correct": "A", "select": 1, "group": "H"}, {"id": "f4-004", "domain": 4, "task_id": "4.5", "objective": "Design efficient batch processing strategies", "situation": "A data pipeline promises stakeholders that classification results for any document will be delivered no later than 48 hours after ingestion. The team plans to route documents through the Message Batches API, which can take up to 24 hours to finish a batch.", "question": "What is the longest interval the team can wait between batch submission cycles while still meeting the 48-hour promise?", "options": [{"letter": "A", "text": "24 hours between submissions, since a document waiting the full interval still finishes within the 24-hour batch window before the 48-hour deadline", "correct": true, "explanation": "According to the Message Batches API documentation, batch processing can take up to 24 hours to complete. To satisfy a 48‑hour delivery promise, the worst‑case wait (the full interval) plus the maximum processing time must not exceed 48 hours. Setting the submission interval to 24 hours gives 24 + 24 = 48 hours, which is the longest interval that still meets the commitment."}, {"letter": "B", "text": "6 hours between submissions, since frequent small batches always complete faster than the standard 24-hour processing window allows", "correct": false, "explanation": "Batch processing time is independent of submission frequency; the API can take up to 24 hours regardless of batch size or frequency. A 6‑hour interval is even shorter than necessary and does not represent the longest interval that meets the 48‑hour deadline."}, {"letter": "C", "text": "48 hours between submissions, since the deadline itself defines how rarely the pipeline needs to kick off a new batch cycle", "correct": false, "explanation": "With a 48‑hour interval, a document that arrives immediately after a submission could wait 48 hours before being batched, then face up to 24 hours of processing. This 72‑hour worst case violates the 48‑hour promise."}, {"letter": "D", "text": "12 hours between submissions, since halving the processing window doubles the safety margin the pipeline needs against demand spikes", "correct": false, "explanation": "A 12‑hour interval would result in a worst‑case delivery time of 36 hours (12 + 24), which is well within the promise. However, the question asks for the longest allowable interval, and 24 hours is feasible, so 12 hours is not the correct answer."}], "correct": "A", "select": 1, "group": "H"}, {"id": "f4-005", "domain": 4, "task_id": "4.2", "objective": "Apply few-shot prompting to improve output consistency and quality", "situation": "An extraction pipeline for supplier contracts frequently returns null for the 'renewal_notice_period' field on contracts where that information is present but phrased unusually, such as buried in a sentence about termination rather than in a clearly labeled 'Renewal' clause. The team has already tried making the field's instruction more explicit with no improvement.", "question": "What should they try next?", "options": [{"letter": "A", "text": "Show an extraction from unusual phrasing plus a case confirming null is correct when unspecified", "correct": true, "explanation": "Examples that demonstrate a correct extraction from unusually phrased text, paired with an example confirming when null is the right answer, directly address empty or null extraction of a required field by showing the model both what to look for and when abstaining is correct, which instructions alone had already failed to convey."}, {"letter": "B", "text": "Instruct the model to scan only clauses whose heading explicitly contains the word 'renewal' or 'termination'", "correct": false, "explanation": "Restricting the scan to headings containing 'renewal' is the same heading-dependent behavior that is already causing the failure, since the scenario states the information is buried outside of a labeled clause."}, {"letter": "C", "text": "Add a fallback default value of thirty days that is used whenever the field would otherwise be left null", "correct": false, "explanation": "Substituting a default value when extraction fails would silently fabricate contract terms that were never actually specified, which is the kind of hallucination the team should be trying to avoid, not accept as a fallback."}, {"letter": "D", "text": "Change the field's data type from a string to a required enumerated value from a fixed set", "correct": false, "explanation": "Restricting the field to a fixed enumeration does not help the model locate the value in unusually phrased text in the first place, and it could force an inaccurate value onto contracts whose actual terms don't match any allowed option."}], "correct": "A", "select": 1, "group": "I"}, {"id": "f4-006", "domain": 4, "task_id": "4.5", "objective": "Design efficient batch processing strategies", "situation": "Before submitting 80,000 support tickets to a batch job that extracts structured fields from each one, a team wants to reduce the odds of an expensive resubmission cycle caused by a poorly tuned prompt.", "question": "What is the most effective step to take first?", "options": [{"letter": "A", "text": "Run the extraction prompt synchronously against a small, representative sample of tickets, refine it until output quality is high, then submit the full 80,000-ticket batch", "correct": true, "explanation": "Validating and refining the prompt on a small sample first surfaces formatting or instruction issues cheaply, before they propagate across tens of thousands of billed batch requests."}, {"letter": "B", "text": "Submit the full 80,000-ticket batch immediately, since any formatting issues can be caught and corrected once the batch results come back", "correct": false, "explanation": "Discovering prompt issues only after the full batch completes means paying to reprocess a large fraction of 80,000 requests, which the sample-first approach is meant to avoid."}, {"letter": "C", "text": "Split the 80,000 tickets into two batches of equal size submitted back to back, since smaller batches are inherently less likely to contain formatting errors", "correct": false, "explanation": "Batch size does not influence the accuracy of the extraction prompt; splitting the same untested prompt across two batches still risks the same failure rate in each."}, {"letter": "D", "text": "Increase max_tokens across the entire 80,000-ticket batch so that longer completions leave less room for the extraction format to be cut off", "correct": false, "explanation": "Raising max_tokens does not address whether the extraction instructions themselves are well-tuned, and does nothing to reduce the risk of systemic formatting failures."}], "correct": "A", "select": 1, "group": "H"}, {"id": "f4-007", "domain": 4, "task_id": "4.3", "objective": "Enforce structured output using tool use and JSON schemas", "situation": "A financial services firm uses a tool-based schema to extract transaction records from PDF statements, including a required amount field and a required transaction_type field. An audit later finds several instances where deposits were mislabeled as withdrawals, even though every amount and transaction_type field is present and passes schema validation.", "question": "What does this best illustrate about JSON schema enforcement via tool use?", "options": [{"letter": "A", "text": "This indicates that the input_schema was likely malformed, since a correctly constructed JSON schema can apply constraints that verify the relationship between amount and transaction_type, preventing such semantic mislabeling.", "correct": false, "explanation": "JSON Schema is designed to validate structure and types, not the factual correctness of values; even a perfectly constructed schema cannot enforce that transaction_type matches the semantic content derived from the document, so such mislabeling would still occur."}, {"letter": "B", "text": "The mislabeling proves that without a tool_choice that forces a specific tool invocation, the model may output transaction_type labels that diverge from the source document's content, leading to semantic inaccuracies in the extracted data.", "correct": false, "explanation": "tool_choice ensures that the model invokes the specified tool, but it does not guarantee that the returned field values are semantically accurate to the source document; the mislabeling is a semantic error independent of tool invocation behavior."}, {"letter": "C", "text": "Schema validation confirms that required fields exist and have the correct types but it cannot verify that the values are semantically correct relative to the source document, so mislabeling errors like this can still occur.", "correct": true, "explanation": "JSON schema validation confirms that required fields exist and have the correct types, but it cannot assess whether the values are semantically accurate relative to the source document; thus, mislabeling errors like deposits as withdrawals can still pass validation."}, {"letter": "D", "text": "This is expected only when strict mode is disabled, because strict mode enforces that field values must satisfy the schema's semantic rules, such as ensuring transaction_type accurately reflects the amount sign, thereby preventing mislabeling.", "correct": false, "explanation": "Strict mode enforces that the tool call output adheres strictly to the schema structure (types, required fields, no extra properties), but it does not add semantic validation such as ensuring transaction_type reflects the amount sign; thus, even with strict mode, mislabeling can occur."}], "correct": "C", "select": 1, "group": "H"}, {"id": "f4-008", "domain": 4, "task_id": "4.6", "objective": "Design multi-instance and multi-pass review architectures", "situation": "A function signature changes in module A, and callers in modules B and C are updated inconsistently, causing a runtime mismatch. Per-file passes each judged their own file correct in isolation and missed the mismatch.", "question": "Which review step should have caught this, and why?", "options": [{"letter": "A", "text": "The self-review step performed by the generator, since it still remembers why it changed the signature and can check callers from memory", "correct": false, "explanation": "The generator's self-review carries the same retained-reasoning limitation and is not the mechanism designed to catch cross-file inconsistencies."}, {"letter": "B", "text": "A second per-file pass over module A alone, since re-reading the same file twice increases the chance of noticing the signature change", "correct": false, "explanation": "Repeating a per-file pass on module A alone still never examines how modules B and C actually call it."}, {"letter": "C", "text": "The cross-file integration pass, because it traces data flow and interface usage across files rather than analyzing each file alone", "correct": true, "explanation": "An integration pass is designed to trace how data and interfaces flow across files, which is exactly the kind of mismatch a per-file pass, scoped to one file at a time, cannot see."}, {"letter": "D", "text": "A longer per-file pass that reads both module A and module B together in one sitting rather than as two separate local passes", "correct": false, "explanation": "Ad hoc combining of two files into one longer per-file pass isn't the designed cross-file mechanism and doesn't scale to the full callers set including module C."}], "correct": "C", "select": 1, "group": "F"}, {"id": "f4-009", "domain": 4, "task_id": "4.3", "objective": "Enforce structured output using tool use and JSON schemas", "situation": "A compliance team needs an audit trail proving that every processed document produced a structured tool call, with no possibility of the model instead returning a plain-text response that silently skips extraction. The current implementation uses tool_choice: {\"type\": \"auto\"} with a single extract_record tool, and spot checks reveal some documents produced only a text response with no tool_use block at all.", "question": "What change directly fixes this compliance gap?", "options": [{"letter": "A", "text": "Add a stronger system-prompt warning that instructs Claude to always invoke the extract_record tool, emphasizing that plain-text responses violate compliance, while leaving tool_choice set to auto.", "correct": false, "explanation": "Adding a stronger system prompt warning is not a guaranteed fix; even with explicit instructions, tool_choice: \"auto\" still allows the model to output plain text, so compliance is not ensured."}, {"letter": "B", "text": "Increase max_tokens on the request to a high value such as 4096 so Claude has enough room to always complete the extract_record tool call instead of truncating early with only plain text.", "correct": false, "explanation": "Increasing max_tokens only affects the total response length, not whether the model invokes a tool. Under tool_choice: \"auto\", the model can still return plain text regardless of the token limit."}, {"letter": "C", "text": "Switch the extract_record tool's input_schema fields such as record_id and content from optional to required so the model must invoke the tool to provide values for every document.", "correct": false, "explanation": "Making fields required in the tool's input schema only applies when the tool is used; it does not compel the model to invoke the tool when tool_choice is \"auto\", so the root compliance gap remains."}, {"letter": "D", "text": "Change tool_choice to {\"type\": \"any\"} (or force the specific tool by name) so Claude must always invoke a tool on every request instead of being able to respond with plain text.", "correct": true, "explanation": "Changing tool_choice to {\"type\": \"any\"} (or forcing a specific tool by name) guarantees that every response includes a tool call; with \"auto\" the model may still respond with plain text, causing the compliance gap."}], "correct": "D", "select": 1, "group": "H"}, {"id": "f4-010", "domain": 4, "task_id": "4.2", "objective": "Apply few-shot prompting to improve output consistency and quality", "situation": "A team is adding few-shot examples to a ticket-triage prompt to fix inconsistent priority assignments. They have dozens of historical tickets available and are deciding how many to include and how to select them.", "question": "Which approach best follows effective few-shot practice while avoiding new failure modes?", "options": [{"letter": "A", "text": "Select only the single ticket that was hardest to triage historically as the key lesson", "correct": false, "explanation": "A single example, even a hard one, cannot demonstrate the range of triage decisions the model needs to generalize across, and it provides no contrast between different plausible outcomes."}, {"letter": "B", "text": "Include as many historical tickets as the context window allows for maximum reliability", "correct": false, "explanation": "Simply maximizing the number of examples risks the model picking up on unintended coincidental patterns across many similar examples, and it increases prompt length and cost without a guaranteed accuracy benefit beyond a well-chosen smaller set."}, {"letter": "C", "text": "Select three to five diverse examples covering distinct edge cases relevant to triage", "correct": true, "explanation": "A small number of examples chosen to be relevant to the task and diverse across the edge cases that matter is the recommended approach: enough examples to demonstrate the pattern, without so many similar ones that the model latches onto an unintended, superficial pattern in the example set."}, {"letter": "D", "text": "Select examples entirely at random from the archive to avoid any selection bias", "correct": false, "explanation": "Random selection risks pulling several near-duplicate tickets by chance and missing important edge cases entirely, whereas deliberately choosing for relevance and diversity better covers the ambiguous cases that matter."}], "correct": "C", "select": 1, "group": "I"}, {"id": "f4-011", "domain": 4, "task_id": "4.6", "objective": "Design multi-instance and multi-pass review architectures", "situation": "A team built a pipeline where the same Claude instance that generates code is then asked, within the same conversation, to 'review your own work for bugs before finishing.' QA later finds subtle issues the model missed during that self-review step.", "question": "Which architectural change is most effective at catching those issues going forward?", "options": [{"letter": "A", "text": "Increase the extended thinking budget for the self-review step so the model reasons longer before finishing", "correct": false, "explanation": "More reasoning tokens inside the same session still operate on top of the reasoning that already justified the original decisions, so the underlying self-review limitation persists."}, {"letter": "B", "text": "Spawn a second, independent Claude instance with no access to the generation session's history to review the code fresh", "correct": true, "explanation": "A fresh instance with no prior reasoning context is not anchored to the generator's original decisions, making it more likely to catch subtle issues than self-review."}, {"letter": "C", "text": "Add stricter self-review instructions to the system prompt so the model scrutinizes its own prior decisions more carefully", "correct": false, "explanation": "Stricter wording in the system prompt doesn't remove the session's retained reasoning context, which is the actual cause of missed self-review findings."}, {"letter": "D", "text": "Ask the same instance to review the code twice in a row within the same session before returning results", "correct": false, "explanation": "Repeating the review within the same session still carries forward the same reasoning that produced the code, so the blind spot remains."}], "correct": "B", "select": 1, "group": "F"}, {"id": "f4-012", "domain": 4, "task_id": "4.3", "objective": "Enforce structured output using tool use and JSON schemas", "situation": "A ticket-classification schema has a category field defined as an enum of five known categories: billing, technical, shipping, account, and refund. After deployment, roughly 8% of real tickets don't cleanly fit any of these categories, and the model is observed forcing them into the closest (often incorrect) enum value.", "question": "Which schema change best addresses this?", "options": [{"letter": "A", "text": "Add a sixth hardcoded category called miscellaneous_unclear_ticket_type_pending_manual_review to the existing enum list", "correct": false, "explanation": "A single rigid extra enum value doesn't provide space to capture the actual nuance of each ambiguous ticket, and the six categories are still a closed, non-extensible set with no detail field."}, {"letter": "B", "text": "Lower the required strictness of the tool by setting strict: false so the model is allowed to skip the category field for edge cases", "correct": false, "explanation": "Making the field skippable would produce missing classifications rather than accurate ones, and does nothing to help distinguish or record what the ambiguous tickets actually are."}, {"letter": "C", "text": "Remove the enum constraint entirely and let category be an unconstrained free-text string so the model can write anything that seems to fit", "correct": false, "explanation": "Removing the enum sacrifices the consistency and downstream queryability that enums provide, and reintroduces the risk of inconsistent free-text values across many tickets."}, {"letter": "D", "text": "Add an \"other\" enum value alongside a separate free-text detail field so ambiguous tickets can be captured without corrupting the five known categories", "correct": true, "explanation": "The \"other\" + detail string pattern lets extensible or ambiguous categories be captured accurately without forcing a misclassification into one of the fixed enum values, while still keeping downstream processing structured."}], "correct": "D", "select": 1, "group": "H"}, {"id": "f4-013", "domain": 4, "task_id": "4.1", "objective": "Design prompts with explicit criteria to improve precision and reduce false positives", "situation": "A code-review assistant reports findings in three categories: security, correctness, and style. The style category has a 60% false positive rate while security and correctness are both above 90% precision. Developers say they now distrust every finding the tool produces, including the security ones.", "question": "Which explanation best accounts for this reaction?", "options": [{"letter": "A", "text": "The tool's overall accuracy score is mathematically dominated by the style category, so the reported precision numbers for security are actually inflated, eroding trust in the tool's accuracy.", "correct": false, "explanation": "The precision of security and correctness categories is reported independently, so the style category's performance does not mathematically inflate their precision figures. The erosion of trust is behavioral, not due to distorted statistical reporting."}, {"letter": "B", "text": "Security and correctness findings are inherently harder to verify than style findings, so developers assume the reported precision figures cannot be trusted, since they are often harder to verify.", "correct": false, "explanation": "The scenario states security and correctness findings both exceed 90% precision, so verification difficulty is not the issue. The described distrust stems from the high noise in the style category spilling over to overall tool perception, not from inherent verification challenges."}, {"letter": "C", "text": "Developers are miscounting the false positive rate because they are including security findings that were later fixed, leading them to distrust the tool's accuracy, even for security findings.", "correct": false, "explanation": "Nothing in the scenario indicates that developers are miscounting false positives or including later-fixed security findings in that count. The reaction is a perceptual spillover effect, not a measurement error."}, {"letter": "D", "text": "A single category with a high false positive rate can undermine confidence in the tool's accurate categories, because developers experience all findings as coming from one undifferentiated source.", "correct": true, "explanation": "Trust in automated reviewers is often holistic; when one category produces many false positives, developers generalize that unreliability across all output, undifferentiating between categories. Thus, a single noisy category can poison confidence in even highly accurate ones."}], "correct": "D", "select": 1, "group": "I"}, {"id": "f4-014", "domain": 4, "task_id": "4.5", "objective": "Design efficient batch processing strategies", "situation": "An agent workflow needs Claude to request a database-lookup tool, receive the tool's result, and then reason over that result before producing a final answer, all within one logical exchange. A developer wants to run this exchange through the Message Batches API to save on cost.", "question": "What is the key limitation that rules this out?", "options": [{"letter": "A", "text": "Batch requests limit each conversation to a single message, so a tool_use block and its follow-up reasoning cannot appear in one batched exchange, since each conversation must be self-contained.", "correct": false, "explanation": "Batch requests can include multiple messages in a conversation, including tool_use and tool_result blocks from prior turns. The constraint is not a single message per conversation, but rather the lack of a live interaction to provide a tool result within a single request."}, {"letter": "B", "text": "A single batch request cannot pause mid-processing to accept an application-supplied tool result, since each request resolves independently with no mid-request round trip.", "correct": true, "explanation": "Each batch request is processed independently and cannot be paused mid-processing for the application to provide a tool result. The model must complete its output in one shot; there is no round-trip within a request."}, {"letter": "C", "text": "The Message Batches API silently strips tool_use content blocks from responses, so the application never learns which tool the model wanted to call, leaving it unable to supply the required result.", "correct": false, "explanation": "The Message Batches API does not strip tool_use content blocks; they are returned in the response like normal. Thus, the application can see which tool was requested, but cannot supply the result within the same batch request."}, {"letter": "D", "text": "Tool definitions cannot be attached to any request submitted through the Message Batches API, so the model never has the option to request a database-lookup tool during processing.", "correct": false, "explanation": "Tool definitions can indeed be attached to requests in the Message Batches API, so the model can request tools. The limitation is not the absence of tool support but the inability to feed back tool results mid-request."}], "correct": "B", "select": 1, "group": "H"}, {"id": "f4-015", "domain": 4, "task_id": "4.4", "objective": "Implement validation, retry, and feedback loops for extraction quality", "situation": "A reimbursement pipeline extracts a required \"project_code\" field from scanned expense reports. On one submission, the employee left the project code blank on the physical form, so it does not appear anywhere in the scanned image. After three retries with error feedback, the field is still empty.", "question": "What should the pipeline do next?", "options": [{"letter": "A", "text": "Switch to a larger model and resend the identical prompt, expecting greater capability to recover the missing value", "correct": false, "explanation": "A more capable model still cannot extract a value that is absent from its input; model capability does not substitute for missing source data."}, {"letter": "B", "text": "Increase the sampling temperature on every retry so the model becomes more likely to locate the missing value in the scan", "correct": false, "explanation": "Raising temperature increases output variance but does not create information that does not exist in the source document; it may even introduce a fabricated value."}, {"letter": "C", "text": "Stop retrying and route the record to a human, since the code is genuinely absent from the source rather than a structural failure", "correct": true, "explanation": "Retries only help when the failure is a format or structural mismatch; when the information simply is not present in the source, no amount of retrying can recover it, so the correct action is to stop and escalate for human resolution."}, {"letter": "D", "text": "Keep retrying while adding progressively more detailed schema instructions about how the project_code field should be structured and formatted", "correct": false, "explanation": "The failure is not a schema or instruction-clarity problem, so further schema-focused instructions cannot surface data that was never captured on the form."}], "correct": "C", "select": 1, "group": "I"}, {"id": "f4-016", "domain": 4, "task_id": "4.5", "objective": "Design efficient batch processing strategies", "situation": "A compliance group reviews 15,000 vendor contracts once a week and publishes a findings report two business days later. Cost per document matters because the review runs across the entire vendor catalog every cycle.", "question": "How should this recurring job be built?", "options": [{"letter": "A", "text": "Submit the contracts through the synchronous Messages API one at a time, since compliance findings require the strict ordering only sequential calls preserve", "correct": false, "explanation": "Result ordering is not guaranteed by either API in a way that requires sequential calls; documents are correlated by custom_id, not by submission order."}, {"letter": "B", "text": "Submit the contracts as a single Message Batch each week, since the two-day turnaround comfortably absorbs the batch processing window and the discount lowers per-cycle spend", "correct": true, "explanation": "A weekly, non-blocking audit with a multi-day turnaround tolerates the batch processing window comfortably while capturing the 50% cost reduction across the full catalog."}, {"letter": "C", "text": "Split the contracts across several small Message Batches submitted every few minutes, since batches must stay under a few hundred requests to process reliably", "correct": false, "explanation": "A single batch can hold up to 100,000 requests, so fragmenting 15,000 contracts into many small submissions adds overhead for no benefit."}, {"letter": "D", "text": "Submit the contracts through the synchronous Messages API in parallel threads, since parallel synchronous calls always finish faster than a queued batch job", "correct": false, "explanation": "Parallel synchronous calls bypass the batch discount entirely and add operational complexity without a latency requirement that demands it."}], "correct": "B", "select": 1, "group": "H"}, {"id": "f4-017", "domain": 4, "task_id": "4.4", "objective": "Implement validation, retry, and feedback loops for extraction quality", "situation": "A medical-intake extractor nests the patient's \"date_of_birth\" field under the wrong parent object in its structured output, causing schema validation to fail. The team wants the next attempt to self-correct the nesting.", "question": "What should the follow-up request include?", "options": [{"letter": "A", "text": "A brand-new prompt that redescribes the schema from scratch, sent without the intake document originally supplied.", "correct": false, "explanation": "A new prompt that only restates the schema does not provide the original source document or the prior failed JSON. The retry-with-error-feedback approach relies on showing the model what it originally extracted and the specific structural error, rather than starting over without the source data."}, {"letter": "B", "text": "The original intake document, the prior failed JSON, and a note that date_of_birth belongs under patient rather than guardian.", "correct": true, "explanation": "Anthropic documentation for validation and retry loops says retry with error feedback should include three pieces: the original document, the failed extraction, and the specific validation error. Including the original intake document, the prior failed JSON, and the explicit instruction to move date_of_birth to patient matches the recommended pattern for misplaced values in structured extraction."}, {"letter": "C", "text": "The validation error code by itself, assuming the model retains full memory of the document across separate turns.", "correct": false, "explanation": "Sending only a validation error code assumes persistent memory across turns, but separate API requests are stateless. The documented retry pattern requires resending the original document, the failed extraction, and the specific validation error so the model has the necessary context to correct the nesting."}, {"letter": "D", "text": "Only the prior failed JSON and a note instructing the model to move date_of_birth from guardian to patient.", "correct": false, "explanation": "This omits the original intake document, which is a required element of the retry-with-error-feedback pattern in the Claude Certification Guide for Subdomain 4.4. Without the source document, the model may not have enough context to verify or correctly re-extract the misplaced value."}], "correct": "B", "select": 1, "group": "I"}, {"id": "f4-018", "domain": 4, "task_id": "4.2", "objective": "Apply few-shot prompting to improve output consistency and quality", "situation": "A content-moderation prompt must decide whether borderline posts (e.g., dark sarcasm about a sensitive topic) should be escalated for human review or allowed to stand. The instructions describe general moderation policy, but the model's escalation decisions on borderline posts are inconsistent between similar sessions. The team wants to add a small number of examples that will generalize well.", "question": "Which set of examples would be most effective?", "options": [{"letter": "A", "text": "Two to four borderline posts, each paired with the escalation decision and a brief rationale for it", "correct": true, "explanation": "A handful of targeted examples that show genuinely ambiguous posts alongside the decision and the reasoning for choosing one path over the other plausible one is exactly the technique for demonstrating ambiguous-case handling, and the added rationale helps the judgment generalize to new borderline posts."}, {"letter": "B", "text": "One example of the single most extreme violation the team has ever seen, as a strong anchor", "correct": false, "explanation": "A single extreme example is not representative of the borderline gray area the model struggles with, so it provides little guidance for the more subtle judgment calls the team actually needs help with."}, {"letter": "C", "text": "A restated summary of the moderation policy, broken into a numbered checklist instead of paragraphs", "correct": false, "explanation": "A checklist restatement is still an instruction, not a demonstrated example, and the scenario states that instructions describing policy have already failed to produce consistent escalation decisions."}, {"letter": "D", "text": "Ten or more posts that are obviously fine, giving the model abundant precedent for the common case", "correct": false, "explanation": "Examples of clearly non-borderline posts don't address the actual failure mode, which is inconsistency on borderline cases; the easy cases already work, so this set wastes the few-shot budget on cases that were never the problem."}], "correct": "A", "select": 1, "group": "I"}, {"id": "f4-019", "domain": 4, "task_id": "4.2", "objective": "Apply few-shot prompting to improve output consistency and quality", "situation": "A team adds five few-shot examples to a document-classification prompt to fix inconsistent labeling. All five examples happen to be English-language emails under 100 words. After deployment, the model performs well on similar short English emails but starts mislabeling longer documents and documents in other languages that it previously handled correctly under the old, example-free prompt.", "question": "What is the most likely cause, and what should the team do?", "options": [{"letter": "A", "text": "The regression is unrelated to the examples and is caused by unrelated model drift", "correct": false, "explanation": "The scenario describes a change coinciding directly with adding a narrow example set, and unintentional pattern-learning from unrepresentative examples is a well-documented risk, making it a far more likely explanation than unrelated drift."}, {"letter": "B", "text": "The examples unintentionally taught an unrelated pattern tied to length and language; diversify them", "correct": true, "explanation": "When every example shares an incidental trait like language or length, the model can pick up on that superficial trait as if it were part of the target pattern; the fix is to make the example set diverse across the dimensions that actually vary in production, not just the dimension the team was trying to fix."}, {"letter": "C", "text": "Few-shot examples are simply incompatible with document classification tasks, so remove them entirely", "correct": false, "explanation": "Few-shot examples are a standard and effective technique for classification tasks; the failure here is attributable to an unrepresentative example set, not to the technique being unsuitable for the task."}, {"letter": "D", "text": "The five examples are too few in number; keep them but duplicate each one three times", "correct": false, "explanation": "Duplicating the same narrow examples would reinforce the unintended pattern rather than correct it, since the underlying problem is lack of diversity, not lack of repetition."}], "correct": "B", "select": 1, "group": "I"}, {"id": "f4-020", "domain": 4, "task_id": "4.3", "objective": "Enforce structured output using tool use and JSON schemas", "situation": "An architect is reviewing a schema-enforced invoice extraction tool. The tool_use response consistently returns syntactically valid JSON with the correct field types, yet a downstream finance audit finds that individual line-item amounts frequently fail to sum to the reported invoice total.", "question": "What should the architect conclude about the current design?", "options": [{"letter": "A", "text": "The schema must be missing a required field constraint, since required fields are the only mechanism that can prevent numeric mismatches between line items and totals", "correct": false, "explanation": "Required fields only ensure a field is present, not that its value satisfies an arithmetic relationship with other fields; JSON Schema has no native way to express \"this field must equal the sum of these other fields.\""}, {"letter": "B", "text": "The model is very likely hallucinating the schema itself at request time, so the input_schema needs to be resent with every single follow-up message", "correct": false, "explanation": "The schema is provided in the tools parameter of the request and is not something the model can lose or forget between calls; the observed issue is a semantic error, not a missing schema."}, {"letter": "C", "text": "The JSON schema enforced by tool use guarantees syntactic validity but does not verify semantic correctness such as arithmetic consistency between related fields", "correct": true, "explanation": "Strict JSON schemas via tool use eliminate syntax errors and enforce types, but they cannot enforce cross-field semantic relationships like a sum constraint, so validation logic for such checks must live outside the schema."}, {"letter": "D", "text": "The tool_choice must be set to auto instead of a forced tool, since forced tool calls are known to skip internal consistency checks on numeric fields", "correct": false, "explanation": "tool_choice controls whether and which tool is called, not whether the model performs internal arithmetic validation; switching to auto would not fix or explain the totals mismatch."}], "correct": "C", "select": 1, "group": "H"}, {"id": "f4-021", "domain": 4, "task_id": "4.1", "objective": "Design prompts with explicit criteria to improve precision and reduce false positives", "situation": "An architect is comparing two candidate prompts for a security-findings category before choosing one for production.", "question": "Prompt X says: \"Flag anything that looks like it could be a security risk.\" Prompt Y says: \"Flag code that writes user-supplied input directly into a SQL query string without parameterization, or that stores a plaintext password.\" Which statement correctly evaluates the two prompts with respect to reducing false positives?", "options": [{"letter": "A", "text": "The two prompts are functionally equivalent, as both ultimately rely on the model's general security knowledge to determine what constitutes a risk, making the specific wording irrelevant to false-positive rates.", "correct": false, "explanation": "The prompts are not equivalent: Prompt Y supplies explicit, checkable criteria that reduce ambiguity, while Prompt X provides no such guidance, making specific wording highly relevant to false-positive rates."}, {"letter": "B", "text": "Prompt Y is worse because its specific examples of SQL injection and plaintext passwords will cause the model to focus on those patterns and miss other risks, leading to more false negatives.", "correct": false, "explanation": "Listing specific conditions does not inherently cause more false negatives; a well-designed prompt can include additional categories as needed. The focus here is on reducing false positives, which Prompt Y achieves through its explicit examples."}, {"letter": "C", "text": "Prompt Y is preferable because it names specific, checkable conditions that define a security issue, while Prompt X relies on an open-ended judgment about what 'looks like' a risk.", "correct": true, "explanation": "Prompt Y names specific, checkable conditions (SQL injection and plaintext passwords) that define security issues clearly, whereas Prompt X uses vague language that invites open-ended judgment. This precision reduces false positives by providing clear criteria for the model to follow."}, {"letter": "D", "text": "Prompt X is preferable because its broad phrasing allows the model to capture a wider range of security threats, including subtle logic flaws and misconfigurations that Prompt Y's specific list might miss.", "correct": false, "explanation": "Broad phrasing like 'anything that looks like a security risk' tends to increase false positives by flagging superficially risky-looking but benign code. Specific criteria, as in Prompt Y, better constrain the model to actual, checkable issues."}], "correct": "C", "select": 1, "group": "I"}, {"id": "f4-022", "domain": 4, "task_id": "4.1", "objective": "Design prompts with explicit criteria to improve precision and reduce false positives", "situation": "An architect wants a Claude-based reviewer to classify findings into severity levels consistently across many pull requests and multiple reviewers on the team.", "question": "Which prompt design best achieves consistent severity classification?", "options": [{"letter": "A", "text": "Define each severity level with a short description plus a concrete code example illustrating what qualifies at that level, so the model has a consistent reference point for every classification.", "correct": true, "explanation": "Providing a short description and a concrete code example for each severity level gives the model an unambiguous, checkable anchor for classification. This fixed reference point promotes consistent severity labels across different pull requests and multiple reviewers."}, {"letter": "B", "text": "Tell the model to default to medium severity for every finding, escalating or de-escalating only when it provides specific technical justification referencing the team's predefined severity criteria.", "correct": false, "explanation": "Defaulting all findings to medium severity and requiring justification to change biases classifications toward that default. This does not provide a reliable method for correctly identifying true severity, as it relies on a threshold that may not align with the actual impact of findings."}, {"letter": "C", "text": "Instruct the model to assign severity based on how urgent the issue feels in the context of the specific pull request, considering the component's criticality, recent commit history, and related incidents.", "correct": false, "explanation": "'How urgent it feels' is a subjective judgment with no fixed reference point, which leads to inconsistent severity labels across different reviews and reviewers. Relying on contextual factors like component criticality and commit history without explicit criteria introduces variability, not consistency."}, {"letter": "D", "text": "Ask the model to compare the current finding to the average severity of findings from the last ten pull requests, using that historical baseline to normalize severity assignments across reviews.", "correct": false, "explanation": "Normalizing against the average severity of the last ten pull requests makes classifications relative to a shifting baseline, not an absolute standard. This approach can cause severity drift over time and fails to ensure uniform application across different review periods."}], "correct": "A", "select": 1, "group": "I"}, {"id": "f4-023", "domain": 4, "task_id": "4.6", "objective": "Design multi-instance and multi-pass review architectures", "situation": "A team runs per-file passes first, producing local findings for each changed file, then runs a separate integration pass over the whole changeset.", "question": "What should the integration pass focus on that the per-file passes are not well suited to catch?", "options": [{"letter": "A", "text": "Syntax errors within an individual file, since a per-file pass already checks whether that specific file compiles and contains valid syntax", "correct": false, "explanation": "Syntax errors within one file are local issues a per-file pass can detect directly; they don't require cross-file integration analysis."}, {"letter": "B", "text": "Duplicate logic within a single file, since per-file passes read files sequentially and cannot notice repeated code blocks in the same file", "correct": false, "explanation": "Duplicate logic within a single file is a local concern a per-file pass can notice while reading that file; it doesn't require the cross-file integration pass."}, {"letter": "C", "text": "Formatting and style issues within a single file, since per-file passes are too narrowly scoped to catch indentation or naming inconsistencies", "correct": false, "explanation": "Formatting and style within a single file are exactly the kind of local issue a per-file pass is well suited to catch, not the integration pass's purpose."}, {"letter": "D", "text": "Inconsistencies in data flow and contracts between files, such as a shared interface used differently across files reviewed in isolation", "correct": true, "explanation": "The integration pass exists specifically to catch cross-file inconsistencies in data flow and shared contracts, which per-file passes structurally cannot see since each is scoped to one file."}], "correct": "D", "select": 1, "group": "F"}, {"id": "f4-024", "domain": 4, "task_id": "4.3", "objective": "Enforce structured output using tool use and JSON schemas", "situation": "An engineering team is deciding whether to expose one single extract_document tool with a very large schema covering invoices, receipts, and purchase orders in one combined structure, or three separate smaller tools (extract_invoice, extract_receipt, extract_purchase_order) selected via tool_choice: \"any\" based on document content. Users upload one document at a time and document type varies per upload.", "question": "Which design better matches the intended use of tool_choice: \"any\" for extraction?", "options": [{"letter": "A", "text": "Three separate, document-type-specific tools with tool_choice: \"any\", so Claude selects the schema matching the actual document avoiding the noise and confusion of an oversized combined schema.", "correct": true, "explanation": "Using three separate tools with tool_choice: \"any\" allows Claude to dynamically select the schema that matches the document type, avoiding the noise and confusion of an oversized combined schema. This approach leverages tool_choice: \"any\" as intended: when multiple extraction schemas exist and the document type varies, Claude picks the most appropriate tool."}, {"letter": "B", "text": "One combined extraction tool with a single schema for all document types, because tool_choice: \"any\" requires exactly one tool to be registered in the tools array for the model to invoke the extraction logic correctly and clearly.", "correct": false, "explanation": "tool_choice: \"any\" works with any number of registered tools and specifically enables Claude to choose among multiple tools, so it does not require exactly one tool. The assumption that a single tool is required is a misunderstanding of the feature."}, {"letter": "C", "text": "Three separate, document-type-specific tools, but with tool_choice: \"any\" forced to extract_invoice as the default since invoices dominate uploads, ensuring the tool always handles the most common case correctly and reliably.", "correct": false, "explanation": "Forcing a default tool like extract_invoice would misclassify receipts and purchase orders, mapping them into an invoice-shaped schema, which tool_choice: \"any\" is designed to avoid when document type varies. This would reduce accuracy for non-invoice documents."}, {"letter": "D", "text": "One combined extraction tool with a unified schema for invoices, receipts, and purchase orders, where tool_choice: \"any\" selects that single tool and ensures consistent field naming across all document types without schema conflicts.", "correct": false, "explanation": "A single combined schema for fundamentally different document types leads to a sprawling structure with many irrelevant fields per document, increasing noise rather than ensuring consistency. The unified schema approach undermines clarity and is likely to produce extraction errors due to field mismatches."}], "correct": "A", "select": 1, "group": "H"}, {"id": "f4-025", "domain": 4, "task_id": "4.4", "objective": "Implement validation, retry, and feedback loops for extraction quality", "situation": "A retail receipt-extraction system reads scanned receipts and populates a \"total\" field. Occasionally the printed total is smudged and misread, producing a plausible but wrong number that still passes schema validation.", "question": "What extraction design best catches this class of error before it reaches downstream accounting?", "options": [{"letter": "A", "text": "Skip extracting individual line items entirely so the pipeline runs faster and only returns the printed total field", "correct": false, "explanation": "Omitting line items removes the very data needed to compute an independent check, making the smudged-total error undetectable rather than caught."}, {"letter": "B", "text": "Trust the printed total field exactly as extracted, since it is the field accounting actually consumes further downstream anyway", "correct": false, "explanation": "Trusting the printed field alone provides no mechanism to catch a misread digit, which is exactly the failure mode described in the scenario."}, {"letter": "C", "text": "Have the model silently overwrite the printed total with whatever value it judges most plausible before it responds", "correct": false, "explanation": "Silently substituting a guessed value removes any auditable signal of disagreement and can introduce a fabricated number that looks authoritative to downstream consumers."}, {"letter": "D", "text": "Extract each line-item price plus the printed total, compute a calculated_total, and flag records where the two figures diverge", "correct": true, "explanation": "Extracting calculated_total alongside stated_total creates an independent cross-check; a discrepancy signals a likely misread and lets the system flag the record instead of passing a wrong total downstream silently."}], "correct": "D", "select": 1, "group": "I"}, {"id": "f4-026", "domain": 4, "task_id": "4.5", "objective": "Design efficient batch processing strategies", "situation": "A demand spike causes a submitted batch to reach its 24-hour expiration before a subset of requests could be sent to the model, and those requests come back with an expired result type. The team is not billed for them.", "question": "What should happen next?", "options": [{"letter": "A", "text": "Switch every future submission for this workload to the synchronous Messages API, since an expiration means batch processing cannot handle this workload at all", "correct": false, "explanation": "A single expiration event caused by a temporary demand spike does not mean the workload is unsuited to batching; resubmitting the small failed subset is the appropriate response."}, {"letter": "B", "text": "Resubmit the full original batch again in its entirety, since expiration means the whole batch's results were discarded and none of it can be trusted", "correct": false, "explanation": "Succeeded requests within the same batch already returned valid, billed results; resubmitting the whole batch would duplicate work and cost unnecessarily."}, {"letter": "C", "text": "Wait for the original batch to automatically requeue the expired requests once demand on the platform decreases, since expired requests remain pending", "correct": false, "explanation": "An expired request is terminal once the batch's 24-hour window closes; it does not automatically requeue and must be explicitly resubmitted."}, {"letter": "D", "text": "Collect the custom_id values for the expired requests and resubmit only those as a new batch, leaving the already-succeeded requests untouched", "correct": true, "explanation": "Only the expired requests failed to process; their custom_id values identify exactly which ones to resubmit, avoiding reprocessing requests that already succeeded."}], "correct": "D", "select": 1, "group": "H"}, {"id": "f4-027", "domain": 4, "task_id": "4.6", "objective": "Design multi-instance and multi-pass review architectures", "situation": "An engineer configures the generation session with a high extended thinking effort and instructs the model to reflect deeply on flaws before submitting its code. Bugs still slip through review.", "question": "Why is extended thinking insufficient as a substitute for an independent review instance here?", "options": [{"letter": "A", "text": "Extended thinking still runs in the same session that produced the code, so it retains the reasoning that justified those decisions", "correct": true, "explanation": "Extended thinking within the same session does not remove the retained reasoning context, which makes the model less likely to question its own prior decisions."}, {"letter": "B", "text": "Extended thinking is capped at a token budget that is too small to cover a second full pass over the generated code", "correct": false, "explanation": "The limitation isn't a budget ceiling; it's that the reflection happens inside the same reasoning context that produced the code."}, {"letter": "C", "text": "Extended thinking spends more tokens on reasoning, but the depth of scrutiny per issue stays roughly the same during the reflection step", "correct": false, "explanation": "The limitation is not about token cost or reasoning depth; extended thinking can genuinely deepen reasoning, it just does so from within a biased vantage point."}, {"letter": "D", "text": "Extended thinking disables tool use during reflection, so the model cannot re-read the files it just wrote to check them", "correct": false, "explanation": "Extended thinking does not disable tool use, and the core issue is retained reasoning context, not file access."}], "correct": "A", "select": 1, "group": "F"}, {"id": "f4-028", "domain": 4, "task_id": "4.1", "objective": "Design prompts with explicit criteria to improve precision and reduce false positives", "situation": "An architect is designing a multi-step pipeline to reduce false positives in a review category: a first API call generates draft findings, and a second API call reviews each draft against explicit criteria before finalizing it.", "question": "Why would this chained approach improve precision compared to a single-pass prompt with the same criteria?", "options": [{"letter": "A", "text": "The second API call is configured to invoke a more capable model by default, so the improved precision comes purely from a model upgrade between calls, not from the explicit two-step review process.", "correct": false, "explanation": "The question does not specify a model change, and the chained approach works effectively with the same model. The precision improvement results from the iterative review, not from using a more powerful model. Anthropic’s examples often use the same model across multiple steps to achieve robust outputs."}, {"letter": "B", "text": "Splitting the task across two calls doubles the amount of context available to the model by exposing all draft findings and the criteria to the second call, which mechanically improves classification accuracy by reducing false positives.", "correct": false, "explanation": "While the second call sees more tokens, simply providing additional context does not automatically improve accuracy. The benefit lies in the deliberate review step rather than in the quantity of context. Anthropic’s prompting best practices emphasize breaking tasks into steps for focused evaluation, not merely increasing input length."}, {"letter": "C", "text": "Chaining calls resets the model's system prompt after the draft generation, which strips any prior contextual cues that could have biased the first pass toward over-flagging benign patterns as findings, reducing false positives.", "correct": false, "explanation": "In prompt chaining, the second API call typically receives the output of the first call as input, so contextual cues are carried forward rather than stripped. The precision gain comes from the focused review step, not from resetting system prompts. Anthropic’s guidance emphasizes using chain-of-thought and structured multi-step prompts rather than relying on prompt resets."}, {"letter": "D", "text": "The second pass gives the model a separate opportunity to check each draft finding against the explicit criteria in isolation, catching cases where the first pass may have misapplied the criteria due to generating a large set of findings in one response.", "correct": true, "explanation": "Anthropic's documentation recommends multi-pass review and prompt chaining as a core pattern for improving accuracy. The second pass provides a dedicated evaluation stage, allowing the model to focus solely on applying the criteria to each finding, which reduces errors such as false positives. This aligns with the Claude Certified Architect Guide’s inclusion of 'multi-pass review' under Prompt Engineering & Structured Output."}], "correct": "D", "select": 1, "group": "I"}, {"id": "f4-029", "domain": 4, "task_id": "4.2", "objective": "Apply few-shot prompting to improve output consistency and quality", "situation": "A pull-request review agent is meant to flag branches (conditional paths) that lack test coverage. In practice it inconsistently flags newly introduced branches that are already exercised indirectly by an existing integration test, producing noisy false positives that erode reviewer trust. Detailed instructions about 'coverage' have not resolved the inconsistency.", "question": "What should the team add to the prompt?", "options": [{"letter": "A", "text": "A couple of examples pairing a diff with a coverage judgment: one branch with no test, one covered indirectly", "correct": true, "explanation": "Showing a genuinely uncovered branch alongside a lookalike branch that is indirectly covered, with the resulting judgment for each, demonstrates the exact distinction the agent needs to generalize, which is the ambiguous-case handling that few-shot examples are designed to convey."}, {"letter": "B", "text": "A rule treating any file changed by fewer than ten lines as automatically having adequate coverage", "correct": false, "explanation": "An arbitrary line-count threshold is unrelated to whether a branch actually has test coverage and would let genuinely uncovered small diffs slip through while ignoring the indirect-coverage ambiguity entirely."}, {"letter": "C", "text": "A requirement to run the full test suite and flag any branch under one hundred percent line coverage", "correct": false, "explanation": "Running the full suite and enforcing a strict coverage percentage swaps in a numeric gate instead of resolving the ambiguous case the prompt is failing on, and it does not use examples to teach the distinction."}, {"letter": "D", "text": "An instruction to flag every new conditional branch in a diff regardless of the surrounding test suite entirely", "correct": false, "explanation": "Flagging every new branch unconditionally removes the judgment call entirely and guarantees the same false-positive pattern the team is trying to eliminate, since indirectly covered branches would still be flagged."}], "correct": "A", "select": 1, "group": "I"}, {"id": "f4-030", "domain": 4, "task_id": "4.2", "objective": "Apply few-shot prompting to improve output consistency and quality", "situation": "A tool extracts a paper's 'sample size' and 'statistical method' fields. Some papers place this information in a clearly labeled Methodology section, while others embed it in a sentence within the Results or Discussion section without any nearby heading. The tool reliably extracts from labeled Methodology sections but frequently returns null when the same information is embedded elsewhere.", "question": "What is the best fix?", "options": [{"letter": "A", "text": "Configure the tool to first search the Methodology section, and only fall back to other sections if the fields are missing.", "correct": false, "explanation": "A fallback strategy does not teach the model how to extract the fields from unstructured text. If the tool cannot extract from embedded sentences, simply telling it to look elsewhere will not solve the underlying pattern-recognition problem. The issue is the extraction logic, not the search order."}, {"letter": "B", "text": "Exclude any paper that lacks a labeled Methodology section from the extraction pipeline.", "correct": false, "explanation": "This is a restrictive workaround that limits the tool's utility and does not improve extraction accuracy for papers with embedded information. Anthropic's approach to robust AI systems involves building models that can handle diverse formats, not excluding edge cases. The recommended fix is to improve the model's capability, not the eligibility criteria."}, {"letter": "C", "text": "Increase the model's context window to ensure it reads the entire paper rather than a truncated excerpt.", "correct": false, "explanation": "The problem is not missing data; the tool already retrieves the text. The failure is in extraction logic when headings are absent. A larger context window may help if text is truncated, but the question indicates the tool returns null when information is embedded, implying it has access but cannot parse it. This change does not address the core issue."}, {"letter": "D", "text": "Provide the model with extraction examples from both a labeled Methodology section and from an embedded sentence in Results, demonstrating how to extract the fields in both cases.", "correct": true, "explanation": "This is a few-shot prompting technique that directly addresses the extraction failure. By showing the model examples of both scenarios, it learns to recognize the relevant information regardless of section heading. Anthropic's documentation on prompt engineering recommends using diverse examples to guide model behavior and improve performance on varied inputs."}], "correct": "D", "select": 1, "group": "I"}, {"id": "f4-031", "domain": 4, "task_id": "4.3", "objective": "Enforce structured output using tool use and JSON schemas", "situation": "A pipeline first needs Claude to run extract_metadata on an uploaded document before any enrichment or summarization steps proceed. Other tools such as translate_text and summarize_document are also registered on the same request, and the team is worried Claude might call one of those first.", "question": "Which tool_choice setting guarantees extract_metadata runs on this turn?", "options": [{"letter": "A", "text": "tool_choice: {\"type\": \"tool\", \"name\": \"extract_metadata\"}", "correct": true, "explanation": "Forcing a specific named tool guarantees that exact tool is called on this turn, ensuring extract_metadata runs before any other registered tool such as translate_text or summarize_document."}, {"letter": "B", "text": "tool_choice: {\"type\": \"any\"}", "correct": false, "explanation": "\"any\" only guarantees that some tool is called; Claude could still pick translate_text or summarize_document instead of extract_metadata."}, {"letter": "C", "text": "tool_choice: {\"type\": \"auto\"}", "correct": false, "explanation": "\"auto\" lets Claude decide freely whether to call a tool at all and which one, so it does not guarantee extract_metadata runs first."}, {"letter": "D", "text": "tool_choice: {\"type\": \"none\"}", "correct": false, "explanation": "\"none\" prevents Claude from calling any tool at all, which would block extract_metadata from running rather than force it."}], "correct": "A", "select": 1, "group": "H"}, {"id": "f4-032", "domain": 4, "task_id": "4.3", "objective": "Enforce structured output using tool use and JSON schemas", "situation": "A team is designing a schema to extract a customer's phone number from scanned support emails. Many older emails in the corpus never mention a phone number at all. In an early version, phone_number was marked as a required string field, and the team noticed Claude sometimes fabricated plausible-looking numbers to satisfy the schema.", "question": "What is the best fix?", "options": [{"letter": "A", "text": "Keep phone_number required but change its type from string to an enum of common area codes so the model has fewer values to guess from", "correct": false, "explanation": "An enum of area codes still forces the model to pick some value, and area codes don't even represent full phone numbers, so fabrication risk remains and the field's meaning becomes wrong."}, {"letter": "B", "text": "Keep phone_number required and add a second required field called phone_number_confidence so low-confidence guesses can be filtered out later", "correct": false, "explanation": "Adding a confidence field doesn't change the fact that phone_number itself is still required, so the model is still forced to produce some value even when none exists in the source."}, {"letter": "C", "text": "Make phone_number an optional, nullable field so Claude can omit it or return null when the source document contains no phone number", "correct": true, "explanation": "Designing fields as optional/nullable when source documents may not contain the information removes the pressure to fabricate a value just to satisfy a required-field constraint."}, {"letter": "D", "text": "Remove the tool definition entirely and ask Claude in plain prose to only include a phone number if it is confident one was found", "correct": false, "explanation": "Dropping tool use altogether reintroduces JSON syntax risk and gives up the schema-compliance guarantees that tool use provides, trading one problem for a worse one."}], "correct": "C", "select": 1, "group": "H"}, {"id": "f4-033", "domain": 4, "task_id": "4.5", "objective": "Design efficient batch processing strategies", "situation": "A developer is assembling a batch of translation requests and needs to assign each one an identifier that the Message Batches API will accept and later use to correlate a result with its original request.", "question": "Which identifier is valid for this purpose?", "options": [{"letter": "A", "text": "doc-2026-report_final, since it uses only letters, digits, hyphens, and underscores, staying under the 64-character limit", "correct": true, "explanation": "custom_id must match ^[a-zA-Z0-9_-]{1,64}$, meaning only alphanumeric characters, hyphens, and underscores are allowed; this identifier satisfies that constraint."}, {"letter": "B", "text": "doc#2026#report#final#v2, since separating each descriptive segment with a distinct punctuation mark avoids any ambiguity about where one part ends", "correct": false, "explanation": "The hash character is not part of the allowed character set for custom_id, so this identifier would fail validation."}, {"letter": "C", "text": "doc 2026 report final v2, since spaces between each descriptive segment keep the identifier readable when scanning a long results file by eye", "correct": false, "explanation": "Spaces are not permitted characters in a custom_id, regardless of how readable they make the identifier."}, {"letter": "D", "text": "doc/2026/report:final, since embedding the ingestion date and a colon-separated section label makes the request easier to trace during a later audit", "correct": false, "explanation": "Slashes and colons are not permitted characters in a custom_id and would cause the request to be rejected."}], "correct": "A", "select": 1, "group": "H"}, {"id": "f4-034", "domain": 4, "task_id": "4.2", "objective": "Apply few-shot prompting to improve output consistency and quality", "situation": "A team built a classification prompt with twenty exact input-output pairs, one for every edge case they had personally encountered in their historical data. The prompt performs well on those twenty inputs but degrades noticeably whenever a customer submits a new input that is similar to, but not identical to, one of the twenty.", "question": "What change would best help the model generalize its judgment to these novel-but-similar inputs?", "options": [{"letter": "A", "text": "Reduce the twenty examples to a small set that makes the underlying decision rule visible to the model", "correct": true, "explanation": "A small number of examples chosen to reveal the reasoning behind the classification, rather than an exhaustive literal catalog, lets the model generalize that judgment to novel inputs instead of only matching the specific cases it was shown."}, {"letter": "B", "text": "Remove the examples entirely and rely on a single instruction sentence describing the desired behavior", "correct": false, "explanation": "Removing examples entirely discards the demonstrated reasoning pattern altogether; instructions alone are typically less effective than well-chosen examples at conveying nuanced judgment for this kind of case."}, {"letter": "C", "text": "Keep appending every newly discovered literal pair to the prompt so every past case is eventually represented", "correct": false, "explanation": "Continuing to add every literal case the team encounters reinforces exact-match behavior rather than the underlying pattern, so novel-but-similar inputs would remain likely to be misclassified."}, {"letter": "D", "text": "Increase the max_tokens parameter so the model has more room to reason before each classification", "correct": false, "explanation": "More token budget for reasoning does not supply the model with the demonstrated judgment pattern it is missing; the failure here is about what the model has been shown, not how much room it has to think."}], "correct": "A", "select": 1, "group": "I"}, {"id": "f4-035", "domain": 4, "task_id": "4.2", "objective": "Apply few-shot prompting to improve output consistency and quality", "situation": "An agent has both a search_docs tool and a search_code tool available. For requests that could plausibly be answered by either tool, such as 'where is the rate limit defined,' the agent picks inconsistently between them across similar sessions. The team wants the agent to make a consistent, well-reasoned choice for this class of ambiguous request.", "question": "What is the most effective change?", "options": [{"letter": "A", "text": "Add a routing step that always calls search_docs first and only falls back to search_code if no results are returned.", "correct": false, "explanation": "A hardcoded docs-first rule is brittle and can choose the wrong source for requests that are ambiguous but better answered from code. Anthropic recommends semantic tool definitions, clear descriptions, enums, and clarifying questions; it does not recommend replacing intent-based selection with a fixed fallback chain."}, {"letter": "B", "text": "Rename the two tools to be more visually distinct (e.g., docs_lookup and code_search) to reduce the chance of confusion.", "correct": false, "explanation": "Distinct, verb-noun tool names can reduce collisions, but simply making names visually different does not explain when each tool should be used for queries like 'where is the rate limit defined.' Official guidance treats clear names as one best practice, not a sufficient fix for semantic ambiguity; disambiguation is better handled through descriptions and input_examples."}, {"letter": "C", "text": "Expand each tool's natural-language description with more adjectives that characterize its typical use cases.", "correct": false, "explanation": "More adjectives add noise rather than precision. Anthropic emphasizes that tool descriptions should be clear, detailed semantic contracts, and it recommends supplementary mechanisms like input_examples for ambiguous tool selection rather than verbose or vague characterization."}, {"letter": "D", "text": "Provide input_examples in the tool definitions that demonstrate for the same ambiguous request which tool should be selected and why.", "correct": true, "explanation": "Anthropic's tool-use documentation recommends input_examples specifically for disambiguation when multiple tools have similar names or functions, such as showing which tool to select for an ambiguous request and why. They are optional but valuable supplements to clear tool descriptions, and Anthropic reports they can improve complex parameter generation accuracy from 72% to 90%. This directly addresses consistent, well-reasoned tool selection for ambiguous requests."}], "correct": "D", "select": 1, "group": "I"}, {"id": "f4-036", "domain": 4, "task_id": "4.6", "objective": "Design multi-instance and multi-pass review architectures", "situation": "An architect wants a review subagent that can analyze generated code but must never modify it.", "question": "How should the subagent be configured?", "options": [{"letter": "A", "text": "Restrict the subagent's tools field to read-only tools like Read, Grep, and Glob, omitting Edit, Write, and Bash", "correct": true, "explanation": "Restricting the tools field to read-only tools such as Read, Grep, and Glob enforces that the subagent can analyze but never modify files, rather than relying on instructions alone."}, {"letter": "B", "text": "Grant the subagent Edit and Write access but set permission mode to require manual approval for every edit it proposes", "correct": false, "explanation": "Granting Edit and Write access still allows modification capability to exist; the goal of read-only analysis is better met by omitting those tools entirely."}, {"letter": "C", "text": "Leave the tools field unset so the subagent inherits every tool from the parent, then instruct it in the prompt not to use Edit or Write", "correct": false, "explanation": "Leaving tools inherited and relying on a prompt instruction is a soft constraint the model could still bypass; tool restriction is the enforced mechanism."}, {"letter": "D", "text": "Grant the subagent Bash access only, since Bash can be used to read file contents without needing dedicated Read or Grep tools", "correct": false, "explanation": "Bash alone isn't a reliable read-only substitute and also gives broader command execution capability than a review-only subagent needs."}], "correct": "A", "select": 1, "group": "F"}, {"id": "f4-037", "domain": 4, "task_id": "4.4", "objective": "Implement validation, retry, and feedback loops for extraction quality", "situation": "A vendor-contract extraction pipeline must extract both a \"start_date\" and an \"end_date\" and ensure the contract term is logically ordered. Occasionally the model extracts an end_date that precedes the start_date even though both individual dates are correctly read from the text.", "question": "What self-correction design best supports catching and resolving this class of issue?", "options": [{"letter": "A", "text": "Instruct the model once, in the original prompt, to be careful with dates, and skip any comparison afterward", "correct": false, "explanation": "A vague, one-time instruction with no post-extraction check provides no reliable detection mechanism and leaves ordering errors to pass through undetected."}, {"letter": "B", "text": "Remove the end_date field from the schema so an out-of-order pair can never be produced by the extractor", "correct": false, "explanation": "Removing the field eliminates the data needed for downstream contract-term calculations entirely, rather than fixing the ordering inconsistency."}, {"letter": "C", "text": "Add an ordering check after extraction, and if end_date precedes start_date, retry with that inconsistency as feedback", "correct": true, "explanation": "Since both dates are individually well-formed, the ordering violation is a semantic issue that only an explicit cross-field comparison can catch; feeding that specific inconsistency back into a retry lets the model self-correct which date it misassigned."}, {"letter": "D", "text": "Rely on the JSON schema's type constraints alone, since both fields are already validated as proper date strings anyway", "correct": false, "explanation": "Type constraints confirm each field is a valid date string but say nothing about the logical relationship between the two fields, so schema validation alone will not catch an inverted date range."}], "correct": "C", "select": 1, "group": "I"}, {"id": "f4-038", "domain": 4, "task_id": "4.1", "objective": "Design prompts with explicit criteria to improve precision and reduce false positives", "situation": "An architect wants to improve consistency in how a review prompt classifies findings as \"bug\" versus \"style,\" and decides to add a small set of worked examples to the prompt in addition to the written criteria.", "question": "Which set of examples best supports this goal?", "options": [{"letter": "A", "text": "Three to five diverse examples, each showing a snippet plus the correct classification and a short reason, covering both clear bugs and clear style issues as well as one borderline case.", "correct": true, "explanation": "Three to five diverse examples covering clear bugs, clear style issues, and a borderline case help the model generalize the classification rule. Including short reasons with each example makes the distinction explicit, which is the recommended approach for improving consistency."}, {"letter": "B", "text": "A single worked example containing only a code snippet and the label 'bug' with no further explanation, so the model must derive the bug-versus-style distinction purely from the example.", "correct": false, "explanation": "A single example with a label and no explanation forces the model to guess the underlying classification rule. This is less reliable for consistent classification than examples paired with explicit reasoning."}, {"letter": "C", "text": "Ten near-duplicate examples of the same kind of off-by-one bug, each with slight variations in context and severity, so the model learns to recognize the pattern in different codebases.", "correct": false, "explanation": "Ten near-duplicate off-by-one bug examples, despite slight variations, still represent a narrow pattern. The model risks overfitting to that specific pattern rather than learning the general bug-versus-style distinction."}, {"letter": "D", "text": "One long example showing the single most severe bug the team has ever found in production, described in exhaustive detail to anchor the model's sense of scale and establish a clear benchmark for bug severity.", "correct": false, "explanation": "One exhaustive example of a severe bug provides no coverage of style issues or borderline cases. It fails to help the model consistently distinguish bugs from style across varied inputs."}], "correct": "A", "select": 1, "group": "I"}, {"id": "f4-039", "domain": 4, "task_id": "4.3", "objective": "Enforce structured output using tool use and JSON schemas", "situation": "A team is building a pipeline that asks Claude to read free-form support tickets and return a JSON object with fields like priority, category, and summary. Early prototypes used a prompt asking Claude to \"reply with only JSON\", but downstream parsing occasionally failed on malformed brackets and stray commentary text.", "question": "Which approach most reliably eliminates these JSON syntax failures?", "options": [{"letter": "A", "text": "Append a stricter instruction to the system prompt demanding that Claude output valid JSON and nothing else, then retry the request whenever parsing fails", "correct": false, "explanation": "Stronger wording in the system prompt can reduce but not eliminate malformed output, and retry loops add latency and cost without guaranteeing correctness."}, {"letter": "B", "text": "Lower the temperature parameter to 0 so that Claude's text completions become more deterministic and less prone to formatting mistakes", "correct": false, "explanation": "Lowering temperature makes responses more consistent but does not enforce a schema; the model can still emit prose, missing brackets, or truncated JSON."}, {"letter": "C", "text": "Ask Claude to wrap its JSON output in triple backticks and strip the backticks during post-processing before parsing the remaining text", "correct": false, "explanation": "Markdown fencing is a formatting convention, not a schema constraint; the enclosed content itself can still be malformed or contain extra commentary."}, {"letter": "D", "text": "Define an extraction tool with an input_schema describing the fields, and parse the structured arguments from the resulting tool_use block instead of parsing free text", "correct": true, "explanation": "Tool use with a JSON schema constrains the model's output to match the schema's structure, guaranteeing well-formed structured data extracted from the tool_use block rather than relying on the model to format free text correctly."}], "correct": "D", "select": 1, "group": "H"}, {"id": "f4-040", "domain": 4, "task_id": "4.2", "objective": "Apply few-shot prompting to improve output consistency and quality", "situation": "An automated code reviewer flags many instances of a pattern (a broad except clause) as issues, but a large fraction of those flags are on lines where the pattern is intentional and acceptable, such as top-level error boundaries that log and re-raise. Reviewers are starting to ignore the tool's output because of the false-positive rate.", "question": "What change would most directly reduce false positives while still catching genuine issues?", "options": [{"letter": "A", "text": "Add paired examples of a genuinely problematic instance and an acceptable instance, each with the correct verdict", "correct": true, "explanation": "Using paired examples with correct verdicts helps the automated reviewer learn the boundary between intentional acceptable uses and genuine violations. Anthropic's guidance for reducing false positives in safety classifiers emphasizes showing both problematic and acceptable instances with labeled verdicts, rather than simply removing or thresholding checks. This approach directly addresses context-based misclassification and preserves detection of real issues."}, {"letter": "B", "text": "Instruct the model to only flag the pattern when it appears more than three times in the same file", "correct": false, "explanation": "A frequency-based threshold is arbitrary and would miss single genuine issues. False positives arise from misclassifying context, not from how many times a pattern appears. This change would reduce true positives without teaching the model the acceptable boundary between intentional top-level error handling and problematic bare except clauses."}, {"letter": "C", "text": "Remove the broad except clause check from the rule set entirely, since it currently produces too many false positives", "correct": false, "explanation": "Removing the check eliminates false positives but also eliminates all genuine detections, creating false negatives. Official safety guidance warns against overly broad removal and emphasizes balancing helpfulness with appropriate limitations. The better approach is to refine the check with labeled examples or adjust its scope, not to remove it wholesale."}, {"letter": "D", "text": "Lower the confidence threshold so only the single highest-confidence finding per file is reported", "correct": false, "explanation": "Lowering the confidence threshold typically increases false positives, not reduces them, and reporting only one finding per file suppresses other genuine issues. The goal is better boundary discrimination through targeted examples or scope refinement, not artificially limiting output volume."}], "correct": "A", "select": 1, "group": "I"}, {"id": "f4-041", "domain": 4, "task_id": "4.6", "objective": "Design multi-instance and multi-pass review architectures", "situation": "Per-file passes on two interdependent files each recommend a different fix for what turns out to be the same underlying data-flow issue, and the two recommendations conflict.", "question": "What architectural step should resolve this rather than picking one per-file recommendation at random?", "options": [{"letter": "A", "text": "A separate cross-file integration pass that examines both files together and produces one recommendation based on the actual data flow", "correct": true, "explanation": "A cross-file integration pass is designed to resolve exactly this kind of conflict by examining how data flows between the interdependent files together, rather than in separate isolated passes."}, {"letter": "B", "text": "Asking the original generator to arbitrate between the two recommendations, since it has full context on why it wrote the code that way", "correct": false, "explanation": "Letting the original generator arbitrate reintroduces the same retained-reasoning bias that independent review is meant to avoid."}, {"letter": "C", "text": "Re-running each per-file pass a second time and keeping whichever recommendation is worded with higher confidence language", "correct": false, "explanation": "Confidence-sounding wording doesn't reflect which recommendation actually accounts for the real cross-file data flow."}, {"letter": "D", "text": "Merging the two files into one before review so a single per-file pass can cover both without needing an integration step", "correct": false, "explanation": "Merging files together is not the designed mechanism and doesn't generalize; the intended fix is a dedicated integration pass, not restructuring the files under review."}], "correct": "A", "select": 1, "group": "F"}, {"id": "f4-042", "domain": 4, "task_id": "4.1", "objective": "Design prompts with explicit criteria to improve precision and reduce false positives", "situation": "A team is writing the \"critical\" severity definition for their review prompt. They want reviewers across the org to converge on the same classification for the same kind of issue.", "question": "Which definition best supports that goal?", "options": [{"letter": "A", "text": "Critical: the change allows unauthenticated access to data that should require authorization, for example a removed permission check before a database query that returns another user's records.", "correct": true, "explanation": "This definition uses clear, objective criteria (unauthenticated access to authorized data) with a concrete example, reducing ambiguity and promoting consistent classification. It aligns with industry standards like the OWASP Broken Access Control category and Anthropic's own treatment of such flaws as critical vulnerabilities (e.g., CVE-2025-49596, which involved unauthenticated command execution). A well-defined threshold helps all reviewers converge on the same severity."}, {"letter": "B", "text": "Critical: the change introduces a risky pattern that would make an experienced engineer feel uneasy, such as a race condition in a payment processing function that could lead to data inconsistency.", "correct": false, "explanation": "This definition relies on subjective feelings (\"feel uneasy\") and a vague pattern (\"risky pattern\"), making consistent classification difficult across different reviewers. While race conditions can be critical, the definition lacks objective, measurable criteria; what makes one engineer uneasy might not register with another. A convergent definition should minimize personal judgment and rely on observable, reproducible conditions."}, {"letter": "C", "text": "Critical: the change is significantly more complex or harder to reason about than the rest of the surrounding code in the same file, such as a refactored function that now uses deeply nested callbacks that obscure the logic.", "correct": false, "explanation": "Complexity alone does not determine severity; well-structured but intricate logic may be safe, while a one-line change could introduce a critical vulnerability. This definition introduces subjectivity (\"significantly more complex\") and does not tie to actual harm or security impact. Consistent classification requires focusing on tangible outcomes like data exposure, integrity loss, or functionality breakage, not subjective code aesthetics."}, {"letter": "D", "text": "Critical: the change touches a module that has historically caused many production incidents, for example the user authentication service that had three outages last quarter due to token validation failures.", "correct": false, "explanation": "Severity should be based on the actual risk introduced by the change itself, not the module's past reliability. A critical defect can occur anywhere, and a minor change in a historically problematic module may not be critical. This definition fails to provide consistent criteria—reviewers might assign 'critical' to any change in such modules, leading to alert fatigue and inconsistent triage. A convergent definition must focus on the nature and potential impact of the change, not its location."}], "correct": "A", "select": 1, "group": "I"}, {"id": "f4-043", "domain": 4, "task_id": "4.6", "objective": "Design multi-instance and multi-pass review architectures", "situation": "A reviewer instance returns 12 findings. The architect wants only high-confidence issues auto-fixed and everything else routed to a human.", "question": "Which review design supports this goal?", "options": [{"letter": "A", "text": "Have the reviewer instance rank findings only by the severity of the underlying bug, then auto-fix the top-ranked items regardless of how certain the reviewer was", "correct": false, "explanation": "Severity alone says nothing about how certain the reviewer is that a finding is real, so it doesn't support routing on confidence."}, {"letter": "B", "text": "Have the reviewer instance tag each finding with a confidence level, then auto-fix the high-confidence findings and send the low-confidence ones to a human", "correct": true, "explanation": "Self-reported confidence alongside each finding is exactly what enables calibrated routing: automate the confident items and send uncertain ones for further scrutiny."}, {"letter": "C", "text": "Have the generator instance re-run its own self-review and quietly discard any finding it personally disagrees with before a human ever sees it", "correct": false, "explanation": "Letting the generator filter findings reintroduces the self-review limitation and removes the human-routing signal entirely."}, {"letter": "D", "text": "Have the reviewer instance combine all findings into one summary paragraph and let a human manually re-derive which findings seem reliable", "correct": false, "explanation": "A single unstructured summary forces the human to manually reconstruct confidence for every finding, defeating the purpose of calibrated routing."}], "correct": "B", "select": 1, "group": "F"}, {"id": "f4-044", "domain": 4, "task_id": "4.5", "objective": "Design efficient batch processing strategies", "situation": "A data-ingestion job needs to classify 150,000 scanned invoices in one nightly run using the Message Batches API.", "question": "What must the team account for given the platform's per-batch limits?", "options": [{"letter": "A", "text": "Use the synchronous Messages API instead of the Batches API, because batches are limited to 500 requests per hour and 150,000 invoices cannot be processed overnight.", "correct": false, "explanation": "The Batches API is not limited to 500 requests per hour; it supports batches of up to 100,000 requests with a processing SLA of up to 24 hours. The synchronous Messages API can be used for real-time calls, but it is not required and is typically less cost-effective for large batch classification."}, {"letter": "B", "text": "Submit all 150,000 invoices in a single batch, because the Batches API has no request limit and only restricts the output size of each response.", "correct": false, "explanation": "The Message Batches API does have explicit request and payload limits: up to 100,000 requests per batch and a total size of 256 MB, whichever comes first. Submitting 150,000 invoices in one batch would exceed the request limit."}, {"letter": "C", "text": "Divide the invoices into batches of exactly 1,000, since this is the maximum batch size allowed by the API for any data type.", "correct": false, "explanation": "The maximum batch size is 100,000 requests, not exactly 1,000. While splitting large workloads into smaller independent batches may improve reliability and incremental processing, it is not a platform maximum."}, {"letter": "D", "text": "Split the 150,000 invoices across at least 2 batch submissions, because a single Message Batch is capped at 100,000 requests or 256 MB total payload, whichever comes first. (More batches may be needed if the total payload exceeds 256 MB.)", "correct": true, "explanation": "According to the current Anthropic Message Batches API documentation, each Message Batch can contain up to 100,000 requests or 256 MB of total payload, whichever limit is reached first. Since 150,000 invoices exceed the request limit, the team must split them into at least 2 batch submissions (150,000 ÷ 100,000 = 1.5, rounded up to 2). If the combined payload of any batch exceeds 256 MB, additional batches will be necessary."}], "correct": "D", "select": 1, "group": "H"}, {"id": "f4-045", "domain": 4, "task_id": "4.5", "objective": "Design efficient batch processing strategies", "situation": "After a batch of 20,000 document-summarization requests finishes, the results show several requests came back with an errored result type carrying an invalid_request_error because those specific documents exceeded the model's context window.", "question": "What is the most efficient way to recover?", "options": [{"letter": "A", "text": "Discard the errored requests permanently, since an invalid_request_error means those documents cannot be processed through the Messages API in any form", "correct": false, "explanation": "The documents are not unprocessable outright; the fix is restructuring them into smaller chunks that fit the context window, not abandoning them."}, {"letter": "B", "text": "Identify the errored requests by their custom_id, split only those oversized documents into smaller chunks, and resubmit just those chunked requests in a new batch", "correct": true, "explanation": "custom_id lets the team isolate exactly which documents failed; chunking only those oversized documents and resubmitting just that small set avoids reprocessing the 19,000-plus requests that already succeeded."}, {"letter": "C", "text": "Reduce max_tokens on every request in a new batch covering all 20,000 documents, since output length is what caused the original context-window errors", "correct": false, "explanation": "A context-window overflow on input is caused by prompt size, not output length, so lowering max_tokens does not address the root cause and needlessly reprocesses documents that already succeeded."}, {"letter": "D", "text": "Resubmit the entire original batch unchanged, since the Batches API automatically retries any request that previously errored before returning final results", "correct": false, "explanation": "The Batches API does not automatically retry errored requests; each request is billed or unbilled and returned as-is, and resubmitting everything would waste cost on already-successful documents."}], "correct": "B", "select": 1, "group": "H"}, {"id": "f4-046", "domain": 4, "task_id": "4.4", "objective": "Implement validation, retry, and feedback loops for extraction quality", "situation": "An employment-contract extractor pulls a \"salary\" field that correctly matches its declared numeric type and required-field constraints, but the value is denominated in the wrong currency because the model misread an ambiguous currency symbol on a multi-currency contract. Schema validation passes without complaint.", "question": "How should this be characterized?", "options": [{"letter": "A", "text": "As a token-limit truncation issue that will be resolved by simply increasing max_tokens on the next call made", "correct": false, "explanation": "The scenario describes a misreading of an ambiguous symbol, not a response cut short by the token limit; increasing max_tokens does not address a misinterpretation of source content."}, {"letter": "B", "text": "As a semantic error outside schema conformance, needing a rule that cross-checks the currency symbol against the number", "correct": true, "explanation": "The salary value is structurally valid but semantically wrong due to a currency misread; this class of error passes schema checks and needs a dedicated semantic validation rule tied to the currency indicator."}, {"letter": "C", "text": "As a refusal, since the model effectively declined to resolve the ambiguous currency symbol on the contract", "correct": false, "explanation": "A refusal is an explicit stop_reason indicating the model declined to respond; here the model produced a complete, schema-valid answer that was simply semantically incorrect."}, {"letter": "D", "text": "As a schema violation that strict tool use should have already prevented from ever reaching the application layer at all", "correct": false, "explanation": "Strict tool use validates structure and type, not whether a correctly typed number carries the intended real-world meaning such as the correct currency."}], "correct": "B", "select": 1, "group": "I"}, {"id": "f4-047", "domain": 4, "task_id": "4.1", "objective": "Design prompts with explicit criteria to improve precision and reduce false positives", "situation": "After temporarily disabling a high false-positive \"performance suggestions\" category and rewriting its criteria with specific, checkable rules, an architect must decide when it is safe to re-enable the category for the whole team.", "question": "What is the most appropriate validation step before re-enabling it broadly?", "options": [{"letter": "A", "text": "Re-enable the category only for pull requests opened by the engineer who reported the false positives, as a limited pilot to verify the criteria, while keeping it disabled for all other contributors.", "correct": false, "explanation": "Restricting the pilot to one engineer's pull requests does not validate the criteria across the diverse code changes made by the entire team. False positives are likely team-wide, so testing on a single contributor fails to assess whether the rewrite resolves the problem for others and delays restoring trust for the rest of the team."}, {"letter": "B", "text": "Run the rewritten prompt against a held-out set of past pull requests with known findings, and confirm its false positive rate has dropped to an acceptable level before re-enabling it for everyone.", "correct": true, "explanation": "Testing the rewritten prompt on a held-out set of historical pull requests with known findings provides direct, quantitative evidence that the false positive rate has dropped to an acceptable level. This data-driven validation ensures the criteria work as intended before risking a team-wide rollout, minimizing the chance of reintroducing noise."}, {"letter": "C", "text": "Ask a single senior engineer to review the new criteria against a small set of past pull requests that triggered false positives, and authorize re-enabling if the criteria appear sound based on that manual check.", "correct": false, "explanation": "A single engineer's subjective review of the criteria against a small, non-representative sample of past pull requests does not provide objective, measurable evidence that the false positive rate has actually improved. This approach relies on human judgment rather than systematic evaluation, which may not catch all issues and lacks statistical confidence."}, {"letter": "D", "text": "Re-enable the category immediately after the new criteria are added to the repository, because the explicit rules themselves demonstrate improved precision without needing any further validation against historic pull requests.", "correct": false, "explanation": "Adding explicit, checkable rules does not guarantee they function correctly in practice; only empirical validation against real-world examples can confirm improved precision. Without testing against historic pull requests, the team risks re-enabling a category that may still produce an unacceptable number of false positives, undermining trust."}], "correct": "B", "select": 1, "group": "I"}, {"id": "f4-048", "domain": 4, "task_id": "4.2", "objective": "Apply few-shot prompting to improve output consistency and quality", "situation": "A security-scanning assistant is supposed to report findings with four consistent fields: location, issue, severity, and suggested fix. Written instructions specify these four fields, but outputs still vary: some findings omit severity, others merge the issue and fix into one sentence.", "question": "What is the most effective way to lock in the desired structure?", "options": [{"letter": "A", "text": "Ask the model to double-check its own output carefully against the four-field requirement before returning it", "correct": false, "explanation": "A self-check instruction can catch some errors, but it relies on the model's own judgment of what 'complete' looks like without ever having seen the target format, which is the same ambiguity that caused the inconsistency in the first place."}, {"letter": "B", "text": "Add a note at the end of the instructions reminding the model not to forget the severity field", "correct": false, "explanation": "A reminder is still a prose instruction, and the scenario already establishes that written instructions describing the four required fields have not produced consistent output, so another reminder is unlikely to fix the pattern."}, {"letter": "C", "text": "Increase the output token limit so the model has room to include every field without truncation", "correct": false, "explanation": "Truncation due to token limits is not indicated as the cause here; the problem described is structural omission and merging of fields, not content being cut off, so raising the token limit would not address the root cause."}, {"letter": "D", "text": "Provide worked examples that render all four fields in the same order, including one low-severity case", "correct": true, "explanation": "Concrete examples that consistently render all four fields, including a lower-stakes example so the model doesn't only see high-severity precedent, demonstrate the exact target format and generalize better than an added reminder in the instructions."}], "correct": "D", "select": 1, "group": "I"}, {"id": "f4-049", "domain": 4, "task_id": "4.5", "objective": "Design efficient batch processing strategies", "situation": "A legal-review pipeline must guarantee that every submitted contract receives a result within 30 hours of arrival, using the Message Batches API's up-to-24-hour processing window.", "question": "How often must the pipeline start a new batch submission cycle to guarantee this SLA in the worst case?", "options": [{"letter": "A", "text": "At least once every 18 hours, since leaving extra headroom beyond the minimum required interval better protects an already generous SLA", "correct": false, "explanation": "An 18-hour interval leaves a worst-case total of 18 + 24 = 42 hours, which breaches the 30-hour SLA rather than protecting it."}, {"letter": "B", "text": "At least once every 24 hours, since that matches the batch processing window and therefore satisfies any SLA built on top of it", "correct": false, "explanation": "A 24-hour interval yields a worst-case total of 24 + 24 = 48 hours, well past the 30-hour promise; matching the processing window is not sufficient here."}, {"letter": "C", "text": "At least once every 30 hours, since the submission cadence should simply mirror the length of the SLA the pipeline promises", "correct": false, "explanation": "Submitting only once every 30 hours would let a worst-case contract wait 30 hours before even entering a batch, then up to 24 more to process, far exceeding the SLA."}, {"letter": "D", "text": "At least once every 6 hours, since a contract's worst-case wait plus the 24-hour processing window still totals no more than 30 hours", "correct": true, "explanation": "Worst case, a contract waits the full submission interval before being included, then up to 24 hours to process: interval + 24 <= 30 means the interval must be 6 hours or less."}], "correct": "D", "select": 1, "group": "H"}, {"id": "f4-050", "domain": 4, "task_id": "4.2", "objective": "Apply few-shot prompting to improve output consistency and quality", "situation": "A team's code-review assistant prompt already spells out an exhaustive, itemized rubric for flagging issues, but reviewers still receive inconsistently formatted findings across runs: some list severity before location, others omit the suggested fix entirely. The team wants the most effective fix for this output-consistency problem.", "question": "What should they do?", "options": [{"letter": "A", "text": "Split the prompt into two calls: one that generates findings and a second that reformats them into the target schema", "correct": false, "explanation": "Adding a second formatting pass adds latency and cost and does not address the root cause: the first call's instructions are still ambiguous about structure, so the second call inherits inconsistent input to reformat."}, {"letter": "B", "text": "Rewrite the rubric as an even longer numbered list that spells out every field, its position, and its formatting rule in detail", "correct": false, "explanation": "Piling on more prose instructions is the approach the team already tried; detailed instructions alone tend to produce inconsistent formatting because the model still has to infer the exact structure from a description rather than seeing it demonstrated."}, {"letter": "C", "text": "Lower the temperature parameter to zero so the same tokens are sampled deterministically across every single run", "correct": false, "explanation": "Temperature controls sampling randomness and can reduce run-to-run variance, but it does not teach the model what the correct field order or completeness should be, so structural inconsistency driven by ambiguous instructions would persist."}, {"letter": "D", "text": "Add three to five worked examples in example tags that each show the exact field order and formatting wanted", "correct": true, "explanation": "When detailed instructions alone produce inconsistent formatting, concrete input-output examples are the most reliable way to pin down the exact structure and field order, and a handful of well-chosen examples covering an edge case generalizes better than more prose."}], "correct": "D", "select": 1, "group": "I"}, {"id": "f4-051", "domain": 4, "task_id": "4.1", "objective": "Design prompts with explicit criteria to improve precision and reduce false positives", "situation": "An architect is structuring a long review prompt that defines separate criteria for security, correctness, and style categories, each with its own inclusion rules and severity examples.", "question": "Which structuring approach best helps Claude apply the right criteria to the right category without cross-contamination?", "options": [{"letter": "A", "text": "Write all criteria as one continuous paragraph of plain prose, trusting that clear sentence structure alone will keep the categories distinct in the model's interpretation.", "correct": false, "explanation": "A single continuous paragraph without structural markers is more prone to the model blending or misapplying criteria across categories compared to explicitly tagged sections."}, {"letter": "B", "text": "Wrap each category's criteria and examples in its own uniquely named XML tag, such as <security_criteria> and <correctness_criteria>, so the boundaries between categories are unambiguous.", "correct": true, "explanation": "Wrapping each category's rules and examples in its own descriptively named XML tag gives the model clear, unambiguous boundaries between categories, which is the recommended way to structure prompts that mix multiple sets of instructions."}, {"letter": "C", "text": "Repeat the full text of every category's criteria at the start of each category's section, so each section is self-contained even if it duplicates content.", "correct": false, "explanation": "Duplicating every category's full criteria in every section adds unnecessary length and redundancy without addressing the actual boundary-clarity problem, which structured tags solve directly."}, {"letter": "D", "text": "List every category's criteria in a single unordered bullet list without headers, relying on bullet order to imply which criteria belong to which category.", "correct": false, "explanation": "Relying on bullet order alone to imply category membership is fragile and ambiguous; there is no explicit marker tying a given bullet to a specific category."}], "correct": "B", "select": 1, "group": "I"}, {"id": "f4-052", "domain": 4, "task_id": "4.4", "objective": "Implement validation, retry, and feedback loops for extraction quality", "situation": "Over several months, a code-review assistant's structured findings include a detected_pattern field, and the team aggregates dismissal rates by pattern value. They discover that findings tagged with detected_pattern \"decorator-wrapped test fixture\" are dismissed over 90% of the time.", "question": "What is the primary value this feedback loop provides?", "options": [{"letter": "A", "text": "It pinpoints one over-triggering pattern so the detection rule can be tuned or suppressed for that construct specifically", "correct": true, "explanation": "Aggregating dismissals by detected_pattern isolates which specific construct is generating disproportionate false positives, giving the team an actionable, targeted signal for tuning that rule rather than adjusting the system broadly."}, {"letter": "B", "text": "It confirms that developers dismiss findings at random and that the review process should therefore be discontinued entirely", "correct": false, "explanation": "A concentrated 90% dismissal rate tied to one specific pattern is a non-random, actionable signal, the opposite of random noise, and argues for refining the process rather than abandoning it."}, {"letter": "C", "text": "It proves the extraction tool's JSON schema itself has a syntax defect that must be patched before the next release", "correct": false, "explanation": "A high dismissal rate for one pattern reflects a detection-logic tuning issue, not a JSON schema syntax defect; schema syntax and dismissal-pattern analysis are unrelated concerns."}, {"letter": "D", "text": "It lets the team automatically close every future finding across all categories without any developer review at all", "correct": false, "explanation": "The high dismissal rate applies specifically to one pattern, not to all findings, so blanket auto-closure across every category would suppress legitimate findings unrelated to that pattern."}], "correct": "A", "select": 1, "group": "I"}, {"id": "f4-053", "domain": 4, "task_id": "4.5", "objective": "Design efficient batch processing strategies", "situation": "An engineer submits a batch of 5,000 requests in a fixed order and, when results come back, zips the results array with the original input list by position, assuming the first result corresponds to the first request submitted. QA later finds several summaries attached to the wrong source document.", "question": "What is the root cause and correct fix?", "options": [{"letter": "A", "text": "The batch contained more than 1,000 requests, which is the point at which the API begins reordering results, so the fix is capping every batch at 1,000 requests", "correct": false, "explanation": "There is no request-count threshold at which the API begins reordering results; ordering is simply never guaranteed, regardless of batch size."}, {"letter": "B", "text": "The results file was read before the batch fully finished processing, so the fix is polling the status endpoint longer before reading any result content", "correct": false, "explanation": "Reading results only after the batch reaches an ended state does not address the underlying issue that result order does not mirror submission order."}, {"letter": "C", "text": "Batch results are not guaranteed to return in submission order, so the engineer must match each result to its request using the shared custom_id rather than list position", "correct": true, "explanation": "The Message Batches API does not guarantee results come back in the same order requests were submitted; custom_id is the documented mechanism for correlating each result to its originating request."}, {"letter": "D", "text": "The original request list must have contained a duplicate document, so the fix is deduplicating inputs before submission rather than changing how results are matched", "correct": false, "explanation": "A duplicate input document would not by itself cause summaries to attach to the wrong document; the mismatch traces to relying on position instead of custom_id."}], "correct": "C", "select": 1, "group": "H"}, {"id": "f4-054", "domain": 4, "task_id": "4.2", "objective": "Apply few-shot prompting to improve output consistency and quality", "situation": "A prompt combines lengthy background context, formatting instructions, several worked examples, and the user's actual request into a single message. The model occasionally treats part of a worked example as if it were the live user request, producing an oddly literal response to sample data instead of the real query.", "question": "What change would most directly resolve this confusion?", "options": [{"letter": "A", "text": "Wrap each example in its own example tag, and the whole set in an outer examples tag", "correct": true, "explanation": "Wrapping examples in dedicated tags gives the model an unambiguous structural signal for where sample content starts and ends, which is the recommended way to prevent the model from confusing example content with the live instructions or request."}, {"letter": "B", "text": "Move all worked examples to the end of the prompt, right after the user's actual request", "correct": false, "explanation": "Placing examples immediately before the live request without any structural marker does not resolve the ambiguity; the model can still fail to distinguish where the last example ends and the real request begins."}, {"letter": "C", "text": "Remove the examples and describe their content in a summary paragraph up front", "correct": false, "explanation": "Summarizing examples in prose forfeits the concrete precedent that makes few-shot examples effective in the first place, trading a structural-confusion problem for a return to the instructions-only inconsistency the team was trying to avoid."}, {"letter": "D", "text": "Rewrite each example as a short bullet point instead of a full input-output pair", "correct": false, "explanation": "Turning full examples into terse bullet points would likely reduce their effectiveness at demonstrating the desired output format, since the model no longer sees a complete example of the target structure, and it still does not clearly demarcate example from request."}], "correct": "A", "select": 1, "group": "I"}, {"id": "f4-055", "domain": 4, "task_id": "4.5", "objective": "Design efficient batch processing strategies", "situation": "A platform team runs an automated code-review gate that blocks a pull request from merging until Claude returns a verdict on the diff, typically within a few seconds.", "question": "Which approach should they use for this workflow?", "options": [{"letter": "A", "text": "Either API works equally well here, because Message Batches results are typically available in under a minute for small request volumes", "correct": false, "explanation": "Batch requests carry no latency guarantee at all, so treating typical fast completion as reliable for a blocking gate is unsafe."}, {"letter": "B", "text": "The Message Batches API, since batching the diff review still returns a verdict well within the few-second window merge gates require", "correct": false, "explanation": "The Batches API has no SLA guaranteeing a few-second turnaround; individual requests can sit for hours during periods of high demand."}, {"letter": "C", "text": "The Message Batches API, since the 50% cost discount outweighs the small delay a blocking merge gate would experience while waiting", "correct": false, "explanation": "A blocking merge gate cannot tolerate an unbounded wait; the cost savings do not offset stalling every pull request pipeline."}, {"letter": "D", "text": "The synchronous Messages API, since the merge gate blocks on an immediate response and the Message Batches API offers no guaranteed latency SLA", "correct": true, "explanation": "Pre-merge checks block on a fast response, and the Batches API has no guaranteed turnaround time (it can take up to 24 hours), so the synchronous API is required."}], "correct": "D", "select": 1, "group": "H"}, {"id": "f4-056", "domain": 4, "task_id": "4.6", "objective": "Design multi-instance and multi-pass review architectures", "situation": "A refactor touches 60 files. A single reviewer instance given the entire diff at once produces contradictory findings between files.", "question": "What change to the review architecture best addresses this?", "options": [{"letter": "A", "text": "Split the review into per-file passes for local issues, plus an integration pass for cross-file consistency", "correct": true, "explanation": "Splitting into per-file local passes plus a separate cross-file integration pass avoids the attention dilution and contradictory findings that come from reviewing everything in one undifferentiated pass."}, {"letter": "B", "text": "Run the same single reviewer instance twice over the full diff, and keep only findings that appear in both runs", "correct": false, "explanation": "Running the same overloaded single pass twice does not fix attention dilution; it just repeats the same structural problem."}, {"letter": "C", "text": "Ask the generator to make smaller, sequential commits, and review only the most recent commit in full each time", "correct": false, "explanation": "Reviewing only the latest commit in isolation loses the cross-file view needed to catch integration issues across the full 60-file change."}, {"letter": "D", "text": "Give the single reviewer instance a much larger context window so it can hold the whole diff in memory during one pass", "correct": false, "explanation": "A larger context window doesn't resolve attention dilution across unrelated files or eliminate contradictory findings between them."}], "correct": "A", "select": 1, "group": "F"}, {"id": "f4-057", "domain": 4, "task_id": "4.4", "objective": "Implement validation, retry, and feedback loops for extraction quality", "situation": "A purchase-order extractor uses a strict JSON schema, so every response has correctly typed fields and no missing keys. On one order, the extracted line-item amounts sum to $940 while the extracted \"order_total\" field reads $980. Both values are individually valid against the schema.", "question": "How should this discrepancy be classified and handled?", "options": [{"letter": "A", "text": "As a semantic error the schema cannot catch, requiring a check that compares a calculated_total against the stated order_total", "correct": true, "explanation": "Schema validation only guarantees type and structural conformance; a mismatch between two independently valid values is a semantic error, best caught by extracting a calculated_total and comparing it to the stated_total."}, {"letter": "B", "text": "As a schema syntax error, since the two numeric fields disagree with each other despite both matching their declared types here", "correct": false, "explanation": "A schema syntax error means a value violates its declared type or required structure; here both fields are individually well-formed, so the disagreement is semantic, not syntactic."}, {"letter": "C", "text": "As a tool-input validation failure that strict schema enforcement should already have blocked before it was returned", "correct": false, "explanation": "Strict schema and tool-input validation confirm structure and type, not that two numeric fields are mutually consistent; this class of error passes structural validation by definition."}, {"letter": "D", "text": "As a transient sampling artifact unlikely to recur, so no additional application-level check is really needed here", "correct": false, "explanation": "Numeric inconsistencies between totals and line items are a recurring class of extraction risk on real documents, not a one-off artifact, so a systematic comparison check is warranted."}], "correct": "A", "select": 1, "group": "I"}, {"id": "f4-058", "domain": 4, "task_id": "4.4", "objective": "Implement validation, retry, and feedback loops for extraction quality", "situation": "A financial-reconciliation pipeline retries an extraction ten times because the extracted transaction list never sums to the extracted statement total. Investigation reveals that the bank statement itself contains a genuine arithmetic error introduced by the issuing bank.", "question": "What should the pipeline do once this is discovered?", "options": [{"letter": "A", "text": "Switch the extraction schema to omit the statement total field so this mismatch can no longer be detected", "correct": false, "explanation": "Removing the total field would hide the inconsistency rather than resolve it, defeating the purpose of the validation check that surfaced the problem."}, {"letter": "B", "text": "Increase max_tokens on every retry, assuming the mismatch is caused by truncation before the total was written", "correct": false, "explanation": "The scenario establishes the cause as a genuine bank arithmetic error, not response truncation, so raising max_tokens does not address the actual root cause."}, {"letter": "C", "text": "Continue retrying indefinitely, since enough attempts will eventually make the model's numbers sum correctly overall anyway", "correct": false, "explanation": "Retrying indefinitely wastes resources chasing a mismatch that stems from the source data, not from model output variance, and will not converge on a consistent answer."}, {"letter": "D", "text": "Stop retrying, since the mismatch comes from an inconsistency in the source rather than a correctable extraction mistake", "correct": true, "explanation": "When the root cause is an inconsistency in the source document itself, no retry can produce line items that both faithfully reflect the source and sum to a total the source itself gets wrong; the correct action is to stop and flag it for human review."}], "correct": "D", "select": 1, "group": "I"}, {"id": "f4-059", "domain": 4, "task_id": "4.1", "objective": "Design prompts with explicit criteria to improve precision and reduce false positives", "situation": "A prompt engineer tries to fix a noisy security-findings category by adding the line \"only report high-confidence findings\" to the system prompt. After a week of testing, the false positive rate is essentially unchanged.", "question": "What is the most likely explanation for why this change failed to improve precision?", "options": [{"letter": "A", "text": "The word \"confidence\" is not in the set of tokens the model is trained to parse for output constraints, so the instruction is treated as decorative text and ignored, leaving the original behavior unchanged.", "correct": false, "explanation": "The model's vocabulary does not cause it to ignore instructions containing 'confidence'; the phrase is processed normally, but it fails to change behavior because it does not provide a concrete, actionable constraint."}, {"letter": "B", "text": "General confidence language gives the model no concrete rule for what to report, so it still applies the same underlying judgment that produced the false positives before the change.", "correct": true, "explanation": "Vague confidence terms give the model no clear, checkable rule for distinguishing true findings from false positives, so it continues applying the same underlying judgment. Precision only improves when explicit, categorical criteria replace subjective confidence thresholds."}, {"letter": "C", "text": "High-confidence phrasing conflicts with the model's safety training, which is designed to avoid under-reporting risks, causing it to over-report findings as a precautionary default across a wider range of inputs.", "correct": false, "explanation": "The model does not have a safety-training conflict that forces over-reporting; it does not default to over-reporting risks regardless of instructions. The instruction's vagueness, not a safety override, explains why false positives persisted."}, {"letter": "D", "text": "Adding any qualifier to a system prompt increases output length, which expands the set of tokens the evaluator inspects and independently raises the chance that a finding is miscategorized as high severity.", "correct": false, "explanation": "The unchanged false positive rate is not caused by increased output length or a larger token set. The added qualifier fails to improve precision because it lacks concrete criteria, not because it makes the output longer."}], "correct": "B", "select": 1, "group": "I"}, {"id": "f4-060", "domain": 4, "task_id": "4.3", "objective": "Enforce structured output using tool use and JSON schemas", "situation": "A team building a resume-parsing tool wants to guarantee that structured candidate data is extracted via a parse_resume tool on the current turn. They also want to know whether Claude can include natural-language reasoning about ambiguous resume sections before that tool call.", "question": "Which tool_choice configuration should be used to guarantee the parse_resume call, and what does Anthropic documentation state about natural-language commentary before a forced tool call?", "options": [{"letter": "A", "text": "tool_choice: {\"type\": \"any\"}, because any allows Claude to freely mix natural-language commentary with the forced tool call in the same response", "correct": false, "explanation": "any forces Claude to use one of the provided tools, but does not force the specific parse_resume tool; Claude may choose a different tool if multiple are available. Moreover, forced tool use (including any) prefills the assistant message to force a tool call, which suppresses any natural-language commentary or explanation before the tool_use content block, contradicting the claim that commentary can be freely mixed."}, {"letter": "B", "text": "tool_choice: {\"type\": \"tool\", \"name\": \"parse_resume\"}, because this is the documented way to force the specific tool; the trade-off is that forced tool use suppresses natural-language text before the tool call", "correct": true, "explanation": "According to Anthropic documentation, setting tool_choice to {\"type\": \"tool\", \"name\": \"parse_resume\"} explicitly forces Claude to invoke the named tool. The documentation states that when tool_choice is tool or any, the API prefills the assistant message to force a tool use, meaning the model will not emit a natural language response or explanation before the tool_use content block, even if explicitly asked. This satisfies the guarantee of the parse_resume call but confirms the limitation that no natural-language reasoning can appear before the forced tool call."}, {"letter": "C", "text": "tool_choice: {\"type\": \"none\"}, so Claude can freely decide in text whether to also produce a parse_resume tool call afterward", "correct": false, "explanation": "tool_choice: {\"type\": \"none\"} explicitly prevents Claude from using any tools, so a parse_resume tool call will not be made at all. This is the opposite of the requirement to guarantee structured candidate data extraction via the tool. Natural-language text may be produced, but no tool call will follow."}, {"letter": "D", "text": "tool_choice: {\"type\": \"auto\"}, combined with an explicit user-message instruction to use the parse_resume tool and share any relevant reasoning as text", "correct": false, "explanation": "auto is the default behavior; Claude decides whether to call any provided tool based on the request and tool descriptions. An explicit instruction may influence the model, but it does not guarantee the parse_resume tool will be called. Natural-language reasoning may occur if Claude chooses to respond in text instead of using a tool, which fails the requirement for guaranteed structured extraction."}], "correct": "B", "select": 1, "group": "H"}, {"id": "f4-061", "domain": 4, "task_id": "4.3", "objective": "Enforce structured output using tool use and JSON schemas", "situation": "A logistics company ingests shipment confirmation emails from many different carriers. Dates appear as 03/14/2026, 14-Mar-2026, and 2026.03.14 depending on the carrier, but the extraction schema defines ship_date as a string with a strict ISO 8601 pattern. Extractions frequently fail schema validation because the source dates don't match the expected format.", "question": "According to the current Anthropic official guidance, what is the most effective fix?", "options": [{"letter": "A", "text": "Split ship_date into three fields such as ship_date_us, ship_date_eu, and ship_date_iso, each expecting a different carrier date format, and populate only the one matching the extracted string.", "correct": false, "explanation": "Splitting into ship_date_us, ship_date_eu, and ship_date_iso still requires the model to correctly classify which format a given source string is in before it can pick the right field, so the same misclassification risk that causes today's validation failures is just relocated rather than removed — and it pushes the job of figuring out which field is populated onto every downstream consumer instead of giving them one normalized value."}, {"letter": "B", "text": "Retain the strict ISO 8601 schema constraint for ship_date and add an explicit description in the JSON schema telling Claude to parse and normalize the carrier date string to ISO 8601 format (e.g., YYYY-MM-DD).", "correct": false, "explanation": "Anthropic's structured outputs documentation lists a fixed set of supported string formats (date-time, time, date, duration, email, hostname, uri, ipv4, ipv6, uuid) and warns that pattern-based constraints have limits: simple regex patterns work well, but complex patterns can trigger 400 errors. Nothing in that feature set lets a schema accept several source formats on input while emitting only ISO 8601 on output — a description merely asks the model to reinterpret free-form carrier text (03/14/2026, 14-Mar-2026, 2026.03.14) correctly on every single call, with no fallback if it misreads an ambiguous format. Tightening the instructions targets the model's judgment, not the actual cause of the validation failures."}, {"letter": "C", "text": "Remove the ship_date field from the extraction schema and infer the shipment date later from other fields such as tracking number lookup or email metadata.", "correct": false, "explanation": "The source emails already state a shipment date explicitly; dropping the field trades a formatting problem for a missing-data problem, since a tracking-number lookup or email-metadata inference is not guaranteed to exist or to be accurate for every carrier. Nothing about the failure described — a format mismatch — is fixed by removing the value that failed to match."}, {"letter": "D", "text": "Loosen the schema to accept any string for ship_date, and add a downstream step that uses a date parser to normalize the value to ISO 8601 format before storing it in the database.", "correct": true, "explanation": "Loosening ship_date to a plain string lets extraction succeed no matter which carrier format appears in the source email, and moving the format conversion to a downstream date-parsing step means the transformation from '14-Mar-2026' or '2026.03.14' into ISO 8601 is applied by deterministic, testable code rather than depending on the model correctly inferring the source format under a strict pattern on every call. It also keeps the original string on hand for auditing if a parse ever needs to be checked by hand."}], "correct": "D", "select": 1, "group": "H"}, {"id": "f4-062", "domain": 4, "task_id": "4.4", "objective": "Implement validation, retry, and feedback loops for extraction quality", "situation": "A tax-document extractor needs a dependent's Social Security number, but the field has been physically redacted with a black marker on the scanned form supplied to the pipeline. Repeated retries with detailed error feedback still return an empty field.", "question": "What does this situation illustrate?", "options": [{"letter": "A", "text": "A tool-input validation failure that will resolve itself once the model is given a strict schema for the SSN field", "correct": false, "explanation": "Strict schema enforcement governs structure and type, not whether a genuinely absent value can be conjured from a redacted image; adding schema strictness will not surface the redacted number."}, {"letter": "B", "text": "A limit of retries: when required information is genuinely absent from the source, feedback retries cannot recover it", "correct": true, "explanation": "Retries address format and structural mistakes, not the fundamental absence of information in the source document; a redacted field cannot be recovered no matter how the retry prompt is worded."}, {"letter": "C", "text": "A prompt-caching issue where the redacted value was cached in a prior turn and needs to be evicted before retrying", "correct": false, "explanation": "The scenario describes a data-availability limitation in the source document, not a caching artifact, so cache eviction is unrelated to the root cause."}, {"letter": "D", "text": "A schema syntax error that structured output enforcement should have already eliminated before it reached this layer entirely", "correct": false, "explanation": "This has nothing to do with JSON syntax or type conformance; the field is empty because the underlying information is physically unavailable in the document, not because the output violates the schema."}], "correct": "B", "select": 1, "group": "I"}, {"id": "f4-063", "domain": 4, "task_id": "4.3", "objective": "Enforce structured output using tool use and JSON schemas", "situation": "A document-processing service receives files that could be invoices, resumes, or contracts, but the type is not known ahead of time. The service defines three separate extraction tools (extract_invoice, extract_resume, extract_contract) and needs Claude to always call exactly one of them so the pipeline never falls back to plain text.", "question": "Which tool_choice configuration should be used?", "options": [{"letter": "A", "text": "Set tool_choice to {\"type\": \"auto\"} so Claude evaluates the document and decides whether calling a tool is appropriate", "correct": false, "explanation": "\"auto\" is the default and permits Claude to respond with plain text instead of calling any tool, which does not guarantee structured output."}, {"letter": "B", "text": "Set tool_choice to {\"type\": \"tool\", \"name\": \"extract_invoice\"} so the same extraction tool always runs regardless of document type", "correct": false, "explanation": "Forcing a single named tool would incorrectly run invoice extraction even on resumes or contracts, since the document type is unknown in advance."}, {"letter": "C", "text": "Omit the tools parameter and instruct Claude in the system prompt to always respond using one of the three named JSON shapes", "correct": false, "explanation": "Without the tools parameter there is no schema enforcement at all; Claude would be generating free-form text that happens to look like JSON, reintroducing syntax risk."}, {"letter": "D", "text": "Set tool_choice to {\"type\": \"any\"} so Claude must call one of the three tools but can pick whichever matches the document", "correct": true, "explanation": "tool_choice: \"any\" forces Claude to call some tool but leaves the choice of which tool up to the model, which is exactly what's needed when the document type is unknown but a tool call must always happen."}], "correct": "D", "select": 1, "group": "H"}, {"id": "f4-064", "domain": 4, "task_id": "4.1", "objective": "Design prompts with explicit criteria to improve precision and reduce false positives", "situation": "A reviewer prompt currently says: \"Only surface issues you are very sure about.\" An architect wants to replace confidence-based filtering with categorical criteria for a bug-detection category specifically.", "question": "Which rewrite achieves that goal?", "options": [{"letter": "A", "text": "Report an issue only if you would personally be willing to bet that a senior engineer on the team would agree it is a genuine problem, after reviewing the code against the team's implicit quality standards and typical bug histories.", "correct": false, "explanation": "This rewording still relies on a subjective confidence bet and hypothetical social validation rather than defining objective, verifiable conditions for a bug. It does not replace internal certainty with explicit categorical criteria."}, {"letter": "B", "text": "Report an issue only when a variable, argument, or return value is used in a way that contradicts its declared type, documented contract, or an explicit precondition stated elsewhere in the code.", "correct": true, "explanation": "This option defines specific, checkable conditions—usage that contradicts a declared type, documented contract, or explicit precondition—that constitute a bug. It replaces subjective certainty with an objective, externally verifiable rule."}, {"letter": "C", "text": "Report an issue whenever your certainty about it being a real bug is above a threshold you judge to be reasonably high for this kind of codebase, calibrated against the defect density you expect for the subsystem.", "correct": false, "explanation": "This remains a confidence-based filter, just with a calibrated threshold, without defining categorical criteria for a bug. It only adjusts the level of subjective certainty, not the basis for identifying an issue."}, {"letter": "D", "text": "Report an issue whenever the surrounding code looks unusual compared to typical patterns you have seen in similar production systems, focusing on deviations from standard naming conventions and control flow idioms.", "correct": false, "explanation": "This approach uses vague heuristics about code unusualness and deviations from common patterns, which are not explicit, checkable criteria for bugs. It fails to establish a categorical rule grounded in concrete code properties."}], "correct": "B", "select": 1, "group": "I"}, {"id": "f4-065", "domain": 4, "task_id": "4.6", "objective": "Design multi-instance and multi-pass review architectures", "situation": "An architect defines a 'code-reviewer' subagent and invokes it from the main agent immediately after code generation.", "question": "Which statement accurately describes what context the reviewer subagent starts with?", "options": [{"letter": "A", "text": "The subagent inherits the parent's reasoning trace but not the actual code files, so the parent must re-describe the implementation in prose", "correct": false, "explanation": "The subagent does not inherit any reasoning trace from the parent; its context is a fresh conversation seeded only by its own prompt and the invocation prompt."}, {"letter": "B", "text": "The subagent automatically receives the parent's entire conversation transcript, including every tool call and result from the generation phase", "correct": false, "explanation": "The parent's conversation history and tool results are explicitly not passed to the subagent; only the Agent tool's prompt string carries information across."}, {"letter": "C", "text": "The subagent starts with its own system prompt plus the Agent tool's prompt string, but not the parent's history or tool results", "correct": true, "explanation": "A subagent's context window starts fresh; it receives its own system prompt and the Agent tool's prompt string, but not the parent's conversation history or tool results."}, {"letter": "D", "text": "The subagent shares the same context window as the parent, so any file the parent read during generation is already visible to the subagent", "correct": false, "explanation": "Subagents run in their own separate context, not a shared context window with the parent."}], "correct": "C", "select": 1, "group": "F"}, {"id": "f4-066", "domain": 4, "task_id": "4.3", "objective": "Enforce structured output using tool use and JSON schemas", "situation": "A developer is choosing between prompting Claude to \"return only a JSON object matching this format\" versus defining a tool with an input_schema and letting Claude populate it via tool_use, with strict enforcement enabled. Both approaches are tested against the same messy scanned-document corpus.", "question": "Which outcome should the developer expect regarding guaranteed schema compliance?", "options": [{"letter": "A", "text": "The prompt-only approach yields more reliable schema-compliant output because it avoids the additional system-prompt instructions introduced by tool definitions, allowing the model to focus directly on the JSON format constraints.", "correct": false, "explanation": "Prompt-only approaches are actually less reliable because they lack any schema validation or enforcement. Tool definitions with strict enforcement are designed precisely to guarantee schema compliance, and any additional system prompt instructions are negligible compared to the benefits of server-side validation."}, {"letter": "B", "text": "Both approaches produce equally reliable schema-compliant output because the language model interprets the format specification identically in each case, generating token sequences that conform to the JSON structure with equal consistency.", "correct": false, "explanation": "This is incorrect. Prompting the model to output JSON does not provide any enforcement mechanism; the model may produce invalid syntax or deviate from the schema. In contrast, strict tool use actively constrains token generation to comply with the defined input_schema, making it far more reliable. The two approaches do not offer equal consistency."}, {"letter": "C", "text": "The tool_use approach reliably produces schema-compliant structured data because the API strictly enforces the input_schema server-side when strict tool use is enabled, while prompt-only JSON requests can still drift into invalid syntax or missing fields.", "correct": true, "explanation": "When tool definitions include strict: true, Anthropic's API guarantees that the arguments match the tool's input_schema through constrained decoding. This server-side enforcement eliminates parsing errors, unlike prompt-only JSON requests, which are considered an unreliable anti-pattern that often results in malformed JSON. The native Structured Outputs feature (output_config.format) offers a similar guarantee for direct JSON output, but for this comparison the strict tool_use approach is the reliable choice."}, {"letter": "D", "text": "Neither approach can guarantee valid structured output when processing messy scanned documents, so a separate JSON-repair library must always be used afterward to correct syntax errors and missing fields, regardless of which method is chosen.", "correct": false, "explanation": "With strict tool use, the model's output is constrained to match the input_schema, ensuring valid, parseable JSON even for messy scanned documents (though content accuracy is not guaranteed). No separate repair library is needed for syntax or missing fields, because the API enforcement eliminates those issues."}], "correct": "C", "select": 1, "group": "H"}, {"id": "f4-067", "domain": 4, "task_id": "4.3", "objective": "Enforce structured output using tool use and JSON schemas", "situation": "A real estate platform extracts property listings from scraped web pages using a single describe_property tool. The square_footage field is defined as a required number, but many older listings state size only in vague prose like \"spacious with room to grow\" and never give a numeric figure. Extraction logs show the model consistently inventing plausible square footage values for these listings.", "question": "Which two schema changes together best resolve this while preserving data quality for downstream reports?", "options": [{"letter": "A", "text": "Remove square_footage from the schema entirely, and rely on a separate keyword-search script to scan raw listing text for numeric patterns and inject the first match into a staging column for reports.", "correct": false, "explanation": "Removing the field from the schema discards the reliability of tool-use extraction for this data point, and relying on a separate keyword-search script reintroduces brittle ad hoc pattern matching, which is less accurate and maintainable."}, {"letter": "B", "text": "Keep square_footage required, but change its type to string so the model can output a placeholder like \"unspecified\" or \"N/A\" instead of fabricating a number, ensuring the field is always present.", "correct": false, "explanation": "Keeping square_footage as required still forces the model to produce a value, which may lead to placeholders that are not useful for numeric reports. Changing the type to string breaks downstream numeric analysis that expects actual numbers when they are available."}, {"letter": "C", "text": "You can make square_footage optional for missing values and add a square_footage_source enum that stores 'stated', 'estimated', or 'unknown' so reports can separate confirmed from absent values.", "correct": true, "explanation": "Making square_footage optional eliminates the requirement to always output a number, preventing fabrications for listings without numeric size data. Adding the square_footage_source enum allows reports to distinguish between confirmed measurements, estimates, and unknown values, preserving data quality for downstream use."}, {"letter": "D", "text": "Keep square_footage required, and add a system prompt instruction (e.g., \"Do not guess; output 'N/A' for missing data\") and a configuration flag to require manual review of any numeric output.", "correct": false, "explanation": "Even with a system prompt instruction not to guess, the schema still defines square_footage as a required number, compelling the model to output a numeric value. The manual review flag adds overhead but doesn't prevent fabricated numbers from being generated initially."}], "correct": "C", "select": 1, "group": "H"}, {"id": "f4-068", "domain": 4, "task_id": "4.4", "objective": "Implement validation, retry, and feedback loops for extraction quality", "situation": "A team building a code-review assistant is deciding how granular to make the detected_pattern field on each structured finding. One option records only a broad category like \"security\" for every finding; another records the specific triggering construct, such as the exact function name, decorator, or regex rule that fired.", "question": "Which choice better supports long-term analysis of developer dismissal patterns?", "options": [{"letter": "A", "text": "Neither option matters much, since dismissal rates should be analyzed through the severity field instead", "correct": false, "explanation": "Severity indicates how serious a finding is, not which specific construct or rule triggered it, so it cannot substitute for a pattern field when the goal is root-cause dismissal analysis."}, {"letter": "B", "text": "The specific-construct option, since it lets the team isolate which exact rule causes a high dismissal rate and tune it", "correct": true, "explanation": "A specific, construct-level detected_pattern value lets analysts pinpoint exactly which rule or code pattern is over-triggering, so the team can tune that particular rule instead of guessing across an entire broad category."}, {"letter": "C", "text": "The broad-category option, since fewer distinct values are easier for a dashboard to render without extra grouping logic", "correct": false, "explanation": "Rendering convenience does not outweigh the analytical value of granularity; a broad category collapses many distinct root causes into one bucket, making targeted tuning impossible."}, {"letter": "D", "text": "The broad-category option, since recording specific constructs would expose proprietary rule names to the team", "correct": false, "explanation": "The review team is the intended consumer of this data for tuning their own rules, so there is no confidentiality barrier that would justify withholding construct-level detail from them."}], "correct": "B", "select": 1, "group": "I"}, {"id": "f4-069", "domain": 4, "task_id": "4.1", "objective": "Design prompts with explicit criteria to improve precision and reduce false positives", "situation": "A team's automated PR-review prompt currently instructs Claude to \"check that comments are accurate.\" The category produces a high volume of false positives on trivial phrasing nitpicks, and developers have started ignoring its output. An architect is rewriting the instruction to raise precision.", "question": "Which replacement instruction best applies the principle of explicit criteria over vague instructions?", "options": [{"letter": "A", "text": "Ask Claude to only flag a comment when it is highly confident the comment makes a factual error about the code, such as claiming a method does not exist when it is clearly present in the codebase, and to ignore borderline cases.", "correct": false, "explanation": "Asking Claude to be 'highly confident' relies on the model's internal certainty rather than providing a concrete, verifiable rule. This confidence-based filter does not define explicit criteria for what constitutes a factual error, so it fails to reliably improve precision."}, {"letter": "B", "text": "Tell Claude to flag any comment that could plausibly be improved in clarity, completeness, or consistency with the team's style guide, such as an ambiguous phrase that might confuse a reader, and to suggest a clearer version.", "correct": false, "explanation": "This instruction is broader and vaguer than the original, explicitly inviting style and clarity nitpicks rather than narrowing scope to factual contradictions. It would increase false positives, not reduce them."}, {"letter": "C", "text": "Flag a comment only when it makes a specific claim about behavior that is contradicted by what the code actually does, such as a docstring stating a function returns None when it always returns a value.", "correct": true, "explanation": "This defines a concrete, checkable criterion (a factual contradiction between a stated claim and observed code behavior) with a worked example, which is exactly the kind of explicit categorical rule that reduces false positives compared to a vague instruction like 'check that comments are accurate.'"}, {"letter": "D", "text": "Instruct Claude to evaluate each comment and flag only those that it deems significant enough to warrant developer attention, such as a comment that could cause a bug if misunderstood, and to ignore trivial wording differences.", "correct": false, "explanation": "This instruction delegates the decision to the model's subjective judgment of 'significance' rather than providing explicit criteria. It relies on the model to determine what warrants attention, which is the same failure mode as the original vague instruction, and does not improve precision."}], "correct": "C", "select": 1, "group": "I"}, {"id": "f4-070", "domain": 4, "task_id": "4.6", "objective": "Design multi-instance and multi-pass review architectures", "situation": "To save tokens, an architect has the generator instance write a summary of its own changes and passes that summary, not the raw diff, to the independent reviewer instance.", "question": "Why does this undermine the value of using a second instance?", "options": [{"letter": "A", "text": "The summary reflects the generator's own framing of its decisions, so the reviewer evaluates that account instead of examining the actual code fresh", "correct": true, "explanation": "A generator-written summary carries the generator's own justification and framing, so the reviewer is no longer examining the actual code independently but is instead anchored to the generator's account of it."}, {"letter": "B", "text": "A reviewer instance can only produce useful findings when it has access to the generator's extended thinking trace, and summaries never include that trace", "correct": false, "explanation": "Independent review instances do not require access to the generator's extended thinking trace; in fact, lacking that trace is exactly what makes them more effective."}, {"letter": "C", "text": "Token savings from summarizing are negligible compared to the cost of running a second instance, so the summarization step provides no benefit either way", "correct": false, "explanation": "The problem is not about whether token savings are worthwhile; it's that a summary substitutes the generator's biased account for the actual code."}, {"letter": "D", "text": "Passing a summary instead of the diff exceeds the maximum prompt length the Agent tool supports, so the reviewer instance would fail to start", "correct": false, "explanation": "Prompt length is not the reason this design fails; the issue is the loss of independent, fresh examination of the actual code."}], "correct": "A", "select": 1, "group": "F"}, {"id": "f4-071", "domain": 4, "task_id": "4.5", "objective": "Design efficient batch processing strategies", "situation": "A QA team wants to generate new regression test cases for 40,000 legacy modules once per night. No developer is waiting on the output, and the team only needs the results within 24 hours for review the following day. Cost efficiency is a priority.", "question": "Which approach best fits this workload?", "options": [{"letter": "A", "text": "The Message Batches API, since batching is required whenever a job processes more than a few hundred requests in one run.", "correct": false, "explanation": "There is no hard rule that batching is 'required' above a certain request count. You can use the synchronous API with parallel async calls for large volumes, but it would be more expensive and less optimized. The primary justification for choosing the Message Batches API is the combination of latency tolerance, the 50% cost discount, and the 24‑hour SLA that meets the team's deadline, not an arbitrary volume threshold."}, {"letter": "B", "text": "The synchronous Messages API, since nightly jobs should minimize total wall-clock time by streaming each response as soon as it is generated.", "correct": false, "explanation": "Streaming reduces the time-to-first-token but does not decrease the total processing time or cost for the entire workload. For 40,000 requests, the synchronous API would still be processing long after the results are needed, and it would cost twice as much as the Batch API. Streaming is beneficial when a user is waiting for interactive output, which is not the case here."}, {"letter": "C", "text": "The synchronous Messages API, since running each request in sequence guarantees every test case is ready before the night's job window closes.", "correct": false, "explanation": "While sequential execution ensures in-order completion, it is significantly slower and more expensive for 40,000 requests. The synchronous API charges standard per-token rates without the 50% batch discount, and processing requests one-by-one at this scale could easily exceed the available window. The team does not require instantaneous responses, so the cost savings and parallel processing of the Batch API are far more appropriate."}, {"letter": "D", "text": "The Message Batches API, since the workload is latency-tolerant and the 50% cost discount scales well across 40,000 requests, while the 24-hour SLA fits the team's deadline.", "correct": true, "explanation": "The Message Batches API is purpose-built for asynchronous, latency-tolerant bulk processing. It provides a 50% cost reduction on input and output tokens compared to the synchronous API, and its 24-hour SLA aligns with the team's requirement to have results ready within 24 hours for next-day review. By splitting the 40,000 requests into multiple batches (up to 10,000 per batch), the workload can be processed in parallel using spare capacity, making it the most cost-effective and appropriate choice."}], "correct": "D", "select": 1, "group": "H"}, {"id": "f4-072", "domain": 4, "task_id": "4.4", "objective": "Implement validation, retry, and feedback loops for extraction quality", "situation": "A utility-bill extraction pipeline has logged the following four distinct failed extractions: 1. The meter_reading value was extracted correctly but placed under billing_address instead of the usage_details object. 2. The account_holder_phone field is blank because no phone number appears anywhere on the scanned bill provided so far. 3. The prior_year_comparison figure is missing because it only appears in an annual letter never supplied to the pipeline. 4. The service_address field holds the mailing address because that is the only address printed on this particular bill.", "question": "Which of these is the one most likely to be fixed by an error-feedback retry, as opposed to requiring a different source document or human escalation?", "options": [{"letter": "A", "text": "The meter_reading value was extracted correctly but placed under billing_address instead of the usage_details object", "correct": true, "explanation": "This is a schema-adherence error, not a missing-information problem. Anthropic’s extraction guidance recommends explicitly defining the output schema (e.g., a usage_details object that contains meter_reading, while billing_address does not), adding clear instructions, and using few-shot examples to keep extracted fields in the correct object. An error-feedback retry can correct the mapping without needing a new source document or human escalation."}, {"letter": "B", "text": "The account_holder_phone field is blank because no phone number appears anywhere on the scanned bill provided so far", "correct": false, "explanation": "The blank field is caused by absent source data rather than an extraction or mapping error. A retry on the same bill cannot invent a phone number; the pipeline needs a document containing the phone number or human escalation to obtain it."}, {"letter": "C", "text": "The prior_year_comparison figure is missing because it only appears in an annual letter never supplied to the pipeline", "correct": false, "explanation": "This is an input coverage gap: the required data is not in the supplied document set. Error-feedback retry cannot generate the prior-year comparison from text it never received, so the fix is to supply the annual letter or escalate for the missing input."}, {"letter": "D", "text": "The service_address field holds the mailing address because that is the only address printed on this particular bill", "correct": false, "explanation": "The model extracted the only available address, so the field contains what the document provides. If the service address is truly required and distinct, a different source document or clarification is needed; retrying on the same input cannot recover an address that was never printed."}], "correct": "A", "select": 1, "group": "I"}, {"id": "f4-073", "domain": 4, "task_id": "4.2", "objective": "Apply few-shot prompting to improve output consistency and quality", "situation": "An extraction pipeline pulls dosage fields from clinical intake notes. Patients often describe amounts informally, such as 'a couple tablets' or 'about half a cup,' and the model sometimes fabricates a precise numeric value where the source text is genuinely vague. The team wants to reduce this fabrication without discarding informal-but-usable descriptions. What is the most effective prompt change?", "question": "Select the single best answer.", "options": [{"letter": "A", "text": "Add a post-processing step that rejects any value failing to match a strict numeric regular expression", "correct": false, "explanation": "A strict numeric regex discards informal-but-usable descriptions like 'a couple tablets' or 'about half a cup,' which the team explicitly wants to avoid. It also only filters outputs and does not prevent upstream fabrication, so it fails the requirement."}, {"letter": "B", "text": "Add examples pairing informal phrases with correct normalization, plus one case where the field is left null", "correct": true, "explanation": "Anthropic's prompt engineering guidance recommends few-shot examples to normalize informal inputs. For extraction, when a value is not present or is genuinely vague, instruct the model to return null (or add a confidence flag for human review). This reduces fabrication while preserving usable informal descriptions such as 'a couple tablets' by showing how to map them rather than rejecting them. This text is the only correct answer; any earlier grading mismatch due to shuffled display letters has been corrected."}, {"letter": "C", "text": "Instruct the model to always convert informal quantities into metric units for consistency across all patient records", "correct": false, "explanation": "Forcing metric conversion is not the recommended mitigation for fabrication and may discard or distort the original informal usage. Anthropic's guidance focuses on examples that normalize informal phrases and on setting missing or vague fields to null, rather than requiring conversions that could introduce additional errors."}, {"letter": "D", "text": "Change the schema so every dosage field is marked required, forcing the model to populate a value", "correct": false, "explanation": "Marking fields required forces the model to produce a value even when the source is genuinely vague, increasing fabrication rather than reducing it. Anthropic's recommended approach for structured extraction is to allow null for absent or uncertain information and to use flags for human review when confidence is low."}], "correct": "B", "select": 1, "group": "I"}, {"id": "f4-074", "domain": 4, "task_id": "4.6", "objective": "Design multi-instance and multi-pass review architectures", "situation": "A substantial multi-file change is nearing merge. The architect wants findings that are independently reproduced and verified before they're reported, and is willing to wait several minutes and spend usage credits for that assurance.", "question": "Which option best fits?", "options": [{"letter": "A", "text": "A single /review <pr> pass, since it applies the same independent verification as the cloud fleet but finishes in seconds", "correct": false, "explanation": "The /review command is a fast, local, single-pass review that does not utilize the cloud sandbox or multiple verification agents. It is not designed for independently reproducing and verifying findings, which is a feature exclusive to /code-review ultra."}, {"letter": "B", "text": "Repeating /code-review three separate times in one session, and keeping only the findings that appear in all three runs", "correct": false, "explanation": "Running the same local review multiple times does not constitute independent reproduction and verification. The model may have shared blind spots, and this approach does not benefit from the dedicated verification agents of /code-review ultra. Anthropic explicitly recommends the multi-agent verification system for high-confidence findings."}, {"letter": "C", "text": "/code-review ultra, since it runs a fleet of reviewer agents that independently reproduce and verify each finding", "correct": true, "explanation": "As per Anthropic's Claude Code documentation, /code-review ultra (or the alias /ultrareview) is a multi-agent automated PR review system that dispatches a fleet of reviewer agents in a cloud sandbox. Each finding is independently reproduced and verified by a separate set of agents, achieving a false positive rate below 1%. This aligns with the architect's requirements for independently verified findings, minutes-long wait, and credit expenditure."}, {"letter": "D", "text": "/code-review at the default effort level, since local reviews already independently reproduce and verify every finding", "correct": false, "explanation": "Standard /code-review without the ultra level performs a local, single-pass review and does not include the independent reproduction and verification step. That capability requires the dedicated cloud-based multi-agent architecture of /code-review ultra."}], "correct": "C", "select": 1, "group": "F"}, {"id": "f4-075", "domain": 4, "task_id": "4.2", "objective": "Apply few-shot prompting to improve output consistency and quality", "situation": "A customer-support routing agent must decide, for ambiguous tickets that mention both a billing and a technical keyword, whether to route to the billing queue or the technical queue. The team wants to add examples that will help the agent generalize its judgment to new ambiguous tickets it has not seen before, not just the exact tickets in the examples.", "question": "What should each example include?", "options": [{"letter": "A", "text": "The ticket text and the queue chosen, with no explanation, so the model infers the pattern from repetition", "correct": false, "explanation": "Providing only the final label without the reasoning gives the model a correct answer for that one ticket but no visible logic to transfer to a differently worded ambiguous ticket, which weakens generalization to novel cases."}, {"letter": "B", "text": "The ticket text, the queue chosen, and a short explanation of why it beat the other plausible queue", "correct": true, "explanation": "Including the reasoning for why one route was chosen over the other plausible one, not just the final label, is what lets the model generalize the underlying judgment to new ambiguous tickets rather than memorizing surface features of the specific examples shown."}, {"letter": "C", "text": "A simplified, idealized ticket rather than an actual historical one, chosen because it is easier to parse quickly", "correct": false, "explanation": "Using an idealized, simplified ticket rather than a realistic one reduces relevance to the actual use case, and the team's goal is explicitly to handle real ambiguous phrasing, not a cleaned-up approximation of it."}, {"letter": "D", "text": "The full text of the routing policy document, repeated in full once inside each individual example", "correct": false, "explanation": "Repeating the full policy document inside each example is redundant with the instructions and does not provide the concrete precedent of an actual decision and its rationale that makes an example effective."}], "correct": "B", "select": 1, "group": "I"}, {"id": "f4-076", "domain": 4, "task_id": "4.6", "objective": "Design multi-instance and multi-pass review architectures", "situation": "A large migration touches 200 files — too many to delegate one at a time through turn-by-turn requests inside a single conversation.", "question": "Which approach best scales the per-file review process to that volume?", "options": [{"letter": "A", "text": "Ask a single subagent to review all 200 files sequentially within one long-running conversation so results stay consistent", "correct": false, "explanation": "Using a single subagent for sequential review can lead to high token costs due to accumulating context, possible goal drift, and agentic laziness. It does not scale efficiently, and Anthropic advises against long-running single-agent conversations without structured delegation."}, {"letter": "B", "text": "Reduce the review to a random sample of 20 files and extrapolate those findings across all the remaining changed files", "correct": false, "explanation": "While sampling might reduce effort, it introduces risk by not reviewing all files. For a comprehensive migration review, Anthropic's recommended approach is to use parallel subagents to efficiently handle the full volume, ensuring thorough coverage without sacrificing accuracy."}, {"letter": "C", "text": "Move orchestration into a workflow tool that runs a script coordinating many subagents, instead of turn-by-turn delegation", "correct": true, "explanation": "Anthropic recommends moving to a workflow tool with a script that coordinates many subagents for complex, large-scale tasks. Dynamic Workflows in Claude Code allow generating orchestration scripts that delegate tasks, run subtasks in parallel, and validate results, providing better structure, parallelization, and cost management than relying on a single agent's unscripted turn-by-turn delegation."}, {"letter": "D", "text": "Increase the generator's own extended thinking effort so it reviews all 200 files itself before returning control to the architect", "correct": false, "explanation": "Increasing extended thinking on a single generator does not provide the parallelization or structured orchestration needed for 200 files. This approach would likely hit context limits and incur high latency, and is not a documented best practice for large-scale migrations."}], "correct": "C", "select": 1, "group": "F"}, {"id": "f4-077", "domain": 4, "task_id": "4.1", "objective": "Design prompts with explicit criteria to improve precision and reduce false positives", "situation": "A reviewer flags a docstring that says \"returns the cached value if present, otherwise fetches from the API,\" but the function under review always calls the API regardless of a cache.", "question": "Under an explicit-criteria rule that only flags comments contradicted by actual code behavior, should this finding be reported?", "options": [{"letter": "A", "text": "No, because docstrings describe intent rather than guaranteed behavior, so a mismatch with the current implementation is not a reportable contradiction.", "correct": false, "explanation": "This treats all docstrings as aspirational rather than descriptive, which is inconsistent with the explicit criterion of flagging comments that make a specific claim contradicted by actual behavior; a docstring stating concrete behavior is exactly the kind of claim that criterion covers."}, {"letter": "B", "text": "No, because caching behavior is an implementation detail, and implementation details are excluded from comment-accuracy review by definition.", "correct": false, "explanation": "The criterion is defined around whether the comment makes a specific claim contradicted by the code, not around whether the described behavior is an 'implementation detail'; caching behavior described in a docstring is a specific claim like any other."}, {"letter": "C", "text": "Yes, because the docstring makes a specific, checkable claim about caching behavior that the code's actual control flow directly contradicts.", "correct": true, "explanation": "The docstring asserts a specific, verifiable claim (that a cache is checked before an API call), and the code's actual control flow contradicts that claim, which is exactly the kind of concrete contradiction the explicit criterion is designed to catch."}, {"letter": "D", "text": "Yes, but only if the function is called from more than one place in the codebase, since single-use functions are exempt from this criterion.", "correct": false, "explanation": "The explicit criterion is based on whether the comment's claim is contradicted by the code's actual behavior, not on how many call sites the function has; call-site count is not part of the defined rule."}], "correct": "C", "select": 1, "group": "I"}, {"id": "f4-078", "domain": 4, "task_id": "4.6", "objective": "Design multi-instance and multi-pass review architectures", "situation": "An architect wants an independent review of generated code. To reuse setup, they resume the same subagent session that just finished generating the code, then ask it to review its own diff.", "question": "Does this qualify as an independent review instance?", "options": [{"letter": "A", "text": "Yes, because the Agent tool automatically strips reasoning context from a resumed session while keeping the generated files visible", "correct": false, "explanation": "Anthropic's Agent tool does not automatically strip reasoning context. When resuming, all context—including reasoning traces—is retained alongside the generated files. Manual intervention (e.g., /clear or /compact) is required to remove prior reasoning."}, {"letter": "B", "text": "No, because resumed sessions cannot access the codebase at all, so the reviewer would have no files to examine regardless of the retained context", "correct": false, "explanation": "Resumed sessions have full access to the codebase and workspace state from the previous session. The issue is not file access, but the persistence of the model's own reasoning history, which compromises independent review."}, {"letter": "C", "text": "No, because resuming the same session retains its full prior conversation history, the same self-review limitation independent instances avoid", "correct": true, "explanation": "According to Anthropic's documentation, resuming a subagent session restores the entire context—including messages, tool calls, and reasoning traces. This preserves the model's prior generation and decision-making, introducing a self-review bias. Independent review requires a fresh context, achieved by clearing history, compacting, or starting a new subagent instance."}, {"letter": "D", "text": "Yes, because resuming a session always clears the model's memory of prior tool calls even though the session id stays the same", "correct": false, "explanation": "Resuming a session does not clear memory; the full conversation history, including all tool calls and results, is restored. The session ID remains constant, but the context is completely preserved, not reset."}], "correct": "C", "select": 1, "group": "F"}, {"id": "f4-079", "domain": 4, "task_id": "4.2", "objective": "Apply few-shot prompting to improve output consistency and quality", "situation": "A document-extraction system pulls structured fields (vendor, total, line items) from invoices that arrive in wildly different layouts: some are tables, some are plain paragraphs, some split totals across multiple currencies. Detailed field-by-field instructions have not stopped the model from occasionally inventing plausible-looking values when a layout doesn't match what the team anticipated.", "question": "What is the best next step?", "options": [{"letter": "A", "text": "Preprocess every document into one canonical plain-text format before it reaches the prompt", "correct": false, "explanation": "Normalizing documents to plain text before extraction can lose structural cues, such as table alignment, that actually help the model locate fields correctly, and it does not address why the model fabricates values when uncertain."}, {"letter": "B", "text": "Add a small set of examples across the layout types actually received, each paired with its correct extraction", "correct": true, "explanation": "Examples that span the actual variety of document structures the pipeline encounters give the model demonstrated precedent for extracting from unfamiliar layouts accurately, which reduces the fabrication that comes from relying on instructions alone to cover every case."}, {"letter": "C", "text": "Write a longer paragraph enumerating every layout variant the team can currently think of and its handling rule", "correct": false, "explanation": "An ever-longer enumerated instruction list still cannot anticipate every layout variant in advance, and prose descriptions of format rules are generally less effective than shown examples at preventing fabrication on unseen structures."}, {"letter": "D", "text": "Reduce the number of required output fields so there are fewer chances to fabricate a value", "correct": false, "explanation": "Removing required fields reduces the surface area for fabrication but also reduces the pipeline's usefulness, and it sidesteps the actual problem of teaching correct behavior on varied layouts."}], "correct": "B", "select": 1, "group": "I"}, {"id": "f4-080", "domain": 4, "task_id": "4.4", "objective": "Implement validation, retry, and feedback loops for extraction quality", "situation": "A logistics team uses strict tool use with strict: true to extract shipment records, guaranteeing that every \"quantity\" field is a well-formed integer as required by the schema. An engineer asks whether this guarantee alone is sufficient to trust the extracted quantities for downstream inventory decisions.", "question": "What is the correct assessment?", "options": [{"letter": "A", "text": "No further check is needed here, but only because this particular field happens to be numeric rather than a string", "correct": false, "explanation": "The field's data type does not change what strict validation checks; numeric fields are just as susceptible to semantically implausible values as any other type."}, {"letter": "B", "text": "Yes, strict tool use guarantees full business-logic correctness, so no further validation of quantity values is required", "correct": false, "explanation": "Strict validation guarantees conformance to the declared schema, not business-logic correctness; a value can be a perfectly valid integer and still be wrong for the shipment in question."}, {"letter": "C", "text": "No, strict validation only guarantees structural and type conformance, so business-plausibility checks are still needed", "correct": true, "explanation": "Strict schema and tool-input validation eliminate type and structural errors, but they cannot confirm that a value is semantically reasonable, such as a negative or implausibly large quantity, so application-level semantic checks remain necessary."}, {"letter": "D", "text": "Yes, because strict mode internally re-runs the extraction against the source until values are semantically confirmed", "correct": false, "explanation": "Strict tool use enforces schema conformance at generation time through constrained decoding; it does not perform an internal semantic re-verification pass against the source document."}], "correct": "C", "select": 1, "group": "I"}, {"id": "f4-081", "domain": 4, "task_id": "4.4", "objective": "Implement validation, retry, and feedback loops for extraction quality", "situation": "A security review assistant extracts structured findings from source code and developers frequently dismiss a large share of the findings tied to one particular helper function used for input sanitization. The team wants to analyze this dismissal trend systematically over time.", "question": "What should be added to the structured finding schema to support this analysis?", "options": [{"letter": "A", "text": "A free-text comments field where each reviewer types their own reasoning for a dismissal in whatever wording feels most natural to them", "correct": false, "explanation": "Unstructured free text cannot be reliably aggregated or queried at scale, making systematic trend analysis across many findings impractical."}, {"letter": "B", "text": "A single overall confidence score per finding, with no record of which construct or rule actually produced that finding", "correct": false, "explanation": "A confidence score without a reference to the triggering construct does not let analysts isolate which specific pattern is responsible for the dismissals."}, {"letter": "C", "text": "A severity field that raises the priority of every incoming finding regardless of which construct triggered it originally", "correct": false, "explanation": "Raising severity indiscriminately does not help identify which construct is generating false positives and would increase noise rather than support the requested analysis."}, {"letter": "D", "text": "A detected_pattern field naming the specific construct that triggered the finding, so dismissals can be grouped by construct", "correct": true, "explanation": "Recording the specific triggering construct in a detected_pattern field lets the team aggregate dismissals by pattern, revealing which constructs generate disproportionate false positives so the detection logic can be tuned."}], "correct": "D", "select": 1, "group": "I"}, {"id": "f4-082", "domain": 4, "task_id": "4.1", "objective": "Design prompts with explicit criteria to improve precision and reduce false positives", "situation": "A team already tried adding \"only report issues you are confident about\" to a noisy category and saw no improvement in precision. An architect now wants to redesign the category's scope entirely rather than continuing to tune confidence language.", "question": "Which redesign reflects the correct lesson from the earlier failed attempt?", "options": [{"letter": "A", "text": "Rephrase the same confidence instruction using stronger emphasis, such as capitalizing key words, so the model treats the requirement as a stricter constraint.", "correct": false, "explanation": "Adding emphasis to the same confidence-based instruction does not change its fundamental nature; it still fails to define what makes an issue reportable, so the same failure mode would likely recur."}, {"letter": "B", "text": "Replace the confidence instruction with a list of the specific issue types that qualify for this category, and explicitly state which related issue types should be skipped.", "correct": true, "explanation": "This replaces the confidence-based filter with categorical criteria defining exactly which issue types are in scope and which are excluded, addressing the root cause of the earlier failure rather than adjusting the phrasing of the same ineffective approach."}, {"letter": "C", "text": "Keep the confidence instruction but add a numeric percentage threshold to it, so the model has a specific number to compare its confidence against.", "correct": false, "explanation": "A numeric threshold still asks the model to introspect on its own confidence rather than apply a defined categorical rule, and confidence self-reports of this kind have already been shown not to move precision in this scenario."}, {"letter": "D", "text": "Move the confidence instruction from the system prompt into the user message instead, on the assumption that message placement was the reason it had no effect.", "correct": false, "explanation": "Message placement is not the reason a vague confidence instruction fails; the lesson from the earlier attempt is that the content of the instruction, not its location, needs to change."}], "correct": "B", "select": 1, "group": "I"}, {"id": "f4-083", "domain": 4, "task_id": "4.4", "objective": "Implement validation, retry, and feedback loops for extraction quality", "situation": "An invoice-extraction pipeline returns structured JSON that is missing the required invoice_number field, even though the number is clearly printed on the source PDF. The team wants to retry the extraction with targeted feedback so the model can correct the omission.", "question": "Which retry design is most likely to succeed?", "options": [{"letter": "A", "text": "Rewrite the system prompt with new wording and resend it with the PDF, without naming the missing field or the earlier attempt", "correct": false, "explanation": "Changing the system prompt wording without explicitly identifying the missing invoice_number field or referencing the earlier failed attempt provides only vague guidance. The model is less likely to correct the omission if the specific validation failure is not communicated."}, {"letter": "B", "text": "Resend the original PDF plus the prior failed JSON, and state that invoice_number is required but was omitted from that attempt", "correct": true, "explanation": "Research-supported validation feedback loops recommend providing the original document plus the identified extraction errors as context when asking the model to re-extract problematic fields. Including the prior failed JSON and explicitly naming invoice_number gives the model both the source data and the specific omission to fix, making this the strongest retry design among the options. For a more prevent-first approach, Anthropic's Structured Outputs feature can enforce required JSON fields from the start."}, {"letter": "C", "text": "Discard the whole conversation and resend the unchanged original prompt, hoping sampling variance yields a different result this time", "correct": false, "explanation": "Resending the same original prompt without any feedback does not address the omission, and relying on sampling variance is not a reliable correction strategy. The model may repeat the same error, especially since it received no new signal that invoice_number is required."}, {"letter": "D", "text": "Send only the validation error text by itself, without re-attaching the source PDF or the earlier failed output from the first pass", "correct": false, "explanation": "Without the source PDF, the model has no document context from which to extract the missing invoice_number. Omitting the earlier failed JSON also removes useful context about what the first pass returned, making it harder for the model to identify and correct the specific omission."}], "correct": "B", "select": 1, "group": "I"}, {"id": "f4-084", "domain": 4, "task_id": "4.1", "objective": "Design prompts with explicit criteria to improve precision and reduce false positives", "situation": "The \"performance suggestions\" category in a code-review tool has a 70% false positive rate, and developers have stopped reading any findings from the tool at all, including from well-performing categories. The team needs to restore trust quickly while a better prompt for that category is developed over the following weeks.", "question": "What is the recommended immediate action?", "options": [{"letter": "A", "text": "Add a disclaimer banner on each code-review result indicating that certain categories have lower precision, so that developers can apply their own filtering criteria when reviewing findings across the project.", "correct": false, "explanation": "A disclaimer does not address the volume of false positives and adds friction without fixing the actual precision problem, failing to restore trust quickly."}, {"letter": "B", "text": "Merge the performance-suggestions category into the correctness category so that developers reviewing correctness findings also encounter performance suggestions, gradually rebuilding trust through repeated exposure.", "correct": false, "explanation": "Merging the noisy category into an accurate one would spread false positives into correctness findings, reducing trust in a previously reliable category. This approach does not rebuild trust but instead contaminates accurate results."}, {"letter": "C", "text": "Lower the overall severity label on every performance finding to \"informational\" so that developers see them as advisory notes and can continue reviewing other accurate categories while the prompt is iterated on.", "correct": false, "explanation": "Relabeling severity does not reduce the false positive rate; developers still encounter a large volume of incorrect findings, just with a different tag, so trust in the tool overall continues to erode."}, {"letter": "D", "text": "Temporarily disable the performance-suggestions category so developers see only the accurate categories, while iterating on that category's prompt separately before re-enabling it.", "correct": true, "explanation": "Temporarily disabling the high false-positive category is the recommended immediate action to prevent its noise from undermining trust in the accurate categories, while the prompt for that category is improved independently before it is re-enabled."}], "correct": "D", "select": 1, "group": "I"}, {"id": "f4-085", "domain": 4, "task_id": "4.5", "objective": "Design efficient batch processing strategies", "situation": "A search-augmentation team wants each batched research request to use a server-side web search tool so Claude can look up current information and incorporate it into the same response, without the application fetching pages or feeding results back itself.", "question": "Is this workable within the Message Batches API?", "options": [{"letter": "A", "text": "Yes, but only if the application also submits a matching synchronous request in parallel so the server tool has a live connection to execute against", "correct": false, "explanation": "No parallel synchronous request is needed; the server tool executes within the batch request's own processing, independent of any other request."}, {"letter": "B", "text": "No, because the Message Batches API rejects any request that references a tool definition, whether the tool executes on the server or on the client", "correct": false, "explanation": "Tool use, including server tools, is explicitly supported in batch requests; only a narrow set of parameters such as streaming are disallowed."}, {"letter": "C", "text": "No, because server tools only function within streamed synchronous responses and streaming is one of the parameters batch requests do not support", "correct": false, "explanation": "Server tools do not require a streaming connection to function; they resolve within the asynchronous batch request just as they would in a non-streamed synchronous call."}, {"letter": "D", "text": "Yes, because server tools such as web search resolve automatically within the request itself, unlike client-side tools that need an application-supplied result", "correct": true, "explanation": "Server tools like web search execute as part of resolving the single batched request, so no application-side round trip is needed; this is distinct from client-side tools where the app must supply a tool_result to continue."}], "correct": "D", "select": 1, "group": "H"}, {"id": "f4-086", "domain": 4, "task_id": "4.2", "objective": "Apply few-shot prompting to improve output consistency and quality", "situation": "A research-summarization tool extracts source citations from academic papers. Some papers cite sources inline in the body text, e.g., '(Smith, 2019)', while others rely entirely on a numbered bibliography referenced by superscript markers. The tool consistently extracts citations correctly from inline-style papers but frequently returns an empty citation list for bibliography-style papers.", "question": "What should be added to the prompt to fix this?", "options": [{"letter": "A", "text": "An instruction noting that citations may appear either inline in the text or in a numbered bibliography", "correct": false, "explanation": "A brief instruction restating that two formats exist does not show the model how to actually perform the bibliography-style extraction; general awareness isn't the missing piece, the mechanics of tracing markers are."}, {"letter": "B", "text": "A preprocessing step that strips all superscript numbers before the document reaches the prompt", "correct": false, "explanation": "Stripping the superscript markers would remove the very signal the model needs to connect in-text references to bibliography entries, making the extraction failure worse rather than better."}, {"letter": "C", "text": "A rule that rejects any paper not using inline citation style, since that format is handled reliably", "correct": false, "explanation": "Rejecting an entire class of legitimate documents avoids the failure rather than fixing it, and it would make the tool unusable for a large share of academic papers."}, {"letter": "D", "text": "An example each for inline and bibliography style, showing how a superscript marker traces to its entry", "correct": true, "explanation": "Showing a worked example for each document structure, including how to connect a superscript marker to its corresponding bibliography entry, gives the model demonstrated precedent for the structure it currently fails on, which is exactly how few-shot examples address varied document structures."}], "correct": "D", "select": 1, "group": "I"}, {"id": "f4-087", "domain": 4, "task_id": "4.1", "objective": "Design prompts with explicit criteria to improve precision and reduce false positives", "situation": "An architect is redesigning review criteria for an internal Claude-based code review agent. The goal is to ensure that critical issues—such as security vulnerabilities and correctness bugs—are never missed, while avoiding noise from subjective preferences. The architect wants to define which issues should always be reported versus always skipped, rather than relying on the model's confidence to decide.", "question": "Which pair of instructions correctly demonstrates this approach?", "options": [{"letter": "A", "text": "Report a finding only after a second independent review pass confirms that the first pass's confidence score exceeds a fixed threshold; skip any finding where the passes disagree or the score is not duplicated.", "correct": false, "explanation": "This approach still depends on confidence scores and adds unnecessary complexity. It does not guarantee that critical issues are reported if the confidence threshold is not met or if passes disagree. The correct strategy uses fixed rules that are not based on confidence."}, {"letter": "B", "text": "Report every change that deviates from the team's officially documented style guide; skip any change the model classifies as a subjective personal taste preference, even if it affects readability.", "correct": false, "explanation": "Reporting every documented style-guide deviation treats a low-severity, cosmetic category the same as security and correctness bugs, so it still generates noise that competes with the critical findings the architect wants surfaced first. The skip half of the rule is not fixed either: it leaves the model to classify whatever the style guide does not cover as a 'subjective personal taste preference,' so an undocumented change can still be waved through or flagged based on the model's own read of it. That is the same reliance on model judgment the architect wants to eliminate, just moved from a confidence score onto a subjectivity label."}, {"letter": "C", "text": "Report any change that introduces a security vulnerability or a correctness bug that breaks existing behavior; skip formatting preferences and deviations from a file's established local conventions.", "correct": true, "explanation": "This matches Anthropic's recommended approach for identifying and reporting critical issues. Official documentation and responsible disclosure policies prioritize security vulnerabilities and correctness bugs that break existing behavior over stylistic concerns. By defining explicit categorical criteria, the architect eliminates reliance on subjective confidence scores, ensuring that critical issues are always reported while non-critical noise is skipped."}, {"letter": "D", "text": "Report any finding where the model's internal confidence score exceeds a fixed threshold of 0.9; skip any finding where the confidence score is below that threshold, treating all categories uniformly.", "correct": false, "explanation": "Relying on a fixed confidence threshold treats all issue categories uniformly, ignoring the varying severity of issues. Confidence scores are subjective and model-dependent; they are not a reliable basis for a deterministic reporting policy. The architect's goal is to catch critical issues regardless of confidence level."}], "correct": "C", "select": 1, "group": "I"}, {"id": "f4-088", "domain": 4, "task_id": "4.5", "objective": "Design efficient batch processing strategies", "situation": "An operations lead argues that because most Message Batches finish processing in under an hour, the team can safely promise customers a fixed 90-minute turnaround for a nightly summarization job. A colleague pushes back on this plan.", "question": "What is the strongest technical objection to the promise?", "options": [{"letter": "A", "text": "Summarization requests inherently take longer to process than other request types, so 90 minutes is too short a window regardless of typical batch completion times", "correct": false, "explanation": "Summarization is not documented as a request type with inherently longer processing times than other Messages API workloads within a batch."}, {"letter": "B", "text": "The Batches API caps total daily throughput per workspace, so a 90-minute promise would only be broken once the workspace exceeds its allotted request volume", "correct": false, "explanation": "The objection to a fixed-time promise is the lack of an SLA, not a workspace-level daily throughput cap being the mechanism that would break it."}, {"letter": "C", "text": "The Batches API carries no guaranteed latency SLA, so a batch can legitimately take up to 24 hours under heavy demand, making a fixed 90-minute promise unreliable", "correct": true, "explanation": "Typical fast completion is not a guarantee; the Batches API can take up to its full 24-hour window when demand is high, so committing to a fixed 90-minute SLA misrepresents the platform's actual guarantee."}, {"letter": "D", "text": "Fixed turnaround promises are incompatible with the Batches API because every batch must be manually retrieved through the console rather than through automated polling", "correct": false, "explanation": "Batch results can be retrieved programmatically through the retrieval endpoint and are not limited to manual console access, so this is not the relevant risk to the promise."}], "correct": "C", "select": 1, "group": "H"}, {"id": "f4-089", "domain": 4, "task_id": "4.3", "objective": "Enforce structured output using tool use and JSON schemas", "situation": "A contract-review tool extracts a governing_law field indicating which jurisdiction's law applies to a contract. In practice, some contracts state this explicitly, some imply it ambiguously across two jurisdictions, and some never mention it at all. The current enum only lists specific jurisdiction names, forcing the model to guess in ambiguous cases.", "question": "How should the schema be revised?", "options": [{"letter": "A", "text": "Change governing_law to a boolean field indicating only whether any jurisdiction is mentioned in the contract text", "correct": false, "explanation": "Collapsing the field to a boolean throws away the actual jurisdiction information for the majority of contracts where it is clearly stated, losing valuable structured data."}, {"letter": "B", "text": "Add an \"unclear\" enum value to distinguish genuinely ambiguous or unstated cases from confidently identified jurisdictions", "correct": true, "explanation": "Adding an explicit \"unclear\" enum value gives the model a truthful, structured way to represent ambiguous or unstated cases instead of being forced into a specific jurisdiction it can't confidently determine."}, {"letter": "C", "text": "Set tool_choice to {\"type\": \"any\"} so the model is forced to pick a jurisdiction value from the enum on every single contract", "correct": false, "explanation": "tool_choice controls whether a tool is invoked at all, not the permissible values within an already-defined field's enum, so it doesn't address the ambiguous-value problem."}, {"letter": "D", "text": "Duplicate every jurisdiction name in the enum with an \"ambiguous_\" prefix so each jurisdiction has an ambiguous counterpart value", "correct": false, "explanation": "Doubling every enum value with an ambiguous variant bloats the schema unnecessarily and still doesn't cover contracts where no jurisdiction is mentioned at all."}], "correct": "B", "select": 1, "group": "H"}, {"id": "f4-090", "domain": 4, "task_id": "4.3", "objective": "Enforce structured output using tool use and JSON schemas", "situation": "A metadata-tagging pipeline runs on Claude Sonnet 4.6 and registers a single tag_document tool. Two pipeline requirements are non-negotiable: every input document must result in a tag_document call, and Claude must remain able to reason before tagging ambiguous documents. The developer sets tool_choice to {\"type\": \"tool\", \"name\": \"tag_document\"} and thinking to {\"type\": \"enabled\", \"budget_tokens\": 10000}, and the API rejects the request with an error saying that thinking may not be enabled when tool_choice forces tool use.", "question": "Which change satisfies both requirements?", "options": [{"letter": "A", "text": "Keep manual extended thinking and change tool_choice to {\"type\": \"any\"}, since any lets Claude choose which registered tool to call instead of naming one", "correct": false, "explanation": "{\"type\": \"any\"} is also a forced tool selection: the documentation names both {\"type\": \"any\"} and {\"type\": \"tool\", \"name\": \"...\"} as the values that fail with manual extended thinking. The request would be rejected with the same error, so neither requirement is met."}, {"letter": "B", "text": "Keep manual extended thinking and switch tool_choice to {\"type\": \"auto\"}, then instruct Claude in the system prompt to call tag_document for every document", "correct": false, "explanation": "Setting tool_choice to {\"type\": \"auto\"} does clear the error, because auto is one of the two values manual extended thinking accepts, but auto leaves the decision to Claude: it may answer directly instead of calling the tool. A system-prompt instruction is guidance, not an API guarantee, so the requirement that every document produce a tag_document call is no longer enforced."}, {"letter": "C", "text": "Keep tool_choice forced to tag_document and remove the thinking parameter from the request, so documents are tagged with thinking turned off", "correct": false, "explanation": "Removing the thinking parameter clears the error and preserves guaranteed tagging, but it turns thinking off entirely, so Claude can no longer reason about an ambiguous document before tagging it. It is also unnecessary on a model that offers adaptive thinking, which is compatible with forced tool use."}, {"letter": "D", "text": "Switch the request to thinking: {\"type\": \"adaptive\"} and keep tool_choice forced to tag_document, since adaptive thinking supports forced tool use", "correct": true, "explanation": "The thinking documentation states that forced tool use is incompatible with manual extended thinking (thinking: {\"type\": \"enabled\"}) but works with adaptive thinking. Claude Sonnet 4.6 supports thinking: {\"type\": \"adaptive\"}, so switching modes keeps the forced tool_choice that guarantees tagging while leaving thinking available for the documents that need it."}], "correct": "D", "select": 1, "group": "H"}, {"id": "f4-091", "domain": 4, "task_id": "4.1", "objective": "Design prompts with explicit criteria to improve precision and reduce false positives", "situation": "A code review agent flags a helper function because its naming does not match the dominant naming convention in the file. However, the file already contains legacy functions with several naming styles, and the helper function's name is consistent with one of those legacy styles but not with the project's canonical naming standard. The architect is using a subagent-based code review workflow and wants a criterion that reduces this kind of false positive without suppressing genuine naming defects.", "question": "Which criterion best addresses this failure mode?", "options": [{"letter": "A", "text": "Skip all naming-related findings across the entire codebase regardless of context, treating naming style as advisory, and focus the review exclusively on validating the code's logic and error handling.", "correct": false, "explanation": "This is overly broad and would suppress genuine naming defects, which the architect explicitly wants to retain. Skipping all naming checks abandons a useful review dimension instead of addressing the specific false positive from pre-existing local style variations. A dedicated subagent with canonical-standard focus is a more targeted approach."}, {"letter": "B", "text": "Ask the model to only report naming issues when its confidence exceeds 90% based on comparing the usage against standard library conventions and the project's own style guide.", "correct": false, "explanation": "A confidence threshold cannot reliably distinguish between a genuine naming defect and a legacy-style mismatch when the file already contains multiple styles. It may suppress some false positives but can also suppress real bugs or still flag correct code that conflicts with an incomplete style guide. The recommended approach is to use a specialized subagent that focuses on the canonical standard, not an arbitrary confidence cutoff."}, {"letter": "C", "text": "Report the naming inconsistency but flag it as low severity in the findings list and include the local style variations that are already present in the file to provide context for the review.", "correct": false, "explanation": "This approach still reports the inconsistency as a finding, so it does not reduce the number of false positives—it only annotates them. The architect asked for a criterion that reduces this failure mode, not one that merely explains it away. Isolating the review in a dedicated subagent with a canonical-standard focus would better reduce false positives."}, {"letter": "D", "text": "Dedicate a subagent to naming review with an isolated context window, a custom system prompt that instructs it to ignore pre-existing local style variations and flag only deviations from the project's canonical naming standard, and least-privilege tool access. This reduces false positives by focusing the review on the canonical standard, though it cannot eliminate all false positives.", "correct": true, "explanation": "Anthropic's documentation recommends isolating focused tasks in separate subagents with their own context windows, custom system prompts, and least-privilege tool access. A naming-review subagent can be instructed to ignore legacy local style noise and flag only deviations from the canonical naming standard, which directly reduces the false positives caused by mixed local styles. Note that this reduces, but does not completely eliminate, false positives—no automated review can guarantee zero false positives."}], "correct": "D", "select": 1, "group": "I"}, {"id": "f4-092", "domain": 4, "task_id": "4.6", "objective": "Design multi-instance and multi-pass review architectures", "situation": "A generator instance is instructed: 'Before you finish, re-read your changes and point out any mistakes.' The team observes this rarely surfaces issues that a fresh reviewer later finds.", "question": "What best explains this?", "options": [{"letter": "A", "text": "The generator is still in the session where it already committed to its design decisions, so it is less likely to question choices it just justified", "correct": true, "explanation": "This is the self-review limitation: retained reasoning context from generation makes the model less likely to question decisions it already made in the same session."}, {"letter": "B", "text": "The instruction is phrased as a command rather than a question, and rephrasing it as a question would make the model more critical of its own output", "correct": false, "explanation": "Phrasing the instruction as a question doesn't remove the retained reasoning context that causes the underlying limitation."}, {"letter": "C", "text": "The generator's extended thinking is disabled by default, so it never allocates any reasoning tokens to the re-read step at all", "correct": false, "explanation": "Extended thinking availability is unrelated to the root cause; even with extended thinking enabled, the same-session bias remains."}, {"letter": "D", "text": "The generator lacks access to the files it just wrote, so it cannot literally re-read the changes it is being asked to critique", "correct": false, "explanation": "The generator typically still has file access; the limitation is about retained reasoning bias, not missing tool access."}], "correct": "A", "select": 1, "group": "F"}, {"id": "f4-093", "domain": 4, "task_id": "4.5", "objective": "Design efficient batch processing strategies", "situation": "A batch processing pipeline must deliver classification results within 36 hours of data ingestion. The pipeline submits jobs to the Anthropic Batches API, which processes each batch in up to 24 hours (worst case). After the API completes, a mandatory 2-hour formatting step runs before the report is finalized.", "question": "To guarantee the 36-hour deadline even under worst-case API timing, what is the longest allowed interval between consecutive batch submission starts?", "options": [{"letter": "A", "text": "At least every 14 hours, since the downstream formatting time can be absorbed by shortening the batch's own processing window instead of the cycle interval.", "correct": false, "explanation": "The Batches API processing time is a guaranteed upper bound (24 hours), not a variable that can be arbitrarily reduced. There is no mechanism to shorten this worst-case time; the formatting step adds a fixed extra 2 hours that must be accommodated by the interval."}, {"letter": "B", "text": "At least every 12 hours, since ignoring the formatting step and matching half the deadline directly against the processing window is sufficient.", "correct": false, "explanation": "This option neglects the 2-hour formatting step. The actual worst-case pipeline time includes that step, so the total would be interval + 24 h + 2 h. With a 12-hour interval, that sum becomes 38 hours, exceeding the 36-hour deadline."}, {"letter": "C", "text": "At least every 10 hours, since the worst-case wait until the next batch (the interval) plus the 24-hour processing and the 2-hour formatting must not exceed 36 hours.", "correct": true, "explanation": "The total pipeline duration for any single data point consists of: (1) the time from ingestion until the next batch submission (which can be as long as the interval between submissions), (2) the Batches API processing time (up to 24 hours, as per the API’s SLA), and (3) the 2-hour formatting step. To guarantee delivery within 36 hours, the worst-case sum must be ≤ 36 h: interval + 24 h + 2 h ≤ 36 h, so interval ≤ 10 h. Therefore, batches must be submitted at least every 10 hours."}, {"letter": "D", "text": "At least every 2 hours, since the formatting step's fixed duration alone should dictate the entire submission cadence regardless of processing time.", "correct": false, "explanation": "While submitting every 2 hours would satisfy the 36‑hour deadline, the question asks for the maximum allowed interval. This option ignores the true processing time and sets an unnecessarily conservative frequency; the correct interval, derived from the full timing, is 10 hours."}], "correct": "C", "select": 1, "group": "H"}, {"id": "f4-094", "domain": 4, "task_id": "4.4", "objective": "Implement validation, retry, and feedback loops for extraction quality", "situation": "A lease-extraction tool processes a residential lease that states a monthly rent of $1,800 in the summary clause on page 1 but $1,850 in the payment schedule on page 4. Both values extract cleanly and both are individually valid numbers.", "question": "What should the extraction output do with this discrepancy?", "options": [{"letter": "A", "text": "Extract only the payment-schedule figure from page four, since it sits within a more detailed section of the lease", "correct": false, "explanation": "Preferring one section over the other without flagging the conflict hides the inconsistency instead of surfacing it, and there is no general rule that later sections are always authoritative."}, {"letter": "B", "text": "Average the two conflicting values together and return that single mean figure as the extracted monthly rent amount", "correct": false, "explanation": "Averaging two conflicting contractual figures manufactures a third number that appears nowhere in the source document and is not a legally meaningful resolution."}, {"letter": "C", "text": "Extract only the first rent figure found on page one of the lease, and quietly discard the later conflicting value found on page four", "correct": false, "explanation": "Silently keeping the first occurrence discards the conflicting information and risks reporting a rent figure that the lease itself does not unambiguously support."}, {"letter": "D", "text": "Extract both rent values, set a conflict_detected boolean to true, and route the record for human reconciliation rather than guessing", "correct": true, "explanation": "When a source document contains genuinely inconsistent data, the extraction should surface both values and a conflict_detected flag rather than guessing, since only a human (or downstream authority) can resolve which figure is authoritative."}], "correct": "D", "select": 1, "group": "I"}, {"id": "g13", "domain": 5, "scenario": "Multi-agent Research System", "situation": "Production monitoring shows inconsistent synthesis quality. When aggregated results are ~75K tokens, the synthesis agent reliably cites information from the first 15K tokens (web-search headlines/snippets) and the last 10K tokens (document analysis conclusions), but often misses critical findings in the middle 50K tokens—even when they directly answer the research question. How should you restructure the aggregated input?", "question": "How should you restructure the aggregated input?", "options": [{"letter": "A", "text": "Summarize all subagent outputs to under 20K tokens before aggregation to keep content within the model’s reliable processing range.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Stream subagent results to the synthesis agent incrementally, processing web-search results first to completion, then adding document analysis results.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Place a key-findings summary at the start of the aggregated input and organize detailed results with explicit section headings for easier navigation.", "correct": true, "explanation": "Putting a key-findings summary at the start leverages primacy effects so critical information sits in the most reliably processed position. Adding explicit section headings throughout helps the model navigate and attend to mid-input content, directly mitigating the “lost in the middle” phenomenon."}, {"letter": "D", "text": "Implement rotation that alternates which subagent’s results appear first across research tasks to ensure both sources get equal top positioning over time.", "correct": false, "explanation": ""}], "correct": "C", "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "group": "G"}, {"id": "g14", "domain": 5, "scenario": "Multi-agent Research System", "situation": "In testing, the combined output of the web-search agent (85K tokens including page content) and the document analysis agent (70K tokens including chains of thought) totals 155K tokens, but the synthesis agent performs best with inputs under 50K tokens. Which solution is most effective?", "question": "Which solution is most effective?", "options": [{"letter": "A", "text": "Modify upstream agents to return structured data (key facts, quotes, relevance scores) instead of verbose content and reasoning.", "correct": true, "explanation": "Modifying upstream agents to return structured data fixes the root cause by reducing token volume at the source while preserving essential information. It avoids passing bulky page content and reasoning traces that inflate tokens without improving the synthesis step."}, {"letter": "B", "text": "Add an intermediate summarization agent that condenses findings before passing them to synthesis.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Have the synthesis agent process findings in sequential batches, maintaining state between calls.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Store findings in a vector database and give the synthesis agent search tools to query during its work.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "group": "B"}, {"id": "g45", "domain": 5, "scenario": "Code Generation with Claude Code", "situation": "You’re adding error-handling wrappers around external API calls across a 120-file codebase. The work has three phases: (1) discover all call sites and patterns, (2) collaboratively design the error-handling approach, and (3) implement wrappers consistently. In Phase 1, Claude generates large output listing hundreds of call sites with context, quickly filling the context window before discovery finishes.", "question": "Which approach is most effective to complete the task while maintaining implementation consistency?", "options": [{"letter": "A", "text": "Use an Explore subagent for Phase 1 to isolate verbose discovery output and return a summary, then continue Phases 2–3 in the main conversation.", "correct": true, "explanation": "An Explore subagent isolates the verbose discovery output in a separate context and returns only a concise summary to the main conversation. This preserves the main context window for the collaborative design and consistent implementation phases where retained context is most valuable."}, {"letter": "B", "text": "Do all phases in the main conversation, periodically using `/compact` to reduce context usage while moving through files.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Switch to headless mode with `--continue`, passing explicit context summaries between batch calls to maintain continuity.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Define the error-handling pattern in CLAUDE.md, then process files in batches across multiple sessions relying on the shared memory file for consistency.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "group": "G"}, {"id": "g54", "domain": 5, "scenario": "Customer Support Agent", "situation": "Production logs show a pattern: customers reference specific amounts (e.g., “the 15% discount I mentioned”), but the agent responds with incorrect values. Investigation shows these details were mentioned 20+ turns ago and condensed into vague summaries like “promotional pricing was discussed.” What fix is most effective?", "question": "What fix is most effective?", "options": [{"letter": "A", "text": "Increase the summarization threshold from 70% to 85% so conversations have more room before summarization triggers.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Store full conversation history in external storage and implement retrieval when the agent detects references like “as I mentioned.”", "correct": false, "explanation": ""}, {"letter": "C", "text": "Extract transactional facts (amounts, dates, order numbers) into a persistent “case facts” block included in every prompt outside the summarized history.", "correct": true, "explanation": "Summarization inherently loses precise details. Extracting transactional facts into a structured “case facts” block outside the summarized history preserves critical information so it’s reliably available in every prompt regardless of how many turns have been summarized."}, {"letter": "D", "text": "Revise the summarization prompt to explicitly preserve all numbers, percentages, dates, and customer-stated expectations verbatim.", "correct": false, "explanation": ""}], "correct": "C", "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "group": "G"}, {"id": "g63", "domain": 5, "scenario": "Conversational AI Architecture Patterns", "situation": "Over several turns discussing investment strategy, a user stated \"I have a very low risk tolerance\" and later \"I want to maximize my returns.\" They now ask: \"What should I invest in?\"", "question": "Which approach best ensures the recommendation aligns with the user's actual priority?", "options": [{"letter": "A", "text": "Surface the contradiction and ask the user to clarify which matters more.", "correct": true, "explanation": "When user preferences directly contradict each other, surfacing the conflict and asking for clarification is the only way to guarantee the recommendation aligns with the user's true intent. Any other approach involves making an assumption that may be wrong—maximizing returns and low risk tolerance are fundamentally incompatible goals that require a human decision."}, {"letter": "B", "text": "Provide separate recommendations for both scenarios.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Proceed with the most recently stated preference.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Recommend a balanced portfolio without addressing the conflict.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "group": "G"}, {"id": "g64", "domain": 5, "scenario": "Conversational AI Architecture Patterns", "situation": "Users refine playlist preferences over multiple conversation turns. Two messages after a user said \"I love jazz,\" Claude asks \"What genres do you enjoy?\"", "question": "What is the most likely cause?", "options": [{"letter": "A", "text": "Claude requires a vector database connection to maintain conversation memory.", "correct": false, "explanation": ""}, {"letter": "B", "text": "The model's context window has been exceeded.", "correct": false, "explanation": ""}, {"letter": "C", "text": "The Claude API requires a `session_id` parameter.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Your application isn't including prior messages in the `messages` array.", "correct": true, "explanation": "Claude has no server-side memory—every API call is stateless. Without including the full conversation history in the `messages` array of each request, Claude has no knowledge of prior turns. Vector databases (A) and `session_id` (C) are not part of Claude's architecture; context window overflow (B) is impossible for two-message exchanges."}], "correct": "D", "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "group": "G"}, {"id": "g65", "domain": 5, "scenario": "Conversational AI Architecture Patterns", "situation": "After a 40-minute cooking session, the conversation reaches 78,000 tokens. History includes allergies, recipe scaling, clarified cooking terms, and general discussion. You must reduce tokens while preserving important information.", "question": "What approach best balances preservation with token reduction?", "options": [{"letter": "A", "text": "Summarize the entire conversation history.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Keep only the most recent 20,000 tokens.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Extract critical structured data (allergies, quantities, preferences), summarize general discussion, and keep recent exchanges verbatim.", "correct": true, "explanation": "The hybrid approach preserves the highest-value information at the lowest cost. Critical facts like allergies and recipe quantities are extracted into a compact structured block (preventing the precision loss that occurs during summarization), general discussion is summarized, and recent exchanges are kept verbatim for conversational coherence. Options A and B risk losing critical dietary information; D is architectural overkill for a single cooking session."}, {"letter": "D", "text": "Store the full conversation externally and retrieve relevant parts via semantic search.", "correct": false, "explanation": ""}], "correct": "C", "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "group": "G"}, {"id": "g66", "domain": 5, "scenario": "Conversational AI Architecture Patterns", "situation": "Users report that during extended conversations the assistant loses track of earlier topics and preferences. Your current implementation keeps only the last 25 message pairs.", "question": "What is the most effective solution?", "options": [{"letter": "A", "text": "Hybrid approach: summarize older messages while keeping recent ones verbatim.", "correct": true, "explanation": "The hybrid approach addresses both dimensions of the problem: retaining exact recent context (critical for conversational coherence) while maintaining a compressed representation of earlier preferences (preventing total loss when pairs are dropped). Increasing the window (C) simply delays the same problem. Vector search (B) may miss important context that isn't semantically similar to the current query. Full per-turn summarization (D) adds overhead and accumulates summarization errors."}, {"letter": "B", "text": "Vector similarity search over the full conversation history.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Increase the window to 50 message pairs.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Summarize dropped messages every turn and prepend the running summary.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "group": "G"}, {"id": "g67", "domain": 5, "scenario": "Conversational AI Architecture Patterns", "situation": "Users report that latency increases and costs rise when conversations exceed 50 turns.", "question": "What is the primary cause?", "options": [{"letter": "A", "text": "The entire conversation history is included with each API request.", "correct": true, "explanation": "Claude's API is fully stateless—every request must include the complete conversation history in the `messages` array. As conversations grow, each request carries more tokens, which directly increases both processing latency and cost. The model does not maintain any internal state between calls (D is false), and response length is not inherently tied to conversation length (B)."}, {"letter": "B", "text": "The model generates progressively longer responses.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Database operations slow down as history grows.", "correct": false, "explanation": ""}, {"letter": "D", "text": "The model builds an internal user profile requiring more processing.", "correct": false, "explanation": ""}], "correct": "A", "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "group": "G"}, {"id": "g68", "domain": 5, "scenario": "Conversational AI Architecture Patterns", "situation": "After three months of weekly sessions, conversation history grows to 85,000 tokens. When a user asks \"What did we conclude about the theme of isolation?\", the assistant gives generic answers instead of referencing previous discussions.", "question": "What is the most effective approach?", "options": [{"letter": "A", "text": "Rolling window truncation.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Progressive summarization capturing key conclusions.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Semantic embeddings with retrieval of relevant exchanges.", "correct": true, "explanation": "Semantic search over conversation history is the only approach that scales to three months of discussion while being able to surface specific relevant exchanges on demand. Rolling window (A) would discard most of the history. Progressive summarization (B) compresses discussions into abstractions that lose the specific conclusions users are asking about. XML tags (D) require restructuring all past content and don't solve the retrieval problem at this scale."}, {"letter": "D", "text": "Add structured XML tags marking discussion conclusions.", "correct": false, "explanation": ""}], "correct": "C", "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "group": "G"}, {"id": "g70", "domain": 5, "scenario": "Conversational AI Architecture Patterns", "situation": "Your AI tutor has a 2,800-token system prompt defining teaching methodology and adaptation rules. After 12 turns, the assistant starts ignoring proficiency levels.", "question": "What is the most effective fix?", "options": [{"letter": "A", "text": "Inject reminders every 4–5 turns.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Replace verbose rules with few-shot examples demonstrating proficiency-level adaptation.", "correct": true, "explanation": "A 2,800-token system prompt with declarative rules is vulnerable to drift because abstract rules require the model to reason about them on every turn. Replacing verbose rules with concrete few-shot examples that demonstrate correct proficiency-level adaptation gives the model clear behavioral patterns to match—this is more reliably followed across many turns than abstract instructions. Reminder injection (A) helps but addresses symptoms; end-placement (C) helps initially but not with turn-level drift; regeneration (D) is expensive and corrective."}, {"letter": "C", "text": "Place critical rules at the end of the system prompt.", "correct": false, "explanation": ""}, {"letter": "D", "text": "Evaluate responses and regenerate if difficulty level mismatches.", "correct": false, "explanation": ""}], "correct": "B", "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "group": "G"}, {"id": "g75", "domain": 5, "scenario": "Conversational AI Architecture Patterns", "situation": "Your assistant uses a contractor-persona system prompt. Early turns follow the rules, but by turn 7 the assistant gives generic advice. Conversation length is only 2,500 tokens.", "question": "What is the most likely cause?", "options": [{"letter": "A", "text": "System prompts only establish initial behavior.", "correct": false, "explanation": ""}, {"letter": "B", "text": "Model attention weakens as turns accumulate.", "correct": false, "explanation": ""}, {"letter": "C", "text": "Accumulated assistant responses dilute system prompt influence.", "correct": true, "explanation": "As assistant responses accumulate in the conversation history, the proportion of text reflecting the system prompt's behavioral constraints decreases relative to the growing body of assistant-generated content. The model increasingly pattern-matches to its own prior outputs rather than the system prompt, compounding drift even at short token lengths. The system prompt is included in every API call (D is false as a standalone explanation), and model attention degradation (B) doesn't operate at 2,500 tokens."}, {"letter": "D", "text": "The system prompt is only sent once.", "correct": false, "explanation": ""}], "correct": "C", "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "group": "G"}, {"id": "m18", "domain": 5, "scenario": null, "situation": "During testing, you observe that in extended exploration sessions (30+ minutes), the agent starts giving inconsistent answers about code structure it discussed earlier. Engineers report having to repeat context about modules they've already explored.", "question": "What's the most effective approach to address this?", "options": [{"letter": "A", "text": "Have the agent maintain a scratchpad file that records key findings, referencing it for subsequent questions.", "correct": true, "explanation": "Correct. A scratchpad offloads findings to durable storage the agent can re-read on demand, giving it a stable 'memory' independent of how crowded the context window gets."}, {"letter": "B", "text": "Switch to a higher-capacity model tier to provide more context window space for accumulated exploration data.", "correct": false, "explanation": "A bigger window delays the problem but doesn't solve attention degradation over long, noisy contexts."}, {"letter": "C", "text": "Implement automatic context clearing every 15 minutes to ensure the agent starts with fresh, uncontaminated context.", "correct": false, "explanation": "Wholesale clearing throws away valid findings — exactly the state engineers complain about having to re-establish."}, {"letter": "D", "text": "Create summaries of all source files before exploration begins, loading only these compressed representations into context.", "correct": false, "explanation": "Pre-summaries lose the detail that makes code exploration useful and often misrepresent files the engineer cares about most."}], "correct": "A", "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "group": "G"}, {"id": "m20", "domain": 5, "scenario": null, "situation": "Your agent has spent 25 minutes exploring a game engine's rendering subsystem—reading shader code, buffer management, and frame synchronization logic. An engineer now asks it to understand how the physics engine integrates with rendering for collision debug overlays. You notice recent responses reference \"typical rendering patterns\" rather than the specific VulkanPipeline and FrameGraph classes it discovered earlier.", "question": "What's the most effective approach?", "options": [{"letter": "A", "text": "Spawn a sub-agent to explore physics independently, then manually synthesize its findings with the rendering knowledge accumulated in the main conversation.", "correct": false, "explanation": "Manual synthesis in the main conversation pushes even more into an already-degrading context."}, {"letter": "B", "text": "Continue in the current context with more targeted prompts referencing the specific classes by name.", "correct": false, "explanation": "The agent is already losing specificity ('typical patterns' instead of VulkanPipeline). Pushing harder on the same degraded context doesn't restore it."}, {"letter": "C", "text": "Summarize key rendering findings, then spawn a sub-agent for physics exploration with that summary in its initial context.", "correct": true, "explanation": "Correct. Condense what you've learned about rendering into a compact summary, then give a fresh subagent that summary plus the physics task — you preserve the important signal and escape the degraded context."}, {"letter": "D", "text": "Use /clear to reset context completely, then start fresh with physics exploration using file paths from the project's CLAUDE.md.", "correct": false, "explanation": "/clear nukes the rendering findings you actually need for the cross-cut question about collision debug overlays."}], "correct": "C", "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "group": "G"}, {"id": "m23", "domain": 5, "scenario": null, "situation": "An engineer asks the agent to understand how the caching layer works before adding a new cache invalidation trigger. After initial Grep searches, the agent has identified that caching logic spans 15 files including decorators, middleware, and service classes (~8,000 lines total).", "question": "What's the most effective next step for building understanding while managing context constraints?", "options": [{"letter": "A", "text": "Use the Read tool to sequentially load all 15 files, building complete understanding across the full caching implementation.", "correct": false, "explanation": "Loading 8,000 lines without structure floods the context and buries the invalidation logic in unrelated code."}, {"letter": "B", "text": "Analyze imports and class hierarchies to identify the base cache class, Read that file to understand the interface, then trace specific invalidation implementations.", "correct": true, "explanation": "Correct. Start from the architectural root (the interface), then navigate only the specific implementations that matter for invalidation — focused reading, low context cost."}, {"letter": "C", "text": "Use Grep to search for \"invalidate\" and \"expire\" patterns across all files, then Read only those specific line ranges with minimal surrounding context.", "correct": false, "explanation": "Keyword-only reading strips the class/method context that tells you what the invalidation actually guards."}, {"letter": "D", "text": "Use Glob to find files matching common caching patterns (cache.py, caching/), prioritize the largest files by reading them first, then check smaller files for gaps.", "correct": false, "explanation": "Size is a poor proxy for importance. You'd waste context on big utility files and miss the small strategic ones."}], "correct": "B", "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "group": "G"}, {"id": "m37", "domain": 5, "scenario": null, "situation": "The agent verifies customer identity through a multi-step process before resetting passwords. During testing, you notice that after the customer answers the third verification question, the agent asks them to provide their name again, as if the earlier exchange never happened.", "question": "What's the most likely cause of this behavior?", "options": [{"letter": "A", "text": "The verification tool is clearing the agent's internal state after each successful validation step.", "correct": false, "explanation": "Tools don't clear agent state — conversation state lives in the messages array you send."}, {"letter": "B", "text": "The prompt lacks instructions telling Claude to remember information across multiple exchanges.", "correct": false, "explanation": "Memory across turns isn't achieved by an instruction — it comes from passing the prior turns into the next request."}, {"letter": "C", "text": "The conversation history isn't being passed in subsequent API requests.", "correct": true, "explanation": "Correct. The API is stateless. Each request must include the full messages array. If you only send the latest turn, the model has no memory of earlier ones — exactly the 'ask for the name again' symptom."}, {"letter": "D", "text": "Claude's memory retention is limited to two conversational turns by default, requiring explicit configuration to extend it.", "correct": false, "explanation": "No such default limit exists. Context is bounded by the window, and history length is under your control."}], "correct": "C", "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "group": "G"}, {"id": "m42", "domain": 5, "scenario": null, "situation": "A customer raises three separate issues during one session: a refund inquiry (turns 1-15), a subscription question (turns 16-30), and a payment method update (turns 31-45). At turn 48, the customer asks \"What happened with my refund?\" The conversation is approaching context limits.", "question": "What strategy best maintains the agent's ability to address all issues throughout the session?", "options": [{"letter": "A", "text": "Extract and persist structured issue data (order IDs, amounts, statuses) into a separate context layer.", "correct": false, "explanation": "A sidecar context layer adds engineering and doesn't naturally preserve the conversational narrative the customer expects."}, {"letter": "B", "text": "Rely on MCP tools to re-fetch relevant information on demand when the customer references earlier issues.", "correct": false, "explanation": "Re-fetching is fine for freshness, but it doesn't solve the context-length problem — and you lose what was said between the customer and the agent."}, {"letter": "C", "text": "Summarize earlier turns into a narrative description, preserving full message history only for the active issue.", "correct": true, "explanation": "Correct. Progressive summarization compresses stable resolved topics while keeping the active thread verbatim — the classic pattern for long multi-issue conversations near the context limit."}, {"letter": "D", "text": "Implement sliding window context that retains the most recent 30 turns.", "correct": false, "explanation": "A pure sliding window silently drops the refund issue (turns 1–15) from context, which is exactly what the customer is asking about."}], "correct": "C", "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "group": "G"}, {"id": "m45", "domain": 5, "scenario": null, "situation": "Your agent has called `lookup_order` multiple times while investigating a customer's return requests. Each response includes 40+ fields (items, shipping details, payment info, status history). Tool outputs now represent the majority of the conversation's context. The customer mentions two more orders they want to discuss.", "question": "What's the most effective approach before making additional lookups?", "options": [{"letter": "A", "text": "Extract only return-relevant fields (items, purchase date, return window, status) from each existing order response, removing verbose details", "correct": true, "explanation": "Correct. Keep the fields that matter for the task and drop the rest. This directly addresses the context-bloat problem before you add two more lookups."}, {"letter": "B", "text": "Have the model generate a natural language summary of each order's key details, replacing structured responses with prose descriptions", "correct": false, "explanation": "Prose summaries lose precision for fields the model may need later (exact dates, amounts). Structured pruning is better than paraphrasing."}, {"letter": "C", "text": "Move all tool responses to a vector database with semantic indexing, retrieving relevant portions as the conversation continues", "correct": false, "explanation": "A vector DB is heavy infrastructure for what is essentially a pruning problem."}, {"letter": "D", "text": "Proceed with additional lookups without modifying the existing tool output context", "correct": false, "explanation": "Does nothing about the bloat — you're heading for context exhaustion."}], "correct": "A", "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "group": "G"}, {"id": "m60", "domain": 5, "scenario": null, "situation": "After your daily batch of 10,000 documents completes, 300 documents (3%) failed with \"`context_length_exceeded`\" errors. The results file identifies each failure by `custom_id`.", "question": "What's the most cost-effective approach to process these failures?", "options": [{"letter": "A", "text": "Reprocess the entire batch with prompt caching enabled to reduce the cost of retrying requests with identical system prompts", "correct": false, "explanation": "Re-running 10,000 successful documents to solve 300 failures is wasteful regardless of caching savings. And caching doesn't fix the length-exceeded error."}, {"letter": "B", "text": "Resubmit only the 300 failed documents after chunking them into smaller pieces, then combine the partial extractions", "correct": true, "explanation": "Correct. Targeted, and it addresses the actual cause (input too long). Chunk the oversized docs, extract per chunk, then merge — minimum tokens, fixes the specific failure mode."}, {"letter": "C", "text": "Resubmit the entire 10,000 document batch using a model tier with a larger context window", "correct": false, "explanation": "Expensive (reprocessing 9,700 successes) and may still fail on truly outsized documents."}, {"letter": "D", "text": "Increase the `max_tokens` parameter for the 300 failed documents and resubmit them in a new batch", "correct": false, "explanation": "`max_tokens` controls output length, not input length. It doesn't solve `context_length_exceeded`, which is about the combined input-plus-output budget."}], "correct": "B", "task_id": "4.3", "objective": "Enforce structured output using tool use and JSON schemas", "group": "H"}, {"id": "f5-001", "domain": 5, "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "situation": "A streaming service agent looks up a customer to resolve a billing question and finds two profiles sharing the same email domain and last name, consistent with a family plan. Usage data shows one profile is far more active than the other.", "question": "What should the agent do?", "options": [{"letter": "A", "text": "Merge billing details from both profiles to answer the question without confirming which one applies", "correct": false, "explanation": "- merging billing details from unconfirmed profiles risks disclosing or acting on the wrong customer's account information."}, {"letter": "B", "text": "Select the profile with the earlier account creation date, since it is likely the primary plan holder", "correct": false, "explanation": "- assuming the earlier-created account is the primary holder is a heuristic guess unrelated to actually verifying which profile applies."}, {"letter": "C", "text": "Select the more active profile, since higher usage suggests it belongs to the customer who initiated contact", "correct": false, "explanation": "- higher usage is a heuristic assumption that could easily point to the wrong family member's profile."}, {"letter": "D", "text": "Ask the customer for an additional identifier, such as the last four digits of the payment card on file", "correct": true, "explanation": "- requesting a distinguishing identifier resolves ambiguity directly, rather than guessing which profile the customer is calling about."}], "correct": "D", "select": 1, "group": "C"}, {"id": "f5-002", "domain": 5, "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "situation": "A synthesis agent is drafting the final section of a multi-source research report on a emerging medical treatment. Several subagents agree on the treatment's basic mechanism of action, but only one subagent found a single small study suggesting a long-term side effect, which no other source corroborates.", "question": "How should the report be structured to reflect this?", "options": [{"letter": "A", "text": "Present the side-effect finding with the same confidence language as the mechanism-of-action findings to keep the tone consistent", "correct": false, "explanation": "Presenting an uncorroborated single-source claim with the same confidence as well-established findings misrepresents its evidentiary strength to the reader."}, {"letter": "B", "text": "Use separate sections that distinguish well-corroborated findings from contested ones, preserving each source's characterization", "correct": true, "explanation": "Structuring the report with explicit sections for well-established versus contested findings preserves the original sources' characterizations and lets readers calibrate their confidence appropriately."}, {"letter": "C", "text": "Exclude the single-source side-effect finding from the report since it lacks corroboration from other subagents", "correct": false, "explanation": "Dropping the finding discards a potentially important signal simply because it is not yet corroborated, rather than presenting it with appropriate caveats."}, {"letter": "D", "text": "Merge all findings into one narrative section so the report reads smoothly without calling attention to which claims are more certain", "correct": false, "explanation": "Blending everything into one narrative erases the distinction between broadly corroborated claims and a single unconfirmed study, making it harder for the reader to judge reliability."}], "correct": "B", "select": 1, "group": "B"}, {"id": "f5-003", "domain": 5, "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "situation": "A cloud support agent attempts three different tool-based remediation steps to fix a customer's stuck deployment, and each attempt fails to resolve the issue. No further remediation options are available to the agent.", "question": "What should happen next?", "options": [{"letter": "A", "text": "Escalate the case to a human agent, since the agent is unable to make further meaningful progress", "correct": true, "explanation": "- inability to make meaningful progress after exhausting available remediation options is a clear escalation trigger."}, {"letter": "B", "text": "Close the case as unresolvable, without any handoff to a human agent for further troubleshooting", "correct": false, "explanation": "- closing the case without escalating abandons the customer instead of routing them to someone who can continue troubleshooting."}, {"letter": "C", "text": "Continue retrying the same remediation steps, since repeated attempts eventually resolve most stuck deployments", "correct": false, "explanation": "- repeating the same failed steps without new options is unlikely to produce a different outcome and delays needed escalation."}, {"letter": "D", "text": "Inform the customer the issue is resolved, since the remediation steps were the correct actions to attempt", "correct": false, "explanation": "- telling the customer the issue is resolved when the deployment is still stuck is factually inaccurate."}], "correct": "A", "select": 1, "group": "C"}, {"id": "f5-004", "domain": 5, "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "situation": "A customer emails a retailer's support agent: 'This is ridiculous, you sent me the wrong size AGAIN,' asking for an exchange for a plain t-shirt order. The exchange itself is a routine request that is fully covered by the store's standard 30-day exchange policy, and the customer has not asked to speak with a human.", "question": "How should the agent respond?", "options": [{"letter": "A", "text": "Ask the customer to confirm they are not requesting a human agent before proceeding with the exchange", "correct": false, "explanation": "The customer has already stated the request (an exchange) and has not asked for a human agent; nothing about what they want is ambiguous. Asking them to first confirm they don't want a human adds a needless step to a request the agent can already resolve directly."}, {"letter": "B", "text": "Acknowledge the customer's frustration and process the exchange now, since the request is within policy and resolvable", "correct": true, "explanation": "The exchange is fully covered by the standard 30-day policy and there is nothing about it the agent cannot resolve on its own — the customer has not asked for a human, and the request is otherwise routine and in-policy. Acknowledging the frustration first keeps the interaction from escalating further while the agent still resolves the case directly, which is what a support agent should optimize for when a request can be handled without a hand-off."}, {"letter": "C", "text": "Escalate to a human agent right away, since the customer's tone signals a case too sensitive to resolve directly", "correct": false, "explanation": "A frustrated tone about a repeated mistake does not, by itself, turn this into a case the agent cannot resolve: the order and the requested exchange are both routine and within policy, and the customer never asked for a human. Escalating a request the agent is fully equipped to complete works against resolving straightforward cases directly and reserving human hand-offs for cases that genuinely can't be completed without one."}, {"letter": "D", "text": "Process the exchange without commenting on the frustration, treating the emotional tone as irrelevant to the resolution", "correct": false, "explanation": "Processing the exchange while ignoring the stated frustration skips the acknowledgment that keeps a repeat-mistake interaction from feeling dismissive. The resolution itself is right, but pairing it with silence on how the customer feels is a worse response than pairing the same resolution with an acknowledgment."}], "correct": "B", "select": 1, "group": "C"}, {"id": "f5-005", "domain": 5, "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "situation": "A subagent researching climate data reports a sea-level rise figure from a study, but its structured output omits the study's publication year. When the coordinator merges this with a more recent study reporting a different figure, it cannot tell whether the discrepancy reflects a real disagreement or simply the passage of time between studies.", "question": "What is the most direct fix to the subagent's output contract?", "options": [{"letter": "A", "text": "Add a required field for the publication or data-collection date next to every extracted figure", "correct": true, "explanation": "This is correct. Adding publication/data-collection dates is a core metadata practice recommended by the FAIR principles to enable data discovery, context, and reproducibility. Anthropic itself includes knowledge cutoff dates in its model cards, demonstrating the importance of temporal information for interpreting figures. Without this metadata, merging studies from different times leads to ambiguity about whether differences are due to actual disagreement or simply the age of the data."}, {"letter": "B", "text": "Add a required field listing the study's funding source so the coordinator can assess potential bias", "correct": false, "explanation": "While funding source can be relevant for evaluating study credibility, it does not address the specific problem of a missing publication year for reconciling two different figures. The most direct fix is to include the date of the data or publication, which immediately clarifies whether a discrepancy is temporal or substantive."}, {"letter": "C", "text": "Add a required field capturing the total word count of the source study so the coordinator can judge its depth", "correct": false, "explanation": "Word count is unrelated to resolving temporal discrepancies between figures from different publication years. It does not help the coordinator understand whether a difference is due to the passage of time. The direct fix requires temporal metadata, not document length."}, {"letter": "D", "text": "Add a required field asking the subagent to guess how the figure might have changed since publication", "correct": false, "explanation": "Guessing introduces speculation rather than providing factual temporal context. The goal is reliable data merging, not conjecture. The correct approach is to include the objective publication date, not to estimate how data may have changed, which would compromise the integrity of the information."}], "correct": "A", "select": 1, "group": "B"}, {"id": "f5-006", "domain": 5, "task_id": "5.5", "objective": "Design human review workflows and confidence calibration", "situation": "A team is designing an ongoing quality-monitoring process for a document-extraction pipeline that has already passed initial validation and moved most high-confidence extractions out of human review. They want to detect new failure modes that didn't appear during initial testing, without re-reviewing every extraction.", "question": "Which sampling approach best fits this goal?", "options": [{"letter": "A", "text": "Review only the extractions that reviewers happen to flag as suspicious while performing unrelated ad hoc spot checks during slow periods, and use those flags to trigger targeted audits for new failure modes.", "correct": false, "explanation": "Relying on ad hoc reviews of extractions that reviewers flag as suspicious introduces bias toward visually obvious errors and offers no systematic coverage. This reactive method cannot provide a reliable estimate of overall error rates or detect subtle, novel failure modes."}, {"letter": "B", "text": "Sample extractions in proportion to how quickly each document type is processed, and use the resulting error counts to monitor for unexpected shifts in extraction accuracy across types.", "correct": false, "explanation": "Sampling in proportion to processing speed does not guarantee a representative sample of extractions, as faster‑processed document types may have different error profiles. Error counts alone, without corresponding totals, cannot yield meaningful error rate estimates to monitor for unexpected shifts."}, {"letter": "C", "text": "Draw a stratified random sample of high-confidence extractions across document types and fields on a recurring basis, and compare observed error rates against the validated baseline.", "correct": true, "explanation": "A recurring stratified random sample across document types and fields provides statistically meaningful, ongoing coverage of high-confidence extractions. This approach allows detection of shifts in error rates or new failure modes by comparing observed error rates against the validated baseline, without needing to review every extraction."}, {"letter": "D", "text": "Review every extraction produced during the first hour of each day across all document types and fields, and compare those error rates against the established baseline to detect emerging failure modes.", "correct": false, "explanation": "Sampling only the first hour of each day is a convenience sample that lacks statistical representativeness. It may miss failure modes that appear later in the day or affect specific document types not processed in that window, so it cannot reliably detect emerging error patterns."}], "correct": "C", "select": 1, "group": "C"}, {"id": "f5-007", "domain": 5, "task_id": "5.5", "objective": "Design human review workflows and confidence calibration", "situation": "A team is stratifying its ongoing sampling of high-confidence extractions across five document types that appear in very different volumes: one type makes up 70% of daily volume, and the other four each make up roughly 7.5%.", "question": "If the team samples strictly in proportion to volume, what risk does this introduce, and how should the sampling plan address it?", "options": [{"letter": "A", "text": "Pure volume-proportional sampling has no drawback here, since sampling proportional to volume always produces the statistically optimal allocation for detecting errors in every segment.", "correct": false, "explanation": "Proportional sampling is not universally optimal for detecting segment-specific errors; it systematically under-represents low-volume segments in absolute sample counts, which weakens the ability to detect error patterns specific to those segments."}, {"letter": "B", "text": "Pure volume-proportional sampling would over-sample the high-volume document type unnecessarily, so the team should exclude it from sampling entirely and focus only on the four smaller types.", "correct": false, "explanation": "Excluding the high-volume document type from sampling entirely removes visibility into the segment that constitutes the majority of daily output, which is not a sound tradeoff even if it's already well understood."}, {"letter": "C", "text": "Pure volume-proportional sampling is only a concern if the four low-volume document types are processed by a different prompt template than the high-volume type.", "correct": false, "explanation": "The under-sampling risk from proportional allocation exists regardless of whether the document types share a prompt template; it's a function of relative volume and sample size, not of prompt architecture."}, {"letter": "D", "text": "Pure volume-proportional sampling would under-sample the four low-volume document types, so the plan should also ensure a minimum sample size per document type regardless of its share of volume.", "correct": true, "explanation": "If samples are drawn strictly proportional to volume, the four smaller document types would each receive a small absolute number of samples, making it hard to detect emerging errors in those segments; ensuring a minimum sample size per document type protects visibility into every segment, not just the dominant one."}], "correct": "D", "select": 1, "group": "C"}, {"id": "f5-008", "domain": 5, "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "situation": "A policyholder has two active insurance claims open in one chat: a water-damage claim (claim #C-1092, estimated $6,400, status: under review) and an auto-glass claim (claim #C-1108, estimated $310, status: approved). The assistant's single blended summary later refers to \"the claim\" and \"the amount\" when answering a follow-up, creating ambiguity about which claim is meant.", "question": "What is the best fix?", "options": [{"letter": "A", "text": "Merge the two claims into a single combined claim number so only one amount and status need to be tracked", "correct": false, "explanation": "Combining unrelated claims into one record loses critical information, undermines data integrity, and prevents accurate reporting. In claims management, each incident should remain a distinct entity to ensure proper processing and auditing."}, {"letter": "B", "text": "Track each claim as its own structured record with claim ID, amount, and status, kept apart from any blended narrative", "correct": true, "explanation": "Maintaining separate, structured records for each claim follows insurance best practices for data management and aligns with Claude's stateless design — the model remembers nothing between calls, so your application must keep explicit context. Structured outputs (e.g., JSON with claim ID, amount, status) prevent ambiguity by clearly identifying each claim in every API call, as recommended in Anthropic's documentation for reliable interactions."}, {"letter": "C", "text": "Summarize only whichever claim was mentioned most recently and assume the other claim is no longer active", "correct": false, "explanation": "Assuming a claim is inactive without confirmation can lead to discarded valid claims and poor customer experience. The assistant must track all active claims explicitly; relying on recency is not a reliable method for ambiguity resolution."}, {"letter": "D", "text": "Require the policyholder to specify which claim number they mean every time they ask a follow-up question", "correct": false, "explanation": "While using claim IDs can disambiguate, forcing the user to always specify is cumbersome and not user‑friendly. A better approach is to store structured claim data in the application and pass only the relevant claim IDs to the model, letting the assistant resolve references automatically."}], "correct": "B", "select": 1, "group": "G"}, {"id": "f5-009", "domain": 5, "task_id": "5.5", "objective": "Design human review workflows and confidence calibration", "situation": "A financial-statement extraction pipeline routes any field scored below 0.85 confidence to human review. After several months, reviewers report that many fields scored at 0.9 or higher are still wrong, and they've begun manually re-checking most extractions regardless of the score, defeating the purpose of the threshold.", "question": "What is the most likely root cause?", "options": [{"letter": "A", "text": "The confidence scores are being computed correctly, but reviewers lack training on how to interpret a 0.9 versus a 0.95 score.", "correct": false, "explanation": "Even if reviewers misinterpret the scores, the fundamental problem remains: the extracted data is incorrect despite high confidence. Adequate training would not fix inaccurate extractions."}, {"letter": "B", "text": "No validation shows that fields scoring above 0.85 are actually correct at an acceptable rate, so a high score does not reliably predict a correct extraction.", "correct": true, "explanation": "Of the four explanations, this is the only one that accounts for the actual symptom described: fields scored 0.9 or higher are still coming out wrong. Raising the threshold, training reviewers to interpret the score more carefully, or switching to a fixed sampling percentage would not change whether a high score reliably means a correct extraction -- none of them touches the relationship between the score and the ground truth. A confidence score only predicts accuracy once it has been checked against labeled, human-verified extractions: Amazon SageMaker Ground Truth's automated-labeling workflow, for example, documents finding its confidence threshold by testing model predictions against a held-out human-annotated validation set so the expected accuracy above that threshold meets a required level, precisely because an untested threshold carries no such guarantee. With fields scored 0.9 or higher still wrong, the most likely explanation is that this pipeline's 0.85 cutoff was never checked against a labeled set in that way, so a 0.9 score cannot be trusted to mean a correct extraction."}, {"letter": "C", "text": "The document set has grown too large for a fixed threshold to remain meaningful, so the team should switch to reviewing a fixed percentage of documents instead.", "correct": false, "explanation": "Growth in document volume does not explain why highly confident fields are erroneous. Switching to a percentage-based review still relies on the same uncalibrated confidence scores and would not resolve the underlying inaccuracy."}, {"letter": "D", "text": "Reviewers are spending too much time on each individual field, so the threshold should be raised to reduce the total number of fields sent to review.", "correct": false, "explanation": "Raising the threshold would only increase the volume of fields sent for human review; it does nothing to address why high-confidence fields are already wrong. The core issue is the reliability of the confidence scores, not the review workload."}], "correct": "B", "select": 1, "group": "C"}, {"id": "f5-010", "domain": 5, "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "situation": "During a competitive-intelligence synthesis, two subagents report different figures for a competitor's annual revenue: one cites a press release stating $420 million, the other cites an analyst report stating $460 million. Both sources are credible.", "question": "How should the synthesis agent handle this discrepancy in the final report?", "options": [{"letter": "A", "text": "Report the analyst report's figure since analyst estimates are generally considered more rigorous than press releases", "correct": false, "explanation": "Silently selecting the analyst figure discards the press release value and hides the disagreement from the reader, who has no way to know a conflict existed."}, {"letter": "B", "text": "Average the two figures into a single blended estimate so the report presents one consistent number", "correct": false, "explanation": "Averaging invents a number that neither source actually reported and obscures the fact that a real disagreement exists between two credible sources."}, {"letter": "C", "text": "Report the press release's figure since it comes directly from the company rather than a third party", "correct": false, "explanation": "Silently selecting the press release figure discards the analyst value and hides the disagreement from the reader, who has no way to know a conflict existed."}, {"letter": "D", "text": "Present both figures side by side, attributed to their sources, and note that the values conflict", "correct": true, "explanation": "When credible sources conflict, the synthesis should annotate the conflict with source attribution for each value rather than arbitrarily choosing one, preserving the reader's ability to judge which figure applies to their use case."}], "correct": "D", "select": 1, "group": "B"}, {"id": "f5-011", "domain": 5, "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "situation": "An SRE assistant aggregates logs and root-cause notes from four different services into one long postmortem draft, arranged in the order the services were investigated. The final postmortem consistently omits the root cause identified from the second service investigated, which sits in the middle of the aggregated text.", "question": "What should the assistant change about how it structures the aggregated input?", "options": [{"letter": "A", "text": "Append a note at the very end of the document reminding the model to check the middle sections carefully", "correct": false, "explanation": "A reminder appended at the end does not restructure the content itself and is not a reliable substitute for surfacing the finding up front."}, {"letter": "B", "text": "Investigate the services in a different order each time so the root cause is not always found by the second service", "correct": false, "explanation": "Randomizing the investigation order does not fix the underlying issue; whichever finding lands in a middle position remains at risk of being dropped."}, {"letter": "C", "text": "Open the aggregated input with a brief summary of each service's root cause, then present detailed notes under clear headings", "correct": true, "explanation": "— an upfront summary of every service's root cause, paired with clearly headed detail sections, ensures the finding is visible regardless of which position it occupies in the investigation order."}, {"letter": "D", "text": "Combine all four services' notes into a single unbroken paragraph to reduce the total document length", "correct": false, "explanation": "Removing paragraph breaks reduces length but removes the structural cues that help the model locate specific findings, making the middle-drop problem more likely, not less."}], "correct": "C", "select": 1, "group": "G"}, {"id": "f5-012", "domain": 5, "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "situation": "A team is designing the system prompt for a support agent and wants Claude to reliably distinguish between cases that should be escalated and cases it can resolve autonomously.", "question": "Which approach best achieves this?", "options": [{"letter": "A", "text": "Add explicit escalation criteria plus few-shot examples showing both escalation and resolution cases", "correct": true, "explanation": "- explicit, example-driven criteria give the agent concrete patterns to distinguish escalation-worthy cases from ones it can resolve directly."}, {"letter": "B", "text": "Rely on the agent's general training to infer appropriate escalation behavior without adding specific guidance", "correct": false, "explanation": "- relying on general training without explicit, scenario-specific guidance leaves escalation behavior inconsistent and unpredictable."}, {"letter": "C", "text": "Instruct the agent to escalate whenever its own self-reported confidence falls below a fixed threshold", "correct": false, "explanation": "- self-reported confidence thresholds are an unreliable proxy for real case complexity and shouldn't drive escalation decisions."}, {"letter": "D", "text": "Instruct the agent to escalate whenever the customer's message contains any negative language or urgency", "correct": false, "explanation": "- negative language or urgency reflects sentiment, which is an unreliable proxy for whether a case actually needs escalation."}], "correct": "A", "select": 1, "group": "C"}, {"id": "f5-013", "domain": 5, "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "situation": "A billing support agent has been summarizing a lengthy chat every few turns to keep the prompt short. After the third summarization pass, the running summary reads \"the customer wants a refund soon and mentioned an order from last month,\" even though the original messages contained an exact order number, a refund amount of $214.88, and a customer-stated deadline of the 15th.", "question": "Which practice best addresses this failure mode going forward?", "options": [{"letter": "A", "text": "Have the customer restate the order number, refund amount, and deadline at the start of every new conversation turn", "correct": false, "explanation": "Repeatedly asking the customer to restate facts already given creates a poor experience and does not fix the underlying context-management gap."}, {"letter": "B", "text": "Increase the frequency of summarization passes so numeric details are refreshed more often before they are dropped", "correct": false, "explanation": "More frequent summarization still compresses details each time and does nothing to stop numeric precision from being lost; it only changes how often the loss happens."}, {"letter": "C", "text": "Switch to a model with the largest available context window so the transaction history never needs summarizing", "correct": false, "explanation": "A larger context window delays when compression is needed but does not prevent numeric detail loss once summarization does occur, and doesn't address the immediate failure already observed."}, {"letter": "D", "text": "Track the order number, exact refund amount, and stated deadline in a persistent case-facts block included in every prompt", "correct": true, "explanation": "— a persistent case-facts block keeps precise transactional details intact regardless of how many times the surrounding narrative gets condensed."}], "correct": "D", "select": 1, "group": "G"}, {"id": "f5-014", "domain": 5, "task_id": "5.5", "objective": "Design human review workflows and confidence calibration", "situation": "A reviewer team has capacity to manually check only 8% of the extractions produced daily by a Claude-based intake pipeline. Extractions vary in model-reported confidence, and some source documents contain contradictory values for the same field across pages.", "question": "How should the team prioritize which extractions receive human review?", "options": [{"letter": "A", "text": "Route low-confidence extractions and documents with ambiguous or contradictory source values to review first, using remaining capacity for a smaller random sample of high-confidence cases.", "correct": true, "explanation": "Reviewer capacity is best used by routing low‑confidence extractions and documents with ambiguous or contradictory source values to review first, since these are the extractions most likely to contain errors. The remaining capacity can then be used for a smaller random sample of high‑confidence cases to detect novel error patterns."}, {"letter": "B", "text": "Review only the extractions with the highest confidence scores, using all available review capacity on the extractions most likely to be trusted downstream, and forgo review of lower-confidence extractions.", "correct": false, "explanation": "Reviewing only the extractions with the highest confidence scores focuses on the cases least likely to contain errors. This leaves risky low‑confidence extractions and contradictory source documents completely unchecked, missing the very errors that review is meant to catch."}, {"letter": "C", "text": "Review a fixed random 8% of all extractions each day, regardless of confidence score or document characteristics, to maintain a statistically representative sample that catches errors uniformly.", "correct": false, "explanation": "A fixed random 8% sample ignores the team's existing signals—low confidence scores and contradictory source values—that indicate likely errors. This approach wastes scarce reviewer capacity on extractions that are probably correct while under‑reviewing the cases that most need human attention."}, {"letter": "D", "text": "Review whichever extractions were processed first each day, since earlier documents are statistically more likely to contain the fields reviewers care about, and defer later extractions to conserve capacity.", "correct": false, "explanation": "The order in which documents are processed has no proven relationship to error likelihood or the presence of important fields. Prioritizing by arrival time is not a principled basis for allocating review capacity and is likely to miss many errors."}], "correct": "A", "select": 1, "group": "C"}, {"id": "f5-015", "domain": 5, "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "situation": "A document-analysis subagent is reviewing internal financial filings and finds that two spreadsheets report different totals for the same quarter's operating expenses. The subagent's task is to hand its findings off to a coordinator before final synthesis occurs.", "question": "What should the subagent do with the conflicting values it found?", "options": [{"letter": "A", "text": "Include both values with clear annotation of which spreadsheet each came from, leaving reconciliation to the coordinator", "correct": true, "explanation": "A document-analysis subagent should surface conflicting values with their sources annotated, letting the coordinator decide how to reconcile them before the numbers reach final synthesis."}, {"letter": "B", "text": "Omit the operating expense figure entirely from its output since the two spreadsheets disagree with each other", "correct": false, "explanation": "Omitting the figure entirely discards potentially useful information and prevents the coordinator from being aware that a discrepancy needs resolution."}, {"letter": "C", "text": "Report only the value from the spreadsheet with the more recent file modification timestamp, omitting the other", "correct": false, "explanation": "File modification timestamps do not necessarily reflect which underlying figure is more accurate, and silently dropping one value hides the conflict from the coordinator."}, {"letter": "D", "text": "Recompute the operating expense total itself from first principles so the coordinator receives one authoritative number", "correct": false, "explanation": "Recomputing a new total invents a figure that appears in neither source document and removes the coordinator's ability to see and weigh the original disagreement."}], "correct": "A", "select": 1, "group": "B"}, {"id": "f5-016", "domain": 5, "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "situation": "A synthesis agent is finalizing a report that combines a subagent's structured findings about a software library's API surface with a subagent's narrative summary of community sentiment about the library. The draft currently renders the API findings as flowing prose paragraphs.", "question": "What change would improve this section?", "options": [{"letter": "A", "text": "Convert the community sentiment summary into a structured list of methods and parameters to match the API section's tone", "correct": false, "explanation": "Community sentiment is qualitative and does not map onto a structured list of methods and parameters, so forcing that format would misrepresent the content."}, {"letter": "B", "text": "Render both the API findings and sentiment summary as a shared table with columns for finding type and source subagent", "correct": false, "explanation": "Cramming both technical details and qualitative sentiment into one table format distorts content that is better served by different, content-appropriate formats."}, {"letter": "C", "text": "Render the API findings as a structured list of methods, parameters, and behaviors, keeping the sentiment summary as prose", "correct": true, "explanation": "Technical findings like API surfaces are best rendered as structured lists for scannability, while narrative sentiment is better preserved as prose, matching format to content type rather than forcing uniformity."}, {"letter": "D", "text": "Merge the API findings and sentiment summary into one continuous paragraph so the report reads as a unified narrative", "correct": false, "explanation": "Merging structured technical details into flowing prose makes it harder to scan for specific methods or parameters than a structured list would."}], "correct": "C", "select": 1, "group": "B"}, {"id": "f5-017", "domain": 5, "task_id": "5.5", "objective": "Design human review workflows and confidence calibration", "situation": "An architect asks Claude to output a confidence score from 0 to 1 for each extracted field in a structured JSON response.", "question": "Before using these scores to decide which fields skip human review, what step should the architect take to make the scores trustworthy for routing decisions?", "options": [{"letter": "A", "text": "Calibrate review thresholds against a labeled validation set, checking whether fields the model scores highly are actually correct at the rate that score implies before trusting it for routing.", "correct": true, "explanation": "A raw confidence score is only useful for routing once it has been calibrated: checking against labeled ground truth whether, say, 0.9-scored fields are actually correct roughly 90% of the time, and adjusting the review threshold based on that empirical relationship."}, {"letter": "B", "text": "Average each field's confidence score with the scores of neighboring fields in the same document to smooth out any single-field scoring noise before applying a threshold.", "correct": false, "explanation": "Averaging a field's score with unrelated neighboring fields conflates independent extraction decisions and provides no evidence that the resulting number reflects actual correctness likelihood."}, {"letter": "C", "text": "Replace numeric confidence scores with a categorical high/medium/low label, since categorical labels are inherently easier for a model to estimate accurately than continuous scores.", "correct": false, "explanation": "Converting to a categorical label doesn't address whether the underlying score is calibrated; an uncalibrated categorical label is just as unreliable for routing as an uncalibrated numeric one."}, {"letter": "D", "text": "Instruct the model to always output scores above 0.9 for fields it extracted successfully, since a consistently high score indicates the extraction pipeline is functioning correctly.", "correct": false, "explanation": "Forcing high scores for successful-looking extractions removes the signal the score is meant to carry; a self-consistently high score doesn't indicate the score maps to real-world correctness rates."}], "correct": "A", "select": 1, "group": "C"}, {"id": "f5-018", "domain": 5, "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "situation": "A research assistant aggregates outputs from five subagents (market sizing, competitor pricing, regulatory risk, customer sentiment, distribution channels) into one long combined document that is then passed to a synthesis step. The synthesis step's final memo omits the regulatory risk finding, which appeared in the third of five sections in the middle of the document.", "question": "What is the best way to structure the aggregated input to prevent this in future runs?", "options": [{"letter": "A", "text": "Reorder the sections so the regulatory risk finding is always discussed last, since Claude retains the most recent content best", "correct": false, "explanation": "Reordering to favor a single finding only protects that one item and still leaves every other middle section vulnerable in different runs."}, {"letter": "B", "text": "Instruct the synthesis step to read through the entire aggregated document twice before drafting its final memo", "correct": false, "explanation": "Rereading does not reliably compensate for position effects and adds latency without a structural fix."}, {"letter": "C", "text": "Place a short key-findings summary at the top of the aggregated document, then present each detailed section beneath an explicit heading", "correct": true, "explanation": "— leading with a concise key-findings summary and using explicit section headers surfaces critical content regardless of where it originated, directly countering the middle-of-document drop-off."}, {"letter": "D", "text": "Split the aggregated document into two shorter documents, without adding any findings summary or section headers", "correct": false, "explanation": "Splitting the document without adding a summary or headers still leaves each half susceptible to internal middle-of-input drop-off."}], "correct": "C", "select": 1, "group": "G"}, {"id": "f5-019", "domain": 5, "task_id": "5.3", "objective": "Implement error propagation strategies across multi-agent systems", "situation": "A legal-research subagent searches a case-law database for precedents matching a narrow fact pattern. The database responds successfully but the search genuinely matches zero cases. Separately, the same subagent's next query fails because the database connection pool is exhausted and refuses new connections.", "question": "The subagent reports both outcomes as 'no results found.' Why is this reporting flawed, and what should change?", "options": [{"letter": "A", "text": "It is flawed because the subagent should have terminated the entire workflow once the pool was exhausted", "correct": false, "explanation": "Terminating the entire workflow is an overreaction and not aligned with best practices. The subagent should report the failure clearly so that the workflow’s coordinator can decide on an appropriate course of action, such as retrying after backoff, switching to an alternate data source, or failing gracefully with a partial result set. Hard failure without notification reduces resilience and usability."}, {"letter": "B", "text": "It is not flawed at all, since both queries ultimately produced zero usable case citations for the coordinator to review", "correct": false, "explanation": "This is incorrect because the two outcomes have fundamentally different causes: one is a valid but empty result set, the other is an operational failure. Treating them identically hides a critical infrastructure issue that may affect reliability and data completeness. The coordinator cannot differentiate between an exhaustive, accurate search and a failed attempt, leading to potential misinterpretation of legal research completeness."}, {"letter": "C", "text": "It is flawed only because 'no results found' is too informal; formal legal wording would fix the issue", "correct": false, "explanation": "The formality of the message is irrelevant to the core problem. The flaw lies in misclassifying a system failure as a successful but empty search, not in the specific phrasing. Adopting formal legal language would not correct the underlying error handling deficiency; the subagent must signal the nature of the failure to enable proper orchestration logic."}, {"letter": "D", "text": "It conflates a completed zero-match search with a connection failure; the latter should be a distinct access failure", "correct": true, "explanation": "Anthropic's error handling distinguishes between successful tool calls that return no data and tool execution failures. A zero-match search is a successful execution that should return a clear 'No results' indication via tool_result without is_error: True. A connection pool exhaustion is a system-level failure that should be reported as an error (e.g., is_error: True with an appropriate error message). Conflating them prevents the coordinator from making informed decisions about retries, fallbacks, or alerting."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f5-020", "domain": 5, "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "situation": "A legal research assistant built on Claude synthesizes findings from three case-law subagents. One subagent flags that a particular precedent has been cited approvingly by every appellate court that has reviewed it, while another flags a claim that appears in only one lower-court opinion and has not been tested elsewhere. Both claims are currently presented in the same paragraph with identical phrasing.", "question": "What revision best serves the report's users?", "options": [{"letter": "A", "text": "Leave both claims in the same paragraph but add a footnote number to each, without changing how confidently either is described", "correct": false, "explanation": "Adding footnote numbers provides a citation trail but does not address the underlying problem that both claims are described with the same confidence despite very different levels of support."}, {"letter": "B", "text": "Separate the claims into sections labeled by evidentiary weight, describing the precedent and single-opinion claim differently", "correct": true, "explanation": "Distinguishing widely affirmed precedent from a single unconfirmed opinion in clearly labeled sections preserves each source's original characterization and gives readers an accurate sense of evidentiary weight."}, {"letter": "C", "text": "Move the single-opinion claim to an appendix without any accompanying language indicating its limited support", "correct": false, "explanation": "Moving the claim to an appendix without any qualifying language still fails to signal that it rests on much thinner support than the widely affirmed precedent."}, {"letter": "D", "text": "Rephrase both claims using identical hedging language such as 'courts have suggested' so neither appears more authoritative", "correct": false, "explanation": "Applying identical hedging language to both claims erases the real difference in evidentiary weight between a widely affirmed precedent and a single unconfirmed opinion."}], "correct": "B", "select": 1, "group": "B"}, {"id": "f5-021", "domain": 5, "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "situation": "A support team relies solely on repeated summarization of long chats to keep prompts short, without any separate structured tracking of transactional details.", "question": "Over many support sessions, what is the most likely consequence for exact figures such as amounts, dates, and order numbers mentioned early in a conversation?", "options": [{"letter": "A", "text": "They are progressively generalized or dropped from the summary, since summarization optimizes for brevity over specific values", "correct": true, "explanation": "— each summarization pass tends to compress specific values into vaguer language, so exact figures are the ones most likely to degrade or disappear across repeated passes without separate structured tracking."}, {"letter": "B", "text": "They remain exactly as stated indefinitely, since summarization only condenses narrative language and never affects numeric values", "correct": false, "explanation": "Summarization inherently condenses text, including numeric details, so exact values are not immune to being generalized away."}, {"letter": "C", "text": "They become more accurate over time as the model has more opportunities to notice and correct any errors in the original figures", "correct": false, "explanation": "Summarization does not add new information or verify prior figures; it condenses existing text, so it cannot make figures more accurate."}, {"letter": "D", "text": "They are automatically moved into a separate memory store by the summarization process without any additional configuration", "correct": false, "explanation": "No automatic mechanism relocates values to persistent storage; without deliberate structured extraction, exact figures are only present in the same summarized narrative as everything else."}], "correct": "A", "select": 1, "group": "G"}, {"id": "f5-022", "domain": 5, "task_id": "5.3", "objective": "Implement error propagation strategies across multi-agent systems", "situation": "A regulatory-filings subagent cannot reach its data provider because the provider's API key expired. Instead of surfacing the failure, the subagent returns an empty findings list with status 'success' so the pipeline continues cleanly. The coordinator later synthesizes a report claiming full regulatory coverage.", "question": "What went wrong?", "options": [{"letter": "A", "text": "The coordinator is at fault for trusting subagent output; every subagent result should require manual verification instead", "correct": false, "explanation": "Requiring manual verification of every subagent result defeats the purpose of automated multi-agent coordination and does not address the actual defect, which is the mislabeled status."}, {"letter": "B", "text": "The subagent silently suppressed an access failure as success, so the coordinator reported false full coverage instead of a real gap", "correct": true, "explanation": "Reporting a genuine access failure as an empty success is a classic anti-pattern: it hides the failure from the coordinator, which then has no way to know its synthesized output has an undisclosed coverage gap."}, {"letter": "C", "text": "The subagent should have terminated the entire multi-agent workflow rather than letting any other subagents keep running", "correct": false, "explanation": "Terminating the whole workflow over one subagent's failure is the opposite anti-pattern; the fix is accurate error reporting, not an all-or-nothing shutdown."}, {"letter": "D", "text": "The subagent correctly avoided alarming the coordinator about a minor credential issue that would resolve on the next scheduled data run", "correct": false, "explanation": "An expired API key is not self-resolving and is exactly the kind of failure that must be surfaced, since the coordinator cannot distinguish it from a legitimately empty result once it is disguised as success."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f5-023", "domain": 5, "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "situation": "A sales assistant queries a CRM contact-lookup tool that returns 60 fields per contact record, including internal lead-scoring metadata, marketing campaign tags, and audit timestamps, when only the contact's name, company, deal stage, and last contact date matter for drafting a follow-up email. After looking up ten contacts, the raw records dominate the context.", "question": "What should the assistant do?", "options": [{"letter": "A", "text": "Look up each contact only once per session and rely on memory of the fields afterward, without keeping the raw output at all", "correct": false, "explanation": "Discarding the lookup output entirely risks losing even the four relevant fields the assistant needs for the current task, since nothing is kept in context."}, {"letter": "B", "text": "Ask the CRM tool to return records in a more compact text format while still including all 60 fields", "correct": false, "explanation": "A more compact text encoding of the same 60 fields still includes irrelevant metadata and does not address which fields are actually needed."}, {"letter": "C", "text": "Store all 60 fields from each lookup in context so the assistant has complete information available for any future question", "correct": false, "explanation": "Retaining all 60 fields for every contact reproduces the same disproportionate token cost the assistant is trying to avoid, just for more contacts."}, {"letter": "D", "text": "Extract only the name, company, deal stage, and last contact date from each record before adding it to context", "correct": true, "explanation": "— keeping only the fields relevant to drafting the follow-up email prevents the assistant's context from being consumed by unrelated CRM metadata."}], "correct": "D", "select": 1, "group": "G"}, {"id": "f5-024", "domain": 5, "task_id": "5.3", "objective": "Implement error propagation strategies across multi-agent systems", "situation": "A synthesis-stage coordinator combines findings from six research subagents into a final report. Two subagents lost access to their assigned sources partway through.", "question": "Which output structure best communicates the reliability of the final report to a downstream reader?", "options": [{"letter": "A", "text": "Append one generic disclaimer noting that some unspecified sources may have been unavailable overall", "correct": false, "explanation": "A generic, non-specific disclaimer doesn't tell the reader which sections are affected, so they cannot judge which claims to trust less."}, {"letter": "B", "text": "Annotate coverage so readers see which sections are well-supported by completed sources and which have gaps", "correct": true, "explanation": "Coverage annotations let the reader see exactly which claims rest on solid evidence and which areas are thin because of source access problems, preserving both value and honesty."}, {"letter": "C", "text": "Omit the two affected sections entirely so the reader only sees content from fully successful subagents", "correct": false, "explanation": "Dropping the affected sections silently removes information the reader might still find useful and hides that a gap exists at all."}, {"letter": "D", "text": "Present all findings as one uniform narrative with no distinction between complete and partial-source sections", "correct": false, "explanation": "A uniform narrative with no distinction implies every section has equal evidentiary support, which misrepresents the sections affected by the access failures."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f5-025", "domain": 5, "task_id": "5.3", "objective": "Implement error propagation strategies across multi-agent systems", "situation": "A financial-data subagent queries a market feed for a ticker's after-hours trades. The feed's cache is stale beyond its allowed threshold, so the subagent's read fails an internal freshness check and aborts. A different subagent queries a competitor's after-hours trades and legitimately finds no trades occurred that session.", "question": "How should the coordinator distinguish these two 'no data' situations?", "options": [{"letter": "A", "text": "Report the stale-cache abort as an access failure eligible for retry, and the no-trades result as a final empty result", "correct": true, "explanation": "The stale-cache abort is an access problem where retrying against a different or refreshed source could yield real data, whereas the no-trades result is a legitimate, complete answer that a retry would not change."}, {"letter": "B", "text": "Report both as plain empty results, since neither subagent has any usable after-hours trade data for the coordinator to review", "correct": false, "explanation": "Treating both as plain empty results hides that one of them is a data-quality failure that retrying could resolve, losing that recovery opportunity."}, {"letter": "C", "text": "Escalate both as unrecoverable errors that halt processing for both tickers until a human resolves them", "correct": false, "explanation": "Escalating the legitimate no-trades result as an unrecoverable error is unnecessary and would halt processing over a correct, already-final answer."}, {"letter": "D", "text": "Report the stale-cache abort as a valid empty result, and the no-trades session as a failure needing retry", "correct": false, "explanation": "This reverses the two cases: the stale cache is the genuine access problem, and the no-trades session is the legitimate empty result, not the other way around."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f5-026", "domain": 5, "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "situation": "A large multi-agent exploration run is paused partway through, with several agents already having completed and returned results. The architect resumes the run within the same session.", "question": "What happens to the agents that had already completed?", "options": [{"letter": "A", "text": "Every agent, including the ones that already completed, is re-run from the beginning of the workflow.", "correct": false, "explanation": "Re-running every agent, including completed ones, wastes the work and cost already spent and is not how a resumed run behaves within the same session."}, {"letter": "B", "text": "Their cached results are reused, and only the remaining, not-yet-completed agents run live.", "correct": true, "explanation": "Resuming within the same session reuses the cached results of agents that already finished and only runs the remaining agents live, avoiding redundant work."}, {"letter": "C", "text": "The run can only be resumed by exiting and starting an entirely new Claude Code installation on the machine.", "correct": false, "explanation": "Resuming happens within the existing session; exiting the session, by contrast, causes the next session to start the workflow fresh instead."}, {"letter": "D", "text": "All previously completed results are discarded and the run cannot be resumed at all after being paused.", "correct": false, "explanation": "Completed results are not discarded on resume; the run is specifically designed to be resumable within the same session rather than starting over."}], "correct": "B", "select": 1, "group": "G"}, {"id": "f5-027", "domain": 5, "task_id": "5.5", "objective": "Design human review workflows and confidence calibration", "situation": "An insurance-claims extraction pipeline processes auto, home, and health claims. The team has a labeled evaluation dataset covering all three types but has only computed an aggregate accuracy figure across the entire dataset.", "question": "Before deciding whether to reduce human review for health claims specifically, what is the most direct way to proceed?", "options": [{"letter": "A", "text": "Increase the aggregate sample size by adding more documents from auto, home, and health claims equally until the combined confidence interval narrows sufficiently to infer that the health‑claims segment also meets the required accuracy threshold.", "correct": false, "explanation": "Expanding the overall sample does not guarantee that the health‑claims subset meets the accuracy bar. The confidence interval of the aggregate metric may shrink, but the health‑claims performance could still be far below the threshold. Direct evaluation of the target segment is necessary, not inference from a mixed population."}, {"letter": "B", "text": "Ask the extraction model to retrospectively estimate its accuracy per document type from its confidence scores, and then use those self‑reported figures for health claims instead of obtaining a labeled ground‑truth dataset.", "correct": false, "explanation": "Model‑reported confidence scores are an unreliable substitute for true ground‑truth evaluation. They are often miscalibrated—especially in unfamiliar data slices—and can overestimate accuracy. Anthropic’s documentation stresses the need for external verification, groundedness checks, and traceability (as seen in partnerships like Allianz). Self‑assessment without labeled data cannot be trusted for a safety‑critical decision like reducing human review in healthcare."}, {"letter": "C", "text": "Compute accuracy separately for the health-claims examples in the existing evaluation set, and only reduce review if that segment's accuracy independently meets the required bar.", "correct": true, "explanation": "This is the recommended approach. Anthropic emphasizes the importance of segment‑specific evaluation, particularly in regulated domains like healthcare. Aggregate accuracy numbers can hide poor performance in a subset. By directly measuring the health‑claims segment against its own ground‑truth labels, you obtain a reliable estimate of that segment’s quality. Only if the segment‑level accuracy meets the predefined threshold should human review be reduced. (See Anthropic’s guidance on ground‑truth sets and segment‑aware evaluation in 'Demystifying Evals for AI Agents' and healthcare‑specific best practices.)"}, {"letter": "D", "text": "Assume that the health‑claims extraction accuracy mirrors the combined aggregate accuracy, because the pipeline uses the same prompt and model across document types, so the overall figure serves as a reliable proxy for the health‑claims segment.", "correct": false, "explanation": "This assumption is unsafe and contradicted by Anthropic’s guidance. A model may perform well on auto claims but poorly on complex health claims that involve clinical terminology and coding. Aggregate accuracy is not a trustworthy proxy for a specific subpopulation. Making this assumption could allow undetected errors in health claims to reach production."}], "correct": "C", "select": 1, "group": "C"}, {"id": "f5-028", "domain": 5, "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "situation": "A bank's monitoring system flags a conversation as having very negative sentiment because the customer used sharp language while asking the agent to reset their online banking password, a routine, fully self-service-eligible request.", "question": "Should the agent escalate based on the sentiment flag?", "options": [{"letter": "A", "text": "Yes, negative sentiment scores reliably indicate that a case is too complex for the agent to resolve", "correct": false, "explanation": "- sentiment scores reflect tone, not the underlying difficulty of the case, so they should not be treated as a reliable complexity signal."}, {"letter": "B", "text": "No, sentiment alone is not a reliable indicator of complexity, and the password reset is within the agent's capability", "correct": true, "explanation": "- sentiment is an unreliable proxy for actual case complexity, and since the request itself is simple and resolvable, it should be handled directly."}, {"letter": "C", "text": "Yes, the negative sentiment flag should always trigger escalation regardless of the actual underlying issue difficulty", "correct": false, "explanation": "- automatically escalating on any sentiment flag ignores whether the underlying issue actually requires human intervention."}, {"letter": "D", "text": "No, but only because password resets are always exempt from any sentiment-based escalation rules", "correct": false, "explanation": "- the reasoning is wrong; the correct issue is that sentiment isn't a reliable complexity signal in general, not that password resets have a special exemption."}], "correct": "B", "select": 1, "group": "C"}, {"id": "f5-029", "domain": 5, "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "situation": "A utility company customer needs a multi-step billing correction involving a meter re-read, a prorated credit, and a plan adjustment. Every step is explicitly detailed in the documented billing policy, and the agent has tools to execute each step.", "question": "Should the agent escalate this case simply because it involves several steps?", "options": [{"letter": "A", "text": "Yes, any case requiring more than one corrective action should be escalated regardless of policy coverage", "correct": false, "explanation": "- escalation should be driven by actual triggers like policy gaps or customer requests, not simply by how many steps a resolution requires."}, {"letter": "B", "text": "Yes, multi-step cases are inherently too complex for an agent to execute reliably without human oversight", "correct": false, "explanation": "- a case being multi-step does not make it inherently unresolvable by the agent when each step is clearly defined by policy."}, {"letter": "C", "text": "No, the case should be resolved directly; step count alone is not an escalation trigger when policy fully covers it", "correct": true, "explanation": "- the number of steps alone doesn't determine whether a case should be escalated; since policy fully covers each step, the agent can resolve it directly."}, {"letter": "D", "text": "No, but only because billing corrections are categorically exempt from any complexity-based escalation rule entirely", "correct": false, "explanation": "- the reasoning is wrong; there's no special exemption for billing corrections, the actual principle is that step count alone isn't a valid escalation trigger."}], "correct": "C", "select": 1, "group": "C"}, {"id": "f5-030", "domain": 5, "task_id": "5.3", "objective": "Implement error propagation strategies across multi-agent systems", "situation": "During a customer-support triage workflow, a knowledge-base subagent cannot reach its vector index because the index service returned a 429 rate-limit response. The subagent has already retried twice locally with exponential backoff and is still being rate-limited.", "question": "What should the subagent do next?", "options": [{"letter": "A", "text": "Immediately fail the entire triage workflow so a human must review every ticket in the current batch manually", "correct": false, "explanation": "Failing the entire batch workflow over one rate-limited subagent needlessly blocks unrelated tickets that have nothing to do with the affected knowledge base."}, {"letter": "B", "text": "Keep retrying locally on the same backoff schedule indefinitely, since a 429 will always eventually resolve given enough retries", "correct": false, "explanation": "Continuing to retry indefinitely without escalating risks stalling the workflow and never gives the coordinator a chance to make a different decision, such as routing around the rate-limited service."}, {"letter": "C", "text": "Escalate with the failure type, the attempted queries, and partial results, since more local retries seem unlikely to help", "correct": true, "explanation": "After exhausting reasonable local retries against a persistent rate limit, the subagent should hand the problem to the coordinator with enough context to decide on next steps, such as trying a different index or deferring the ticket."}, {"letter": "D", "text": "Return an empty result and mark the ticket resolved, since the knowledge base could not be reached in time", "correct": false, "explanation": "Marking the ticket resolved when the knowledge base was never actually queried successfully is the silent-suppression anti-pattern and would leave the customer's issue unaddressed."}], "correct": "C", "select": 1, "group": "A"}, {"id": "f5-031", "domain": 5, "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "situation": "Three market-research subagents each report a total addressable market figure for the same industry, but none includes the publication date or methodology of the source they used. The synthesis step cannot tell whether the figures are comparable or which is most current, and the final report cites an outdated figure.", "question": "What subagent output requirement would have prevented this?", "options": [{"letter": "A", "text": "Require each subagent to report a single average figure computed across all of the sources it found", "correct": false, "explanation": "An averaged figure blends comparable and incomparable estimates together, and would not have flagged that one source was outdated to begin with."}, {"letter": "B", "text": "Require each subagent to include the source's publication date and methodological context alongside its reported figure", "correct": true, "explanation": "— publication date and methodology let the synthesis step judge comparability and recency, directly preventing an outdated figure from being cited as current."}, {"letter": "C", "text": "Require the synthesis step to always prefer whichever figure the subagents reported as the highest", "correct": false, "explanation": "Always preferring the highest figure is an arbitrary rule unrelated to recency or comparability, and could just as easily select an outdated figure."}, {"letter": "D", "text": "Require each subagent to describe in more detail its reasoning process for how it searched for information", "correct": false, "explanation": "A more detailed search narrative does not supply the specific date and methodology needed to judge whether a figure is current."}], "correct": "B", "select": 1, "group": "G"}, {"id": "f5-032", "domain": 5, "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "situation": "A legal-research orchestrator combines the outputs of eight subagents, each of which returns several paragraphs of step-by-step reasoning plus their conclusion. The combined output exceeds what the downstream drafting agent can process within its allotted context budget, forcing it to drop some subagent findings entirely before drafting.", "question": "What change addresses this?", "options": [{"letter": "A", "text": "Have the drafting agent read the subagent outputs in two passes to catch findings it missed the first time", "correct": false, "explanation": "A second read-through does not reduce the volume of text competing for the budget and does not guarantee previously dropped findings are recovered."}, {"letter": "B", "text": "Modify the subagents to return only key facts, citations, and relevance scores instead of their full reasoning narratives", "correct": true, "explanation": "— returning structured facts, citations, and relevance scores instead of verbose reasoning trims each subagent's contribution to what the drafting agent actually needs, fitting more findings within its context budget."}, {"letter": "C", "text": "Instruct the orchestrator to forward only the subagent outputs that arrive first, discarding later ones", "correct": false, "explanation": "Discarding findings based on arrival order rather than relevance risks dropping the most important legal conclusions simply because they came in late."}, {"letter": "D", "text": "Increase the number of subagents so each one covers a narrower topic while still returning full reasoning chains", "correct": false, "explanation": "Adding more subagents while keeping full reasoning chains increases, rather than reduces, the total volume the drafting agent must fit into its limited budget."}], "correct": "B", "select": 1, "group": "G"}, {"id": "f5-033", "domain": 5, "task_id": "5.3", "objective": "Implement error propagation strategies across multi-agent systems", "situation": "A coordinator dispatches a pricing-lookup subagent against an internal catalog API. The subagent's HTTP call hangs and exceeds its timeout budget before any response arrives. A separate subagent queries the same catalog for a discontinued SKU and receives a 200 response with zero matching rows.", "question": "How should the two outcomes be reported to the coordinator?", "options": [{"letter": "A", "text": "Report the timeout as an access failure eligible for retry, and the zero-row response as a valid empty result needing no retry", "correct": true, "explanation": "Timeout errors are transient failures that should be retried with exponential backoff, as recommended by Anthropic's official documentation. A 200 response with zero matching rows for a discontinued SKU is an expected, valid empty result and does not warrant a retry. The application should handle the empty state gracefully."}, {"letter": "B", "text": "Report the timeout as a valid empty result and the zero-row response as an access failure that needs a retry", "correct": false, "explanation": "This reverses the correct classification. A timeout should be treated as a failure eligible for retry, while a 200 OK with zero rows is a successful empty response that should not trigger retries."}, {"letter": "C", "text": "Report both outcomes as access failures so the coordinator retries each lookup the same fixed number of times", "correct": false, "explanation": "Classifying a valid zero-row response as an access failure leads to unnecessary retries that waste resources. Retries are appropriate only for transient errors like timeouts, not for successful but empty responses."}, {"letter": "D", "text": "Report both outcomes as empty results, since neither subagent returned any usable pricing data for the coordinator to act upon", "correct": false, "explanation": "A timeout is not an empty result; it is a failure indicating the request did not complete, which is distinct from a successful response with no data. Treating it as empty would miss the opportunity to retry and potentially recover."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f5-034", "domain": 5, "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "situation": "An architect reviews a synthesis agent's prompt and finds that it instructs the agent to 'write a concise unified summary of all subagent findings.' Reports produced under this prompt read smoothly but reviewers can no longer tell which subagent, document, or date each specific claim traces back to.", "question": "Which prompt revision most directly addresses this problem?", "options": [{"letter": "A", "text": "Instruct the synthesis agent to shorten the report further so that reviewers can read through it more quickly", "correct": false, "explanation": "Shortening the report further would likely compress content even more and make it harder, not easier, to trace claims back to their sources."}, {"letter": "B", "text": "Instruct the synthesis agent to write in a more formal register so the report appears more authoritative to reviewers", "correct": false, "explanation": "A more formal writing register does not restore any connection between individual claims and their originating sources or dates."}, {"letter": "C", "text": "Instruct the synthesis agent to carry forward each claim's source and date metadata from subagent outputs into the report", "correct": true, "explanation": "The root cause is that the synthesis prompt only asks for a unified summary without directing the agent to retain source and date metadata; explicitly requiring that carry-through preserves provenance in the final report."}, {"letter": "D", "text": "Instruct the synthesis agent to add a general disclaimer at the end stating that sources are available upon request", "correct": false, "explanation": "A generic disclaimer that sources exist elsewhere does not tell reviewers which source supports which specific claim in the report itself."}], "correct": "C", "select": 1, "group": "B"}, {"id": "f5-035", "domain": 5, "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "situation": "An insurance customer asks whether their policy covers a rental car while their electric vehicle's battery is being replaced under a separate manufacturer recall. The policy documentation addresses rental reimbursement only for collision and comprehensive claims, and does not mention manufacturer recall repairs.", "question": "What should the agent do?", "options": [{"letter": "A", "text": "Approve rental reimbursement, since recall repairs are similar enough in nature to comprehensive collision claims", "correct": false, "explanation": "- approving reimbursement by analogy to comprehensive claims extends coverage beyond what the documented policy states."}, {"letter": "B", "text": "Escalate the question, since the policy documentation does not address this specific recall-related scenario", "correct": true, "explanation": "- the policy is silent on rental coverage during recall repairs, a gap that should be escalated rather than resolved by the agent's own judgment."}, {"letter": "C", "text": "Deny rental reimbursement, since the recall repair is not listed among the covered claim types", "correct": false, "explanation": "- treating an unaddressed scenario as automatically excluded asserts a coverage decision the policy never actually makes."}, {"letter": "D", "text": "Direct the customer to the manufacturer, since the recall is the underlying cause of the repair", "correct": false, "explanation": "- redirecting the customer to the manufacturer avoids answering the coverage question rather than resolving the policy gap appropriately."}], "correct": "B", "select": 1, "group": "C"}, {"id": "f5-036", "domain": 5, "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "situation": "During a long exploration session, an architect notices the context window filling with verbose file dumps and command output from the current phase, and the key findings from that phase have not yet been written anywhere durable. The architect wants to reclaim context space before continuing.", "question": "What should they do first?", "options": [{"letter": "A", "text": "End the current session entirely and start a brand new one with no reference to any prior exploration work at all.", "correct": false, "explanation": "Ending the session discards all context, including findings that were never persisted, forcing the architect to redo the exploration."}, {"letter": "B", "text": "Continue exploring without compacting anything or recording findings until the session eventually runs entirely out of context space.", "correct": false, "explanation": "Continuing without freeing space eventually forces an uncontrolled compaction or truncation at an inconvenient point, rather than a deliberate one after findings are recorded."}, {"letter": "C", "text": "Persist the key findings from the current phase to a scratchpad file, and only then compact the conversation to reclaim space.", "correct": true, "explanation": "Compaction replaces the verbatim conversation with a summary, so writing key findings to a scratchpad file first ensures they survive compaction intact."}, {"letter": "D", "text": "Immediately compact the conversation without recording anything, trusting the details stay accessible afterward.", "correct": false, "explanation": "Compacting without saving findings risks losing exact details, such as specific file paths or discovered class names, that a summary may compress or omit."}], "correct": "C", "select": 1, "group": "G"}, {"id": "f5-037", "domain": 5, "task_id": "5.5", "objective": "Design human review workflows and confidence calibration", "situation": "A team runs an initial validation showing 95% accuracy across all document types and removes human review for high-confidence extractions. A stakeholder proposes stopping ongoing sampling altogether, arguing that the initial validation already proved the pipeline works.", "question": "What is the strongest argument against permanently stopping sampling after initial validation?", "options": [{"letter": "A", "text": "Initial validation only measures performance on the document population and conditions present at that time, and ongoing stratified sampling is needed to detect later shifts in error rates or new failure patterns as inputs change.", "correct": true, "explanation": "Initial validation captures a snapshot of performance, but document populations and conditions can evolve, leading to new error patterns or shifts in error rates. Ongoing stratified sampling is necessary to detect such drift and ensure sustained reliability."}, {"letter": "B", "text": "Ongoing sampling provides the essential evidence that auditors and regulators require to confirm the pipeline operates within acceptable error bounds, and without it the organization may face compliance findings regardless of the initial validation.", "correct": false, "explanation": "This argument emphasizes compliance and audit requirements, but the strongest reason to continue sampling is the technical need to monitor for real accuracy degradation, not just to satisfy regulators. Relying solely on compliance framing overlooks the primary purpose of detecting drift."}, {"letter": "C", "text": "Stopping sampling would violate a regulation that requires continuous sampling at a prescribed rate for all automated extraction systems, and without such sampling the pipeline would be non-compliant even if its initial performance appeared satisfactory.", "correct": false, "explanation": "There is no universal regulation that mandates continuous sampling at a prescribed rate for all automated extraction systems. The justification for ongoing sampling is the practical necessity of detecting drift and novel failure patterns, not an imagined blanket legal requirement."}, {"letter": "D", "text": "Sampling should continue indefinitely because the extraction model’s performance inevitably degrades over time due to accumulated data drift, much like mechanical components wear out, making periodic checks essential to catch hidden accuracy drops.", "correct": false, "explanation": "A machine learning model does not physically degrade or wear out over time like mechanical components; instead, the risk stems from changes in the input data distribution (data drift). Periodic sampling is essential to catch these shifts, not because the model inherently degrades with usage."}], "correct": "A", "select": 1, "group": "C"}, {"id": "f5-038", "domain": 5, "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "situation": "Several subagents are running concurrently, each investigating a different module of a large codebase and recording findings as they go. The architect must prevent one subagent's findings from being overwritten by another before the coordinator can reliably aggregate the results.", "question": "Which scratchpad convention best achieves this?", "options": [{"letter": "A", "text": "Skip scratchpad files entirely and have each subagent only report findings verbally in its final response.", "correct": false, "explanation": "Relying solely on verbal reporting risks context pollution and makes it harder to preserve intermediate results. The recommended pattern is to persist findings to isolated scratchpads or external storage, which the coordinator can then reliably read and synthesize."}, {"letter": "B", "text": "Have every subagent overwrite the same fixed scratchpad filename each time it records a newly discovered finding.", "correct": false, "explanation": "Using a single fixed filename results in the last write from any subagent overwriting all previous findings, directly causing the problem the architect wants to avoid. Unique filenames per subagent are required for safe concurrent recording."}, {"letter": "C", "text": "Have all subagents write their findings to one single shared scratchpad file at the same time as they each discover them.", "correct": false, "explanation": "A single shared scratchpad file creates race conditions: concurrent writes from multiple subagents can lead to lost or corrupted data before the coordinator can aggregate the results. Isolated files per subagent are necessary to prevent such conflicts."}, {"letter": "D", "text": "Give each subagent its own uniquely named scratchpad file, e.g., one per module, and have the coordinator aggregate all files only after all subagents have finished their investigations.", "correct": true, "explanation": "Anthropic's multi-agent architecture guidelines recommend isolating subagent output to prevent interference. Each subagent should write to its own scratchpad or external memory (e.g., uniquely named files). The coordinator then aggregates these separate outputs once all subagents complete, avoiding race conditions and ensuring data integrity."}], "correct": "D", "select": 1, "group": "G"}, {"id": "f5-039", "domain": 5, "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "situation": "A patient messages a healthcare scheduling assistant in a clearly irritated tone about needing to reschedule a routine appointment with no clinical urgency. The assistant can complete the rescheduling directly, and the patient has not asked to speak with a staff member.", "question": "What should the assistant do?", "options": [{"letter": "A", "text": "Acknowledge the patient's frustration, signal that the rescheduling can be handled right away, and also offer to escalate to a staff member if the patient prefers.", "correct": true, "explanation": "This reflects the recommended approach from user feedback and aligns with Anthropic's guidance that healthcare scheduling assistants should provide empathetic, human-like interactions while maintaining clear pathways for human escalation when the patient wants it. Acknowledging frustration reduces friction, signalling immediate resolution addresses the request, and offering escalation respects patient autonomy without forcing them to choose. It does not escalate solely based on tone, nor does it withhold an option that may be comforting."}, {"letter": "B", "text": "Escalate to a staff member immediately, since the irritated tone suggests the situation needs human handling.", "correct": false, "explanation": "Immediate escalation based only on an irritated tone is not supported. Escalation is appropriate when the agent cannot complete the request, when clinical judgment is required, or when the patient explicitly requests a human. Tone alone is not a trigger for immediate handoff; the correct action is to acknowledge, handle the request, and offer escalation as an option."}, {"letter": "C", "text": "Reschedule the appointment without acknowledging the tone, treating it as unrelated to completing the request directly.", "correct": false, "explanation": "This is incomplete. The agent should complete the rescheduling because it is within capability, but ignoring the patient's frustration misses the recommended patient-centered communication that acknowledges the sentiment and reduces friction. Anthropic emphasizes human-like, meaningful interactions; failing to acknowledge tone feels robotic and may increase dissatisfaction."}, {"letter": "D", "text": "Ask the patient to confirm they don't want to speak with a staff member before proceeding with the reschedule.", "correct": false, "explanation": "This adds unnecessary friction. Since the request is within capability, the assistant should proceed with acknowledgment and direct rescheduling while actively offering escalation as an option. Requiring confirmation before acting imposes an extra step that the patient did not ask for and may worsen frustration."}], "correct": "A", "select": 1, "group": "C"}, {"id": "f5-040", "domain": 5, "task_id": "5.5", "objective": "Design human review workflows and confidence calibration", "situation": "During calibration, a team finds that fields the model scores at 0.95 confidence are correct only 78% of the time, while fields scored at 0.6 confidence are correct 90% of the time.", "question": "What does this pattern indicate, and what should the team do?", "options": [{"letter": "A", "text": "The validation set is too small to draw any conclusion about field-level confidence calibration, so the team should discard field-level confidence scoring entirely and rely only on document-type stratified sampling, redirecting all fields to human review.", "correct": false, "explanation": "A calibration discrepancy does not automatically imply the validation set is too small to be meaningful. Discarding field-level confidence scoring entirely and redirecting all fields to human review is an overreaction; the team should first investigate the specific pattern before abandoning the approach."}, {"letter": "B", "text": "This pattern is expected behavior for well-calibrated models, since lower scores naturally correspond to higher observed accuracy on any validation set, and the team should therefore continue using the raw confidence scores as a routing threshold without any recalibration.", "correct": false, "explanation": "In a well-calibrated model, higher confidence should correspond to higher observed accuracy, not an inverse relationship. This pattern is a clear sign of calibration error, so continuing to use raw scores without recalibration would be misguided."}, {"letter": "C", "text": "The 0.6-confidence fields must belong to an easier field type, so the team should raise the routing threshold to 0.96, effectively routing all fields with scores below 0.96 to human review without adjusting the model's calibration.", "correct": false, "explanation": "Assuming without evidence that the 0.6-confidence fields belong to an easier field type is an unverified assumption. Raising the threshold to 0.96 as a blanket fix treats the miscalibration symptom as benign, risking incorrect routing based on a flawed premise."}, {"letter": "D", "text": "The model's confidence scores are miscalibrated and inversely related to actual correctness for these ranges, so the team should not use the raw scores directly to set a simple 'route below X' threshold without further investigation.", "correct": true, "explanation": "The observed pattern shows higher confidence scores associated with lower actual correctness, indicating miscalibration. The team should not treat the raw scores as reliable for setting a simple threshold and must investigate the underlying cause before making routing decisions."}], "correct": "D", "select": 1, "group": "C"}, {"id": "f5-041", "domain": 5, "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "situation": "A synthesis agent is producing a final report that combines a subagent's stock price and revenue trend findings with a subagent's qualitative summary of recent news coverage about the same company. Both are converted into the same bulleted list format in the draft report.", "question": "What change would most improve how this content is rendered?", "options": [{"letter": "A", "text": "Convert the financial figures into prose paragraphs so they read alongside the news coverage without a jarring format change", "correct": false, "explanation": "Converting tabular financial data into prose paragraphs makes it harder to scan and compare figures than a table would."}, {"letter": "B", "text": "Render the financial figures as a table and keep the news coverage as prose, matching each content type to its natural format", "correct": true, "explanation": "Financial data is best rendered as a table for easy comparison, while news content reads more naturally as prose; matching format to content type preserves clarity instead of flattening everything into one uniform shape."}, {"letter": "C", "text": "Combine the figures and news coverage into a single bulleted list that interleaves numbers and narrative sentences item by item", "correct": false, "explanation": "Interleaving numeric data and narrative sentences in one list mixes formats that serve different purposes and makes both harder to read."}, {"letter": "D", "text": "Convert the news coverage into the same bulleted list format as the financial figures so the whole report has one consistent style", "correct": false, "explanation": "Forcing narrative news coverage into a bulleted list strips out the connective context and nuance that prose naturally conveys."}], "correct": "B", "select": 1, "group": "B"}, {"id": "f5-042", "domain": 5, "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "situation": "Late in a multi-hour exploration session, an architect asks the agent to describe the error-handling pattern in the billing service. Earlier, a subagent had discovered that the service uses a custom Result type instead of exceptions. The agent now answers that the service 'typically uses try/catch with logging,' contradicting the earlier finding.", "question": "What is the most likely cause, and what should the architect have done to prevent it?", "options": [{"letter": "A", "text": "A network interruption corrupted the agent's understanding of the billing service partway through the session.", "correct": false, "explanation": "A network interruption would more likely cause a dropped connection or failed request, not a subtle drift toward generic language about a well-documented finding."}, {"letter": "B", "text": "The agent lacked permission to read the billing service's files, so it fabricated a plausible-sounding answer instead.", "correct": false, "explanation": "A permissions issue would typically produce an explicit error or refusal, not a confident but generic answer that contradicts an earlier specific finding."}, {"letter": "C", "text": "The model experienced a technical malfunction that can only be resolved by restarting the Claude Code application entirely.", "correct": false, "explanation": "This is not a technical malfunction requiring a restart; it is a predictable consequence of long-session context pressure, addressed through context management rather than restarting."}, {"letter": "D", "text": "The finding degraded out of active context over the long session; a scratchpad record would have kept the answer grounded.", "correct": true, "explanation": "This is the characteristic pattern of context degradation: a specific fact fades and is replaced by a generic 'typical pattern' answer. Persisting it to a scratchpad prevents this drift."}], "correct": "D", "select": 1, "group": "G"}, {"id": "f5-043", "domain": 5, "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "situation": "A returns-processing agent calls an order-lookup tool that returns a JSON payload with 40+ fields (shipping carrier metadata, internal warehouse codes, marketing tags, etc.) for every order it checks, and after a dozen lookups the raw payloads dominate the context window even though only 5 fields (order status, purchase date, item, amount, return-window deadline) matter for return eligibility.", "question": "What should the agent do before adding each lookup result to context?", "options": [{"letter": "A", "text": "Truncate each payload to a fixed character length, regardless of which specific fields fall inside that cutoff limit", "correct": false, "explanation": "A fixed character cutoff can just as easily cut off the return-window deadline as it can an irrelevant warehouse code, since it ignores field relevance entirely."}, {"letter": "B", "text": "Run every raw payload through a separate summarization call to shrink it before it is added to context", "correct": false, "explanation": "An extra summarization call adds cost and latency and risks paraphrasing exact fields like the deadline or amount imprecisely, which are exactly the values that must stay exact."}, {"letter": "C", "text": "Cache the full raw JSON payload from each lookup so it can be reused later without recalling the tool again", "correct": false, "explanation": "Caching the full payload preserves the same disproportionate token cost every time it is reused; it never reduces what occupies context."}, {"letter": "D", "text": "Keep only the order status, purchase date, item, amount, and return-window deadline before the result enters context", "correct": true, "explanation": "— extracting only the fields relevant to the current task keeps context usage proportional to relevance rather than to the tool's total response size."}], "correct": "D", "select": 1, "group": "G"}, {"id": "f5-044", "domain": 5, "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "situation": "An architect needs to understand how a refund flow touches the payments, ledger, and notification services before proposing a redesign. The main agent should stay focused on synthesizing the redesign proposal rather than tracing every function call itself.", "question": "How should the architect structure this investigation?", "options": [{"letter": "A", "text": "Ask the main agent to guess at the refund flow's dependencies based on similar flows it recalls from other codebases.", "correct": false, "explanation": "Guessing based on other codebases risks the same generic, unverified answers that context degradation produces; the flow must be traced in this specific codebase."}, {"letter": "B", "text": "Delegate a subagent with the bounded question of tracing refund flow dependencies across the three services and report a distilled summary.", "correct": true, "explanation": "A bounded, specific delegation lets the subagent do the verbose tracing work in its own context and return only a distilled summary, keeping the main agent free to coordinate the redesign."}, {"letter": "C", "text": "Have the main agent open every file in the three services one at a time and keep each file's full contents in the conversation.", "correct": false, "explanation": "Reading every file's full contents into the main agent's context reintroduces the exact problem delegation is meant to solve, consuming space needed for the redesign work."}, {"letter": "D", "text": "Spawn a subagent with no specific instructions at all and simply let it independently decide which parts of the codebase are relevant to refunds.", "correct": false, "explanation": "An unscoped subagent without a specific question tends to explore broadly and return diffuse output, undermining the benefit of isolating a focused investigation."}], "correct": "B", "select": 1, "group": "G"}, {"id": "f5-045", "domain": 5, "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "situation": "An architect needs to quickly locate the single function that computes a discount rate in a small module before making one edit. Spawning a subagent for this lookup would add coordination overhead disproportionate to the task.", "question": "What should the architect do instead?", "options": [{"letter": "A", "text": "Disable all direct tool use in the main agent so that every lookup, however small, must go through delegation.", "correct": false, "explanation": "This would force even trivial lookups through subagent overhead and is not recommended. Anthropic agents can use tools directly when operating within permissions and isolation; subagents are not a required mediation layer for all tool use."}, {"letter": "B", "text": "Always spawn a subagent regardless of task size, since delegation is universally preferable for any exploration step.", "correct": false, "explanation": "This overstates the guidance. Subagents are useful for parallelizable work, context isolation, and specialized instructions, but they add coordination overhead. For a single, small lookup, spawning a subagent is disproportionate."}, {"letter": "C", "text": "Spawn several subagents in parallel to search for the same function redundantly, to cross-check each other's answers.", "correct": false, "explanation": "Parallel subagents are intended for independent subtasks, not duplicate attempts at the same lookup. Redundant subagents increase cost, latency, and coordination without providing a documented cross-check benefit in this context."}, {"letter": "D", "text": "Have the main agent search and read the relevant file directly within its existing permitted and isolated context, because the lookup is small and targeted enough that subagent coordination overhead is not warranted.", "correct": true, "explanation": "Direct main-agent tool use is appropriate for a quick, targeted file lookup. Anthropic's multi-agent guidance supports delegating to subagents for parallelizable, isolated, or specialized work, but not for every small step. Even so, security guidance still requires the main agent to operate under explicit permissions and isolation; the task being small does not waive isolation."}], "correct": "D", "select": 1, "group": "G"}, {"id": "f5-046", "domain": 5, "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "situation": "An airline support agent searches for a passenger by name and flight route to check a baggage claim. The lookup tool returns three passenger records with the same name on that route, each with a different booking reference.", "question": "What should the agent do?", "options": [{"letter": "A", "text": "Select the record with the most complete profile information, since it suggests an established customer", "correct": false, "explanation": "- profile completeness has no bearing on which record actually belongs to the customer being helped."}, {"letter": "B", "text": "Select the record with the most recently booked flight date, since it is the most likely match for the claim", "correct": false, "explanation": "- picking the most recent flight date is a heuristic guess that could easily point to the wrong passenger record."}, {"letter": "C", "text": "Ask the customer for their booking confirmation number or another identifier to pinpoint the correct record", "correct": true, "explanation": "- when multiple records match, the agent should request an additional identifier from the customer rather than heuristically guessing which record applies."}, {"letter": "D", "text": "Merge the relevant details from all three records to construct a single response to the baggage claim", "correct": false, "explanation": "- merging details from multiple unverified records risks combining information from the wrong customer's account."}], "correct": "C", "select": 1, "group": "C"}, {"id": "f5-047", "domain": 5, "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "situation": "A clinic intake assistant summarizes a long patient conversation every few exchanges. After several rounds, the running summary says \"patient has some allergies and wants an appointment soon,\" even though the original messages specified a penicillin allergy and a requested appointment date of the 22nd.", "question": "What should the assistant do differently?", "options": [{"letter": "A", "text": "Maintain a persistent facts block listing the specific allergy and requested date, included in every prompt outside the summary", "correct": true, "explanation": "— a persistent facts block preserves the exact allergy and date outside the summarized narrative, so they survive regardless of how often the rest of the conversation is condensed."}, {"letter": "B", "text": "End every single conversation by asking the patient to confirm their allergy and requested appointment date one final time", "correct": false, "explanation": "A final confirmation catches errors only at the very end and does not protect these details from being misused earlier in the same conversation."}, {"letter": "C", "text": "Store the full unsummarized transcript for later manual review by clinic staff instead of using any summary", "correct": false, "explanation": "Keeping the full transcript for manual review does not help the assistant itself use the correct details during the live conversation."}, {"letter": "D", "text": "Summarize the conversation less frequently so the allergy and date have more turns before being condensed", "correct": false, "explanation": "Reducing summarization frequency delays the loss but does not prevent the specific allergy and date from eventually being generalized away."}], "correct": "A", "select": 1, "group": "G"}, {"id": "f5-048", "domain": 5, "task_id": "5.5", "objective": "Design human review workflows and confidence calibration", "situation": "A logistics company's bill-of-lading extraction pipeline shows 94% field accuracy in aggregate. A new architect discovers that the 'weight' field is correct only 70% of the time specifically when the source document's units are ambiguous (for example, a number with no unit label present). All other conditions for the weight field exceed 95%.", "question": "What review policy should be applied to the weight field going forward?", "options": [{"letter": "A", "text": "Remove the weight field from automated extraction entirely and require full manual entry for every bill of lading, regardless of whether units are specified.", "correct": false, "explanation": "Removing the field from automation entirely discards the fact that unambiguous cases already exceed 95% accuracy, which is an overcorrection that wastes reviewer capacity on cases that don't need it."}, {"letter": "B", "text": "Increase the model's temperature setting when extracting the weight field so it generates a wider range of candidate values for reviewers to choose from.", "correct": false, "explanation": "Raising temperature increases output variability rather than addressing the root cause, which is that the source document itself lacks the information needed to resolve the unit ambiguity; this would not improve accuracy or reviewer usefulness."}, {"letter": "C", "text": "Leave the weight field's review policy unchanged for all documents, since the field's blended 94% aggregate accuracy across the pipeline already meets the general accuracy bar.", "correct": false, "explanation": "Citing the blended aggregate ignores the documented sub-condition where accuracy drops to 70%, which is well below an acceptable bar for unreviewed automation; leaving the policy unchanged would let that specific failure mode continue unflagged."}, {"letter": "D", "text": "Route the weight field to human review whenever the source document does not clearly specify units, while allowing high-confidence, unambiguous cases to bypass review.", "correct": true, "explanation": "The failure is specifically tied to source ambiguity (missing units), so the appropriate policy routes exactly those ambiguous cases to human review while letting unambiguous, well-performing cases continue with reduced review, matching review effort to the identified risk."}], "correct": "D", "select": 1, "group": "C"}, {"id": "f5-049", "domain": 5, "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "situation": "An orchestrator dispatches three research subagents to gather competitor pricing information; each returns only a plain-text paragraph of conclusions with no indication of which source or date the pricing came from. When two subagents report conflicting prices, the orchestrator cannot tell which is current.", "question": "What change to the subagent output requirements would prevent this?", "options": [{"letter": "A", "text": "Require each subagent to assign a confidence score to its conclusion without citing where the information came from", "correct": false, "explanation": "A confidence score indicates certainty but does not provide source location or retrieval date. An outdated price given with high confidence is still outdated, so the orchestrator cannot judge recency."}, {"letter": "B", "text": "Require each subagent to attach metadata such as source location and retrieval date in a structured format (e.g., JSON) alongside every reported fact", "correct": true, "explanation": "This directly addresses the root cause by giving the orchestrator the source and retrieval date. Using a structured format ensures the metadata is reliably machine-parseable, preventing the ambiguity that arises from free-text descriptions. Anthropic strongly recommends structured outputs for critical data fields in agentic workflows to guarantee consistent formatting, as opposed to relying on plain-text instructions which can be inconsistently followed."}, {"letter": "C", "text": "Require the orchestrator to average together the conflicting prices reported by the three subagents", "correct": false, "explanation": "Averaging conflicting prices does not reveal which price is actually the most current. It simply computes a mathematical mean that may be inaccurate and fails to address the underlying metadata deficiency."}, {"letter": "D", "text": "Require each subagent to write a longer, more detailed paragraph explaining its reasoning process for each price found", "correct": false, "explanation": "A longer reasoning paragraph might explain how a price was derived but still lacks explicit source or date metadata. Without that, the orchestrator cannot determine which reported price is more current, so the conflict remains."}], "correct": "B", "select": 1, "group": "G"}, {"id": "f5-050", "domain": 5, "task_id": "5.5", "objective": "Design human review workflows and confidence calibration", "situation": "A team building a resume-parsing pipeline wants Claude to output a confidence score for each extracted field so reviewers can prioritize their limited time.", "question": "During prompt design, which approach best supports later calibration of these scores against a labeled validation set?", "options": [{"letter": "A", "text": "Instruct the model to output a confidence score only for fields it judges to be difficult to extract, omitting scores for fields it judges to be straightforward.", "correct": false, "explanation": "Omitting scores for fields judged 'easy' removes exactly the data needed to verify that assumption; some of those omitted fields could still have meaningfully lower accuracy that only a labeled validation set would reveal."}, {"letter": "B", "text": "Instruct the model to output a per-field confidence score alongside each extracted value, using a consistent numeric scale across all fields and documents.", "correct": true, "explanation": "Field-level scores on a consistent scale allow the team to compare each field's stated confidence against its actual observed correctness rate in a labeled validation set, which is exactly what's needed to calibrate thresholds and route review attention per field."}, {"letter": "C", "text": "Instruct the model to output a single overall document-level confidence score that summarizes its certainty about the entire resume at once.", "correct": false, "explanation": "A single document-level score cannot be calibrated against or used to route individual fields; a resume with one bad field and nine good ones would get one blended score, defeating the purpose of field-level review routing."}, {"letter": "D", "text": "Instruct the model to output a textual explanation of its reasoning for each field instead of a numeric score, since qualitative reasoning is easier for reviewers to act on.", "correct": false, "explanation": "A free-text explanation without a numeric or ordinal score cannot be directly compared against a labeled validation set to compute a calibration curve or set a quantitative routing threshold."}], "correct": "B", "select": 1, "group": "C"}, {"id": "f5-051", "domain": 5, "task_id": "5.3", "objective": "Implement error propagation strategies across multi-agent systems", "situation": "An incident-response coordinator dispatches subagents to pull logs from four services.", "question": "The subagent for the payments service reports a structured failure: 'access_failure, attempted last-15-minutes log pull, partial results: 3 of 4 pods returned data, alternative: retry against read replica.' How should the coordinator most effectively use this information?", "options": [{"letter": "A", "text": "Add the three pods' data to the timeline now, and separately decide whether to retry the missing pod via replica", "correct": true, "explanation": "Best practice in incident response is to immediately document all available facts in the timeline, even if partial, to maintain an accurate record and aid diagnosis. The coordinator can then independently assess whether to attempt the retry against the read replica or rely on Kubernetes’ self-healing (e.g., a Deployment controller replacing the missing pod), maximizing the value of the structured failure report."}, {"letter": "B", "text": "Treat the structured message as equivalent to a plain 'error' status and abandon the payments-service investigation entirely", "correct": false, "explanation": "A structured failure with partial results and a specific alternative offers actionable information. Reducing it to a plain error and abandoning the investigation ignores the recovered data and the possibility of obtaining the missing logs, which could provide critical insights."}, {"letter": "C", "text": "Discard the entire payments-service contribution until all four pods can be pulled in one atomic retry", "correct": false, "explanation": "Discarding valid partial data delays incident diagnosis and wastes information already gathered. Incident timelines should include all factual data as it arrives; waiting for a fully atomic retry is unnecessary and may be impossible if the missing pod is transiently unavailable."}, {"letter": "D", "text": "Ignore the suggested read-replica alternative, since coordinators should never act on subagent suggestions", "correct": false, "explanation": "Coordinators should evaluate subagent suggestions as part of their decision-making. Ignoring a valid suggestion like retrying via a read replica may cause the loss of an opportunity to fill the data gap; the coordinator retains final authority but should use all inputs."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f5-052", "domain": 5, "task_id": "5.3", "objective": "Implement error propagation strategies across multi-agent systems", "situation": "A competitive-intelligence coordinator asks three subagents to gather pricing pages from three competitor websites. One competitor's site returns a 403 because the subagent's fetch was blocked by anti-bot protection after two local retry attempts. The other two subagents succeed.", "question": "When the coordinator synthesizes the final comparison report, what should it do about the blocked competitor?", "options": [{"letter": "A", "text": "Include the two successful comparisons and annotate that the third competitor's pricing is unavailable", "correct": true, "explanation": "Annotating the gap keeps the report honest about its coverage while still delivering the value of the two successful comparisons, consistent with structured, transparent error propagation."}, {"letter": "B", "text": "Present a full three-way comparison without any note, since two of three competitors is close enough", "correct": false, "explanation": "Presenting the comparison as complete without noting the gap misrepresents the report's actual coverage to the reader."}, {"letter": "C", "text": "Withhold the entire report until the anti-bot block can somehow be resolved, however long that may take", "correct": false, "explanation": "Withholding the whole report over one blocked source discards two valid, successful results and may delay useful intelligence indefinitely."}, {"letter": "D", "text": "Present a full three-way comparison, filling the blocked competitor's pricing with a market-trend estimate", "correct": false, "explanation": "Substituting an estimate for real data and presenting it as if it were fetched pricing risks the reader treating a guess as verified fact."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f5-053", "domain": 5, "task_id": "5.5", "objective": "Design human review workflows and confidence calibration", "situation": "A tax-form processing team wants to determine whether it's safe to reduce human review specifically for the 'filing status' field, which currently sits at 99.2% aggregate accuracy.", "question": "Before making that decision, what additional analysis should the team perform to ensure this specific reduction is safe?", "options": [{"letter": "A", "text": "Compare the filing-status field's 99.2% aggregate accuracy against the average accuracy of all other fields on the form, such as by computing the overall mean accuracy, and reduce review if it falls above the median of all other fields.", "correct": false, "explanation": "Comparing the field's accuracy against the average or median of other fields does not expose hidden failure modes within the filing-status field itself. A field can outperform peers overall while still failing on critical edge cases like amended returns."}, {"letter": "B", "text": "Survey reviewers about their confidence in reviewing the filing-status field across different document types and edge cases (such as amended returns or joint filings), and reduce review only if a majority report low concern about that field.", "correct": false, "explanation": "Reviewer confidence surveys capture subjective sentiment, not objective extraction performance. A majority reporting low concern does not quantify whether edge cases like joint filings with differing names are actually extracted correctly, so the analysis would not reliably confirm safety."}, {"letter": "C", "text": "Confirm that the filing-status field has the lowest average character length of any field on the form, by inspecting a sample of completed forms, since shorter fields are generally understood to be easier for models to extract correctly.", "correct": false, "explanation": "Character length is not a validated proxy for extraction accuracy on edge-case conditions, so confirming it is the shortest field does not ensure safety. Even short fields can be mis-extracted in complex scenarios, and visual inspection of a sample cannot substitute for quantitative accuracy analysis."}, {"letter": "D", "text": "Break down the filing-status field's accuracy further by document type and by any known edge-case conditions (such as amended returns or joint filings with differing last names) to confirm no sub-segment falls well below the aggregate.", "correct": true, "explanation": "Breaking down the field's accuracy by document type and edge-case conditions reveals whether any sub-segment performs poorly despite the high aggregate. Without this finer analysis, a specific, risky segment could escape detection, making a reduction unsafe."}], "correct": "D", "select": 1, "group": "C"}, {"id": "f5-054", "domain": 5, "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "situation": "A retail customer asks a support agent to match a lower price they found on a competitor's website. The store's documented policy only describes price adjustments when the store's own website lowers a price within 14 days of purchase; it does not mention competitor pricing at all.", "question": "How should the agent proceed?", "options": [{"letter": "A", "text": "Decline the request, since the policy's own-site provision implies competitor price matches are not permitted", "correct": false, "explanation": "- treating silence as an implicit denial fabricates a policy position that was never actually stated."}, {"letter": "B", "text": "Approve the competitor price match by analogy to the store's own-site adjustment provision instead", "correct": false, "explanation": "- approving by analogy invents an extension of policy the agent isn't authorized to make."}, {"letter": "C", "text": "Ask the customer to submit the competitor's listing as proof before independently approving the match", "correct": false, "explanation": "- collecting proof and then independently approving still involves making a policy decision the agent isn't authorized to make on an unaddressed scenario."}, {"letter": "D", "text": "Escalate the request, since the documented policy is silent on competitor price matching entirely", "correct": true, "explanation": "- the policy only addresses own-site adjustments and is silent on competitor pricing, so this gap should be escalated rather than resolved by inference."}], "correct": "D", "select": 1, "group": "C"}, {"id": "f5-055", "domain": 5, "task_id": "5.5", "objective": "Design human review workflows and confidence calibration", "situation": "A medical-records extraction system reports 96% overall field accuracy. When an architect breaks the results down further, the 'medication dosage' field is only 81% accurate on handwritten prescription forms, while every other field and document type exceeds 97%. The team is deciding whether to reduce human review of the pipeline overall.", "question": "What is the correct action?", "options": [{"letter": "A", "text": "Keep human review at the current level for all fields and document types until the medication-dosage field's accuracy on handwritten forms is separately investigated and improved.", "correct": true, "explanation": "Anthropic's official guidance and HIPAA-compliant deployments require human-in-the-loop for clinical data. The medication-dosage field on handwritten forms is a high-risk area where accuracy must be near-perfect. Pausing any reduction in review until that specific field is investigated and improved aligns with Anthropic's emphasis on safety, clinician sign-off, and targeted validation of weak spots."}, {"letter": "B", "text": "Reduce human review for every field and document type except handwritten prescriptions in general, treating the entire document type as unreliable rather than isolating the specific field.", "correct": false, "explanation": "While handwritten forms are challenging (OCR accuracy 94-97%), the issue is field-specific (medication dosage at 81%) not the entire document type. Isolating and improving the problematic field is more precise and maintains necessary review. Anthropic recommends targeted validation, not broad exclusions, to ensure safety and efficiency."}, {"letter": "C", "text": "Remove human review only from the medication-dosage field on handwritten forms, since that field's absolute accuracy is still above chance level and the errors are likely evenly distributed.", "correct": false, "explanation": "Removing human review from the lowest-accuracy, highest-risk field directly contradicts Anthropic's safety posture. Medication dosage errors can have life-threatening consequences, and accuracy must be 'effectively zero' error, not just above chance. Human oversight is mandatory for such critical extractions."}, {"letter": "D", "text": "Reduce human review across the entire pipeline uniformly, since the 96% overall figure already reflects the presence of the weaker medication-dosage field.", "correct": false, "explanation": "Anthropic emphasizes that for high-risk healthcare use cases, a qualified professional must review AI-generated outputs. A 96% overall accuracy masks a critical 81% accuracy on medication dosage from handwritten prescriptions. Medication errors have an acceptable rate of 'effectively zero,' so reducing review without addressing this weak spot violates responsible AI principles."}], "correct": "A", "select": 1, "group": "C"}, {"id": "f5-056", "domain": 5, "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "situation": "A team is redesigning a multi-day customer-support assistant that currently relies only on rolling summarization of the transcript and unfiltered tool outputs, and has been losing exact figures, mixing up simultaneous issues, and missing details from lengthy middle sections of aggregated reports.", "question": "Which combination of changes would most directly address all of these failure modes together?", "options": [{"letter": "A", "text": "Switch to the model with the largest context window and stop summarizing at all, while keeping the current tool-output and aggregation structure unchanged", "correct": false, "explanation": "A larger context window may delay when limits are hit but does not fix unfiltered tool-output bloat, blended multi-issue narratives, or the placement of key findings within a document."}, {"letter": "B", "text": "Summarize the transcript more frequently, keep all raw tool output for completeness, and add a closing reminder to double-check earlier sections", "correct": false, "explanation": "More frequent summarization still compresses figures each time, keeping all raw tool output reintroduces the token-bloat problem, and a closing reminder does not restructure content to fix position effects."}, {"letter": "C", "text": "Maintain a case-facts block, keep separate records per active issue, trim tool outputs to relevant fields, and lead aggregated input with headed key findings", "correct": true, "explanation": "— each element (persistent case facts, per-issue structured records, trimmed tool outputs, and up-front findings with headers) targets one of the specific failure modes described, and together they address numeric loss, issue mixing, token bloat, and middle-of-document omissions."}, {"letter": "D", "text": "Ask customers to restate key details periodically, forward full subagent reasoning chains unchanged, and rely on default handling of long documents", "correct": false, "explanation": "Relying on customers to restate details burdens users rather than fixing the system, and forwarding unfiltered subagent reasoning reintroduces the exact token-budget problem structured outputs are meant to solve."}], "correct": "C", "select": 1, "group": "G"}, {"id": "f5-057", "domain": 5, "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "situation": "A brokerage's support workflow includes a configured escalation rule: honor a customer's explicit request for a human agent immediately, without attempting to resolve the underlying issue first. A customer messages support: 'I want a real person, not a bot,' regarding a routine request to reset their account password. The agent has not yet attempted any troubleshooting.", "question": "Per the configured escalation rule, what is the appropriate response?", "options": [{"letter": "A", "text": "Ask the customer to explain why a human agent is preferred before deciding how to proceed", "correct": false, "explanation": "- asking the customer to justify the request before proceeding delays applying a rule that already specifies an immediate response to an explicit request."}, {"letter": "B", "text": "Escalate to a human agent right away, honoring the request without first attempting to resolve it", "correct": true, "explanation": "- the configured escalation rule requires honoring an explicit request for a human agent immediately, regardless of how routine the underlying issue is."}, {"letter": "C", "text": "Walk the customer through the password reset steps first, since the process is quick and routine", "correct": false, "explanation": "- resolving the issue first violates the configured rule, which makes no exception for issues that are quick or routine."}, {"letter": "D", "text": "Offer to reset the password immediately and escalate only if the customer repeats the request afterward", "correct": false, "explanation": "- the 'offer first, escalate only if repeated' pattern does not match a rule that triggers on the customer's first explicit request, not on repetition."}], "correct": "B", "select": 1, "group": "C"}, {"id": "f5-058", "domain": 5, "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "situation": "An architect wants to free up conversation context space after a long debugging detour, but also wants to make sure specific facts the agent uncovered, such as an exact configuration value, are not lost to summarization.", "question": "Which combination of practices best achieves both goals?", "options": [{"letter": "A", "text": "Write the exact facts to a durable scratchpad file first, then compact the conversation to reclaim space afterward.", "correct": true, "explanation": "Because compaction condenses the verbatim conversation into a summary, exact details are best preserved by writing them to a durable scratchpad file beforehand, then compacting safely."}, {"letter": "B", "text": "Compact the conversation immediately without recording anything, since compaction is guaranteed to preserve every exact fact.", "correct": false, "explanation": "Compaction condenses the conversation into a summary and is not guaranteed to retain every exact detail, so compacting without first recording facts risks losing precision."}, {"letter": "C", "text": "Delete the conversation and restart from an empty session, since that is the only way to guarantee the facts remain available.", "correct": false, "explanation": "Deleting the conversation and restarting discards the facts entirely rather than preserving them, which is the opposite of the goal."}, {"letter": "D", "text": "Avoid compacting entirely and let the session continue to accumulate context indefinitely for the rest of the exploration.", "correct": false, "explanation": "Never compacting eventually leads to an overfull context window that degrades performance or forces an uncontrolled truncation at an inconvenient point."}], "correct": "A", "select": 1, "group": "G"}, {"id": "f5-059", "domain": 5, "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "situation": "A financial-advisory assistant lets users resume a previous session using a session ID. A developer implements resumption by sending only the user's new message along with the session ID, assuming the API will look up the prior conversation automatically. Users report the assistant \"forgets\" earlier discussed risk tolerance and investment goals when they resume.", "question": "What is the most direct explanation and fix?", "options": [{"letter": "A", "text": "The developer should shorten new messages so there is more room for the model to recall earlier details on its own", "correct": false, "explanation": "Shortening new messages does not cause the model to recall content that was never supplied to it."}, {"letter": "B", "text": "The model should be swapped for one with a larger context window so it can recall the earlier session automatically", "correct": false, "explanation": "A larger context window does not help if the earlier turns are never included in the request in the first place."}, {"letter": "C", "text": "The message history from the earlier session must be resolved into the request; a session ID alone does not supply prior turns to the model", "correct": true, "explanation": "— without the actual prior turns present in the request, directly or via a mechanism that reconstructs them, the model has no access to earlier risk tolerance or goals, since a session ID by itself carries no conversational content."}, {"letter": "D", "text": "The assistant should ask users to restate their risk tolerance and goals immediately after resuming, since this is unavoidable", "correct": false, "explanation": "Repeated restatement is a workaround for a solvable implementation gap, not a necessary limitation of the system."}], "correct": "C", "select": 1, "group": "G"}, {"id": "f5-060", "domain": 5, "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "situation": "A main agent is coordinating exploration of a repository with over 50,000 files and needs a complete list of every test file before planning a coverage audit. Running the search directly would flood the main conversation with thousands of matched paths it will not need again.", "question": "What best keeps the main agent's context focused on high-level coordination while still producing the list?", "options": [{"letter": "A", "text": "Ask the user to manually paste the list of test file paths into the conversation before continuing with the audit.", "correct": false, "explanation": "Manually collecting the list defeats the purpose of using the agent for exploration and adds unnecessary user burden for a task the agent can perform directly."}, {"letter": "B", "text": "Read every file in the repository sequentially in the main agent to identify which ones are tests by inspecting their contents.", "correct": false, "explanation": "Reading every file's contents sequentially in the main agent is far more expensive than a targeted search and consumes even more context than the raw search output would."}, {"letter": "C", "text": "Run the search directly in the main agent and keep the entire raw output in the conversation in case it needs to be referenced again later.", "correct": false, "explanation": "Keeping the full raw search output in the main conversation is exactly the pattern that floods the coordinating agent's context with data it will not use again."}, {"letter": "D", "text": "Spawn a subagent to search for test files and return only the consolidated list, leaving the raw output isolated in its own context.", "correct": true, "explanation": "Delegating the search isolates the verbose discovery output in the subagent's own context window; the main agent receives only the distilled list and keeps its context free for coordination."}], "correct": "D", "select": 1, "group": "G"}, {"id": "f5-061", "domain": 5, "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "situation": "An architect is designing the output schema for research subagents that will feed a downstream synthesis agent producing a due-diligence report. The architect wants downstream agents to be able to verify and re-attribute every claim without re-reading the original documents.", "question": "What should each subagent's structured output include for every extracted claim?", "options": [{"letter": "A", "text": "The claim text and a short paraphrase of the surrounding paragraph, without naming the specific source document", "correct": false, "explanation": "A paraphrase without the source name loses the ability to trace the claim back to a specific document, which defeats the purpose of preserving provenance."}, {"letter": "B", "text": "The claim text, the source URL or document name it came from, and a relevant excerpt supporting the claim", "correct": true, "explanation": "Pairing each claim with its source identifier and a supporting excerpt gives downstream agents everything needed to verify and re-attribute the claim without returning to the original documents."}, {"letter": "C", "text": "The claim text and the name of the subagent that produced it, so the coordinator knows which subagent to re-query if needed", "correct": false, "explanation": "Knowing which subagent produced a claim does not identify the underlying source document, so verification still requires re-querying rather than reading the excerpt directly."}, {"letter": "D", "text": "The claim text and a numeric confidence score assigned by the subagent based on its own judgment of source reliability", "correct": false, "explanation": "A self-assigned confidence score reflects the subagent's opinion but does not tell downstream agents which document or excerpt the claim actually came from."}], "correct": "B", "select": 1, "group": "B"}, {"id": "f5-062", "domain": 5, "task_id": "5.3", "objective": "Implement error propagation strategies across multi-agent systems", "situation": "A coordinator receives this message from a subagent: 'Query failed.' No other detail is provided. The coordinator must decide whether to retry the subagent, try an alternative subagent, or give up on that portion of the task.", "question": "What is the core limitation of this message for making that decision?", "options": [{"letter": "A", "text": "It is too long for the coordinator to process efficiently within its available context window during a multi-agent run", "correct": false, "explanation": "The message is extremely short, not long, so length is not the limitation; the problem is missing content, not excess content."}, {"letter": "B", "text": "It is missing a timestamp, and timestamps alone let a coordinator choose between retrying and escalating", "correct": false, "explanation": "A timestamp tells the coordinator when the failure happened but nothing about why, so it cannot substitute for the missing failure-type and context information."}, {"letter": "C", "text": "It fails to include the exact HTTP status code, which is the only detail coordinators ever really require", "correct": false, "explanation": "An HTTP status code alone, while useful, is not 'the only detail' needed; partial results and alternatives matter just as much for a full recovery decision."}, {"letter": "D", "text": "It omits the failure type and any partial results or alternatives, leaving no basis to choose a recovery path", "correct": true, "explanation": "Without failure type, partial results, or alternatives, the coordinator is reduced to guessing whether a retry might help, which defeats the purpose of structured error propagation."}], "correct": "D", "select": 1, "group": "A"}, {"id": "f5-063", "domain": 5, "task_id": "5.3", "objective": "Implement error propagation strategies across multi-agent systems", "situation": "In a research pipeline, a coordinator dispatches five subagents to gather sources on different aspects of a market analysis. One subagent's web-fetch tool raises an unrecoverable connection error. The coordinator's current implementation aborts the entire pipeline and discards the four completed subagent results.", "question": "What is the better design?", "options": [{"letter": "A", "text": "Still abort the pipeline entirely, since a single subagent failure means the overall research quality cannot be trusted", "correct": false, "explanation": "Aborting on a single subagent failure is the anti-pattern being illustrated, since it throws away four subagents' worth of valid, already-completed work."}, {"letter": "B", "text": "Replace the failed subagent's section with fabricated but plausible content so the report shows no visible gaps", "correct": false, "explanation": "Fabricating content to hide a coverage gap is worse than aborting, because it presents unsupported information as if it were verified research."}, {"letter": "C", "text": "Synthesize from the four completed results and use the failed subagent's error context to flag the coverage gap", "correct": true, "explanation": "Terminating the whole workflow on one subagent failure discards useful completed work; the coordinator should use the partial results it has and rely on the failed subagent's error context to annotate what's missing."}, {"letter": "D", "text": "Re-run all five subagents from scratch, including the four that already succeeded, to keep the batch consistent", "correct": false, "explanation": "Re-running the four successful subagents wastes time and cost for no benefit, since their results are already valid and do not need to be regenerated."}], "correct": "C", "select": 1, "group": "A"}, {"id": "f5-064", "domain": 5, "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "situation": "An architect is about to spawn a new subagent to investigate the caching layer, but suspects a similar investigation may already have been done earlier in the exploration.", "question": "What should the architect check first to avoid redundant exploration?", "options": [{"letter": "A", "text": "Nothing; spawn the new subagent immediately regardless of what earlier phases may have already found.", "correct": false, "explanation": "Spawning without checking risks duplicating an investigation that has already been completed, wasting time and tokens."}, {"letter": "B", "text": "The user's personal notes taken outside the Claude Code session, since those are the only reliable record of prior work.", "correct": false, "explanation": "Personal notes outside the session are not integrated into the agent workflow; the scratchpad and manifest records are the mechanism designed for this purpose."}, {"letter": "C", "text": "The main agent's unaided memory of the conversation from several hours ago, trusting it to recall every earlier finding.", "correct": false, "explanation": "Relying on the main agent's unaided recall from hours earlier is exactly the kind of context degradation the scratchpad and manifest are meant to counteract."}, {"letter": "D", "text": "The manifest or scratchpad records from earlier phases, to see whether the caching layer was already investigated.", "correct": true, "explanation": "Checking the manifest or scratchpad records shows what has already been discovered and recorded, avoiding the cost of spawning a subagent to redo captured work."}], "correct": "D", "select": 1, "group": "G"}, {"id": "f5-065", "domain": 5, "task_id": "5.5", "objective": "Design human review workflows and confidence calibration", "situation": "An invoice-extraction pipeline built on Claude reports 97% overall field accuracy across all document types, and the team is preparing to remove human review for any extraction above a fixed confidence threshold. Before doing so, an architect wants to confirm the 97% figure isn't hiding a problem.", "question": "What is the most important check to run first?", "options": [{"letter": "A", "text": "Ask the model to self-report its own estimated accuracy by examining a sample of its outputs, such as a subset of invoices, and compare that self-estimate against the measured 97%.", "correct": false, "explanation": "A model’s self‑estimated accuracy on its own outputs is not a reliable substitute for stratified evaluation against labeled ground truth. Self‑estimates can be uncalibrated and would not expose segment‑level accuracy gaps that the 97% aggregate might conceal."}, {"letter": "B", "text": "Break down accuracy by document type and field to check whether any segment, such as handwritten receipts or a specific vendor's layout, performs far worse than the aggregate.", "correct": true, "explanation": "Breaking down accuracy by document type and field reveals whether the aggregate 97% masks poor performance on specific segments like handwritten receipts or particular vendors. This segmentation is essential to ensure that no minority subset with high error rates slips through automated acceptance."}, {"letter": "C", "text": "Re-run the same evaluation set through the model a second time, using identical configurations, and confirm the overall accuracy score stays within one percentage point of 97%.", "correct": false, "explanation": "Re‑running the same evaluation set under identical configurations only checks measurement stability, not whether the 97% accuracy is evenly distributed across document types and fields. It cannot surface hidden weak spots that would persist in repeated runs."}, {"letter": "D", "text": "Increase the size of the evaluation set by sampling more document types, so the 97% aggregate figure is computed from a larger and therefore more statistically significant sample.", "correct": false, "explanation": "Increasing the evaluation set size yields a more statistically significant aggregate number, but that single blended figure can still obscure poorly performing segments. A larger sample without stratification does not help detect problematic sub-categories."}], "correct": "B", "select": 1, "group": "C"}, {"id": "f5-066", "domain": 5, "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "situation": "An airline customer messages support: 'I was assigned a middle seat and I want to talk to an actual person about it.' The agent responds by offering to move the customer to an aisle seat directly, without addressing the request to speak with a person.", "question": "The customer replies: 'No, I said I want to talk to an actual person about this.' What should the agent do now?", "options": [{"letter": "A", "text": "Escalate the conversation to a human agent now, since the customer has reiterated their preference for one", "correct": true, "explanation": "- the customer's first message already asked to speak with a person, the agent's offer ignored that request, and the customer's reply explicitly restates it. Agent-design guidance treats an explicit, ignored request for a human as a blocker the agent should pause for rather than continue past on its own."}, {"letter": "B", "text": "Ask the customer to clarify why the automated resolution is not acceptable before escalating further", "correct": false, "explanation": "- demanding a justification after the customer has already restated the request for a human adds friction instead of honoring a clearly repeated request."}, {"letter": "C", "text": "Repeat the offer to fix the seat assignment, since the request is simple and within the agent's capability", "correct": false, "explanation": "- repeating an offer the customer already declined ignores the explicit, restated request for a human agent instead of treating it as a blocker requiring escalation."}, {"letter": "D", "text": "Proceed to change the seat assignment and inform the customer of the change without further discussion", "correct": false, "explanation": "- proceeding unilaterally after the customer restated a request for a human disregards the explicit blocker the customer raised."}], "correct": "A", "select": 1, "group": "C"}, {"id": "f5-067", "domain": 5, "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "situation": "A customer contacts a company's Claude-powered support agent about a $45 API overage charge and says: \"I don't want to discuss this with a bot, connect me to a human agent right now.\" The agent has not yet reviewed the account.", "question": "Following recommended escalation design practice for AI support agents, what should the agent do?", "options": [{"letter": "A", "text": "Review the account's usage history and attempt to resolve the overage charge before responding to the customer's escalation request", "correct": false, "explanation": "Reviewing usage history before responding treats an explicit demand for a human as an input to a decision rather than as the escalation trigger it already is, and it leaves the customer's stated request unanswered while the agent investigates. Investigating first is appropriate only when no explicit request for a human has been made."}, {"letter": "B", "text": "Offer to explain the overage charge in detail, attempt to resolve it, and escalate to a human only if the customer repeats the request or the issue remains unresolved", "correct": false, "explanation": "This is the acknowledge-frustration-and-offer-resolution pattern, which the exam guide reserves for a customer who is frustrated or whose issue is straightforward but who has not explicitly demanded a human — there the agent offers help and escalates only if the customer reiterates the preference. That the $45 overage is simple enough to explain does not move this case into that branch: the guide draws the line at whether the customer explicitly demands a human, and once the request is explicit, waiting for the customer to repeat it is exactly the behaviour it rules out."}, {"letter": "C", "text": "Escalate the conversation to a human agent immediately, without first investigating the overage charge", "correct": true, "explanation": "The customer has made an explicit, unambiguous demand for a human (\"connect me to a human agent right now\"), and a customer request for a human is itself an appropriate escalation trigger. Anthropic's Claude Certified Architect – Foundations exam guide states the required skill as honoring explicit customer requests for human agents immediately without first attempting investigation, so the handoff happens now rather than after the agent has looked into the charge."}, {"letter": "D", "text": "Ask the customer to first explain why they don't want to work with an automated system before escalating", "correct": false, "explanation": "No escalation guidance requires a customer to justify preferring a human, and demanding that justification adds an intake step on top of an escalation that has already been triggered. Asking for more input from the customer is reserved for genuinely ambiguous cases, such as a lookup returning multiple matching customer records where additional identifiers are needed."}], "correct": "C", "select": 1, "group": "C"}, {"id": "f5-068", "domain": 5, "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "situation": "A long-running research agent uses server-side context compaction (context_management with beta header compact-2026-01-12) to keep an extended multi-turn investigation within the context window. After several rounds of compaction, the architect notices that citations linking earlier findings to their original source documents have disappeared from the working context, even though the findings themselves survived.", "question": "What is the most likely cause of this gap?", "options": [{"letter": "A", "text": "The compaction step condensed earlier turns without explicitly preserving claim-source mappings alongside the findings", "correct": true, "explanation": "Anthropic's server-side compaction is lossy by design: it summarizes older context into a high-fidelity summary that preserves key facts, decisions, and constraints, but discards redundant or detailed content. If citation/source mappings are not explicitly prioritized as critical facts or structured metadata, they can be condensed away while the substantive finding remains. This matches the documented behavior of compaction as a lossy summarization step rather than a verbatim retention of original text."}, {"letter": "B", "text": "Citations are stored in a separate ephemeral cache that is cleared automatically once a conversation exceeds a fixed number of turns", "correct": false, "explanation": "Citations are not documented as being kept in a separate ephemeral cache with a fixed turn-based clearing rule. The loss is explained by the summarization process itself: detailed, sentence-level citation markers do not survive lossy compaction when they are not explicitly preserved."}, {"letter": "C", "text": "Compaction only operates on tool results and never touches any text the model itself generated, including citations", "correct": false, "explanation": "Server-side compaction is not limited to tool_result content; it summarizes older conversation context, including model-generated text and agent reasoning, while preserving tool_use / tool_result pairing integrity. It can therefore affect citations and other model-generated details as part of compressing the conversation trajectory."}, {"letter": "D", "text": "The model's context window silently shrank between turns, causing the oldest citations to be truncated regardless of compaction", "correct": false, "explanation": "There is no documented behavior of the context window silently shrinking between turns; instead, compaction extends the effective context length by automatically summarizing older content as the token limit is approached. Losing the oldest citations while the findings survive is not explained by silent truncation, but by lossy summarization."}], "correct": "A", "select": 1, "group": "B"}, {"id": "f5-069", "domain": 5, "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "situation": "Phase 1 of a codebase exploration used several subagents to map the authentication module's structure. The architect is now ready to start phase 2, spawning subagents to investigate how other services integrate with authentication.", "question": "What should happen between the two phases to keep phase 2 grounded in phase 1's discoveries?", "options": [{"letter": "A", "text": "Forward the complete raw transcripts of every phase 1 subagent directly into each phase 2 subagent's initial prompt.", "correct": false, "explanation": "Forwarding full raw transcripts reintroduces the verbose output that delegation was meant to isolate, burdening each phase 2 subagent's context unnecessarily."}, {"letter": "B", "text": "Wait until a phase 2 subagent asks a clarifying question before providing any information about the authentication module.", "correct": false, "explanation": "Waiting for a clarifying question is unreliable and delays the investigation; the coordinator should proactively supply relevant context at the start."}, {"letter": "C", "text": "Start phase 2 subagents with no reference to phase 1 and let them rediscover the authentication module's structure independently.", "correct": false, "explanation": "Letting phase 2 subagents rediscover the module wastes the work already done in phase 1 and risks conclusions inconsistent with the earlier findings."}, {"letter": "D", "text": "Synthesize the phase 1 findings into a concise summary and inject that summary into each phase 2 subagent's initial prompt.", "correct": true, "explanation": "Condensing findings into a summary before the next phase preserves the key facts discovered without carrying forward verbose exploration output, grounding phase 2 subagents efficiently."}], "correct": "D", "select": 1, "group": "G"}, {"id": "f5-070", "domain": 5, "task_id": "5.3", "objective": "Implement error propagation strategies across multi-agent systems", "situation": "A code-review coordinator dispatches a security-scanning subagent and a style-checking subagent in parallel. The security subagent's static-analysis tool crashes partway through, after already flagging two vulnerabilities in files it processed before the crash.", "question": "What should the subagent send to the coordinator?", "options": [{"letter": "A", "text": "Nothing at all until the scan can be fully restarted and completed without any interruption whatsoever", "correct": false, "explanation": "Waiting for a full restart delays the coordinator's ability to act on the two vulnerabilities that were already confirmed, with no benefit to accuracy."}, {"letter": "B", "text": "The two vulnerabilities found, the crash's failure type, and which files were left unscanned by the tool", "correct": true, "explanation": "Surfacing both the partial results and the specifics of what wasn't covered lets the coordinator use the findings it has while deciding whether to rescan the remaining files."}, {"letter": "C", "text": "The two vulnerabilities found, reported as the complete final result with no note about unscanned files", "correct": false, "explanation": "Presenting partial results as complete misleads the coordinator into believing the whole codebase was scanned when several files were never checked."}, {"letter": "D", "text": "Only a generic failure notice, since including partial findings could give a false sense of completeness", "correct": false, "explanation": "Withholding real findings the tool already produced discards useful, verified information the coordinator could act on immediately."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f5-071", "domain": 5, "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "situation": "A telecom support agent is helping a customer dispute a charge. An account lookup by last name and area code returns two accounts with similar details, and neither the customer's phone number nor account number has been confirmed yet.", "question": "What should the agent do?", "options": [{"letter": "A", "text": "Proceed with the account that has the more recent billing activity, assuming it reflects the active dispute", "correct": false, "explanation": "Recent billing activity is a heuristic, not an identity check. Two records sharing a last name and area code can both show recent activity, so acting on the more recently billed one risks exposing or altering the wrong customer's billing data. The documented pattern is to request a discriminating identifier instead."}, {"letter": "B", "text": "Ask the customer general questions about their service plan and infer the correct account from the answers", "correct": false, "explanation": "General questions about a service plan produce answers that both similar accounts could plausibly match, so the ambiguity remains unresolved and the agent is still guessing. Resolving it requires a discriminating identifier such as the account number or the PIN on file."}, {"letter": "C", "text": "Ask the customer to provide their account number or the PIN on file to confirm which account is theirs", "correct": true, "explanation": "Requesting an additional identifier is the documented pattern for this situation: when a lookup returns multiple matches, the agent should ask for additional identifiers rather than selecting based on heuristics. An account number or the account PIN on file uniquely distinguishes the two similar records, so the agent can confirm ownership before discussing or adjusting any charge."}, {"letter": "D", "text": "Proceed with the account showing the higher outstanding balance, since disputes typically involve larger charges", "correct": false, "explanation": "An outstanding balance says nothing about which record belongs to the caller, and a dispute can just as easily concern the smaller charge. Choosing the record with the larger balance is exactly the kind of heuristic selection that the escalation and ambiguity guidance rules out when a lookup returns multiple matches."}], "correct": "C", "select": 1, "group": "C"}, {"id": "f5-072", "domain": 5, "task_id": "5.5", "objective": "Design human review workflows and confidence calibration", "situation": "A legal-document review pipeline processes contracts where, in some cases, two clauses on different pages state contradictory terms for the same provision (for example, differing renewal notice periods). The model extracts a single value for the field without flagging the contradiction.", "question": "What review-routing behavior should the team implement for this scenario?", "options": [{"letter": "A", "text": "Have the model detect when source values conflict across the document and route those specific extractions to human review, even if its confidence in the single value it chose is high.", "correct": true, "explanation": "Ambiguous or contradictory source documents are a distinct risk signal from low model confidence, and extractions from such documents should be routed to human review regardless of the model's stated confidence in the single value it produced, since a confident-sounding answer can still be based on an unresolved conflict in the source."}, {"letter": "B", "text": "Trust the model's single extracted value whenever its reported confidence score is above the routing threshold, since the score already accounts for any conflicting source text.", "correct": false, "explanation": "A model's confidence score reflects its certainty in the value it output, not necessarily an accurate signal that the source document itself was unambiguous; a model can be confident while silently picking one of two conflicting values."}, {"letter": "C", "text": "Extract only the value from whichever page appears first in the document, since earlier clauses are conventionally assumed to take precedence in contract structure.", "correct": false, "explanation": "There is no general convention that an earlier clause automatically overrides a later one in contract interpretation; resolving a genuine contradiction between clauses requires human legal judgment, not a positional heuristic."}, {"letter": "D", "text": "Average the two conflicting values from the document to produce a single extracted number that falls between them, then route that averaged value through normal processing.", "correct": false, "explanation": "Averaging two conflicting contractual terms, such as notice periods, produces a legally meaningless number that matches neither clause and does not resolve which term is actually binding; this masks the conflict instead of surfacing it for human resolution."}], "correct": "A", "select": 1, "group": "C"}, {"id": "f5-073", "domain": 5, "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "situation": "An architect is setting conventions for what agents should write into scratchpad files during a large codebase exploration.", "question": "Which content best serves the goal of counteracting context degradation on later questions?", "options": [{"letter": "A", "text": "A single vague note such as 'explored the codebase and things generally look fine,' without further detail.", "correct": false, "explanation": "A vague summary provides nothing specific to anchor later answers and does not prevent the agent from falling back on generic assumptions."}, {"letter": "B", "text": "Concrete facts such as exact class names, file paths, and discovered mechanisms, recorded as they are found.", "correct": true, "explanation": "Specific, concrete findings are exactly what later questions need to avoid drifting into generic 'typical pattern' language, giving the agent something precise to reference."}, {"letter": "C", "text": "Only a timestamp of when each exploration step occurred, without any description of what was found.", "correct": false, "explanation": "Timestamps without any description of findings give the coordinator no usable information about what was actually discovered."}, {"letter": "D", "text": "The complete raw contents of every file the agent opened during the investigation, copied in full into the scratchpad.", "correct": false, "explanation": "Copying full raw file contents defeats the purpose of a scratchpad, reintroducing the same volume of data the agent was trying to avoid keeping in active context."}], "correct": "B", "select": 1, "group": "G"}, {"id": "f5-074", "domain": 5, "task_id": "5.3", "objective": "Implement error propagation strategies across multi-agent systems", "situation": "A translation subagent is asked to localize a product description into a language for which no glossary entry exists in the terminology database. The lookup completes successfully but returns no matching glossary terms, which is expected for a newly supported language.", "question": "How should this be distinguished from a scenario where the terminology database itself is offline?", "options": [{"letter": "A", "text": "Report the missing-glossary case as a valid empty result, and the database-offline case as a retryable access failure", "correct": true, "explanation": "A successful lookup that legitimately finds nothing is different from a lookup that never happened because the database was unreachable; only the latter is worth retrying."}, {"letter": "B", "text": "Skip reporting either case separately, since the coordinator can translate regardless of glossary availability", "correct": false, "explanation": "The two cases have different implications for recovery, so treating them as unimportant to distinguish would prevent the coordinator from retrying the genuinely fixable database outage."}, {"letter": "C", "text": "Report both cases the same way, since in each case the subagent ends up without any glossary terms to work with at all", "correct": false, "explanation": "Reporting both identically hides that the offline database is a fixable problem, while the missing glossary for a new language may simply be an expected, permanent state."}, {"letter": "D", "text": "Report the missing-glossary case as an access failure, since finding no matching terms means the lookup failed", "correct": false, "explanation": "A successful lookup with no matches is not evidence of a broken lookup; conflating the two would cause needless retries against a database that is working correctly."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f5-075", "domain": 5, "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "situation": "A coordinator process orchestrating five exploration subagents crashes after three of them finished and exported their findings to known file locations. The coordinator restarts.", "question": "Following a structured state persistence design, what should it do?", "options": [{"letter": "A", "text": "Discard the three completed agents' exported findings and ask the user to describe what those agents had found from memory.", "correct": false, "explanation": "Discarding exported findings and relying on the user's memory abandons the durable state that was specifically saved to survive exactly this kind of crash."}, {"letter": "B", "text": "Re-run all five agents completely from scratch, discarding the exported findings from the three that already completed before the crash.", "correct": false, "explanation": "Re-running completed agents discards useful work already exported to disk, wasting the time and cost already spent and defeating the purpose of crash recovery."}, {"letter": "C", "text": "Wait indefinitely for the crashed agents to resume on their own without taking any action to reload prior state.", "correct": false, "explanation": "Waiting indefinitely takes no advantage of the state already persisted to disk and leaves the run permanently stalled."}, {"letter": "D", "text": "Load the manifest of agent states, skip re-running the three completed agents, and inject their findings into the remaining prompts.", "correct": true, "explanation": "This is the purpose of the manifest: the coordinator loads what each agent already exported, avoids redoing completed work, and injects recovered state into the agents still to run."}], "correct": "D", "select": 1, "group": "G"}, {"id": "f5-076", "domain": 5, "task_id": "5.5", "objective": "Design human review workflows and confidence calibration", "situation": "A customer-onboarding pipeline extracts identity fields from submitted documents and uses a calibrated confidence threshold to decide which extractions bypass human review. A new document scanner is deployed that produces slightly lower-resolution images than before.", "question": "What is the correct response to this change from a review-workflow design perspective?", "options": [{"letter": "A", "text": "Continue using the existing calibrated threshold unchanged, since confidence thresholds are calibrated against the model's behavior and are independent of image quality or scanning hardware.", "correct": false, "explanation": "Confidence calibration is measured against a specific input distribution; if the input images change (e.g., lower resolution), the previously calibrated relationship between score and actual accuracy is not guaranteed to still hold."}, {"letter": "B", "text": "Treat the new scanner output as a potential shift in the input population and re-validate accuracy and confidence calibration on a sample of documents captured with it before trusting the existing threshold.", "correct": true, "explanation": "A change in scanning hardware can alter the distribution of input document quality, which is exactly the kind of shift that stratified sampling and calibration are meant to catch; the safe approach is to re-validate accuracy and recalibrate the threshold on documents from the new scanner rather than assuming the old calibration still holds."}, {"letter": "C", "text": "Disable the confidence threshold and route all extractions from the new scanner to human review permanently, since any hardware change should be assumed to make automation unsafe going forward.", "correct": false, "explanation": "Permanently routing all extractions to review is an overcorrection that assumes accuracy definitely degraded without measuring it, discarding the efficiency benefits of automation without evidence that they're actually unsafe with the new scanner."}, {"letter": "D", "text": "Automatically lower the confidence threshold by a fixed amount whenever any hardware change is deployed, since lower image quality always reduces model confidence by a predictable margin.", "correct": false, "explanation": "There is no established fixed, predictable relationship between a hardware change and a specific threshold adjustment; applying an arbitrary fixed offset without measurement is not a validated approach."}], "correct": "B", "select": 1, "group": "C"}, {"id": "f5-077", "domain": 5, "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "situation": "A coordinator agent receives structured claim-source mappings from four subagents researching the same topic from different angles. During merging, several claims are near-duplicates reported by multiple subagents with slightly different wording, and the supporting source citations are not identical across the reports.", "question": "What is the best way to merge these without losing provenance?", "options": [{"letter": "A", "text": "Keep only the version of the claim reported by the subagent that produced its output first, discarding the duplicates", "correct": false, "explanation": "This approach arbitrarily prioritizes one subagent and discards source citations contributed by the other subagents. The synthesis agent is expected to preserve and merge claim-source mappings, not delete them based on response order, because doing so loses traceability for the near-duplicate claims."}, {"letter": "B", "text": "Delete all but one occurrence of the claim and drop its source citations, since the claim is now well established", "correct": false, "explanation": "Even if the claim is well established, dropping source citations removes provenance. The recommended approach is to consolidate the near-duplicates but retain the full set of supporting source citations so the merged entry remains auditable and compliant with provenance requirements."}, {"letter": "C", "text": "Rewrite the duplicate claims into a single new sentence that references none of the original subagents' citations", "correct": false, "explanation": "Rewriting without retaining any source citations destroys the original claim-source mappings and makes the provenance untraceable. The correct approach is to consolidate the claim entry while preserving all source attributions that support it."}, {"letter": "D", "text": "Consolidate the duplicates into one entry while retaining the full set of source citations that support it", "correct": true, "explanation": "The CCAR-F exam subdomain 5.6 on Preserve information provenance and handle uncertainty in multi-source synthesis requires maintaining structured claim-source mappings when combining findings. Because the near-duplicates are supported by different source citations, consolidating them into a single claim entry while retaining the full set of distinct supporting citations keeps every attribution traceable and auditable. Dropping any subagent's citation would lose provenance."}], "correct": "D", "select": 1, "group": "B"}, {"id": "f5-078", "domain": 5, "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "situation": "An architect is exploring a 200,000-line codebase to produce a high-level dependency map for a migration plan. Some tasks involve exhaustively searching many directories and generating large volumes of matched output; other tasks involve synthesizing an overall migration strategy from what was found.", "question": "How should responsibility be split between the main agent and subagents?", "options": [{"letter": "A", "text": "Have the main agent perform the exhaustive multi-directory searches itself so it retains every matched result in its own context.", "correct": false, "explanation": "Having the main agent perform the exhaustive searches itself reintroduces the exact volume of output delegation is meant to isolate, crowding out room for synthesis."}, {"letter": "B", "text": "Delegate the exhaustive searches that generate verbose output to subagents, and let the main agent synthesize their summaries.", "correct": true, "explanation": "This matches the intended split: subagents absorb verbose, high-volume exploration output in their own context, while the main agent preserves capacity for high-level synthesis."}, {"letter": "C", "text": "Avoid any delegation and keep the entire investigation, including all search output, within a single agent's context throughout.", "correct": false, "explanation": "Keeping everything in a single agent's context is the pattern most likely to trigger context degradation on a codebase this large."}, {"letter": "D", "text": "Delegate the high-level synthesis of the migration strategy to a subagent while the main agent performs the exhaustive searches.", "correct": false, "explanation": "Synthesis benefits from the coordinator's accumulated high-level view across all findings; delegating it away while the main agent does verbose searching inverts the intended roles."}], "correct": "B", "select": 1, "group": "G"}, {"id": "f5-079", "domain": 5, "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "situation": "A SaaS support agent generates a self-reported confidence score of 45% for its answer to a billing proration question, even though the retrieved documentation and account data clearly and unambiguously support that answer.", "question": "Should the agent escalate based on the low confidence score?", "options": [{"letter": "A", "text": "No, a self-reported score isn't a reliable complexity proxy when the evidence clearly supports the answer", "correct": true, "explanation": "- self-reported confidence scores are unreliable proxies for actual case complexity, and here the supporting evidence is clear, so the case should be resolved directly."}, {"letter": "B", "text": "Yes, low self-reported confidence always indicates the underlying case is factually ambiguous", "correct": false, "explanation": "- a low confidence score does not necessarily reflect factual ambiguity in the underlying case, as shown by the clear supporting evidence here."}, {"letter": "C", "text": "Yes, any confidence score below 50% should automatically trigger escalation to a human agent", "correct": false, "explanation": "- treating any low score as an automatic escalation trigger substitutes an unreliable heuristic for actual evaluation of the case."}, {"letter": "D", "text": "No, but only because billing proration questions are categorically excluded from confidence-based escalation", "correct": false, "explanation": "- the reasoning is wrong; the issue is that confidence scores generally aren't reliable complexity signals, not that this topic has a special exemption."}], "correct": "A", "select": 1, "group": "C"}, {"id": "f5-080", "domain": 5, "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "situation": "A retrieval-augmented pipeline uses Claude's search results feature to ground answers in a custom knowledge base of internal policy documents. The architect wants the final synthesis to retain proper source attribution comparable to what web search citations provide.", "question": "What does the search results feature primarily enable in this scenario?", "options": [{"letter": "A", "text": "A guarantee that every retrieved passage is the single most recent version of the relevant internal policy document", "correct": false, "explanation": "The feature does not verify recency or guarantee that a retrieved passage is the latest version; that determination still depends on the underlying document metadata."}, {"letter": "B", "text": "Automatic reconciliation of any conflicting policy statements found across different internal documents in the knowledge base", "correct": false, "explanation": "The feature provides attribution for retrieved content; it does not itself resolve disagreements between conflicting internal documents."}, {"letter": "C", "text": "Automatic conversion of all retrieved policy text into a single normalized document format before synthesis begins", "correct": false, "explanation": "The feature attributes retrieved passages to their sources rather than normalizing or reformatting the underlying documents."}, {"letter": "D", "text": "Natural citations with proper source attribution for the custom knowledge base, similar in quality to web search citations", "correct": true, "explanation": "Search results are designed to enable natural citations with proper source attribution for custom knowledge bases and tools, achieving citation quality comparable to web search."}], "correct": "D", "select": 1, "group": "B"}, {"id": "f5-081", "domain": 5, "task_id": "5.3", "objective": "Implement error propagation strategies across multi-agent systems", "situation": "A coordinator is choosing between two error-reporting schemas for its subagents. Schema A returns a single string status like 'error' or 'ok'. Schema B returns a structured object with failure type, attempted query, partial results, and suggested alternatives.", "question": "During an incident where three of eight subagents fail for different reasons, which schema lets the coordinator respond most effectively, and why?", "options": [{"letter": "A", "text": "Schema B, since the coordinator can inspect each failure's type and results to decide whether to retry or proceed", "correct": true, "explanation": "With three different failure causes among eight subagents, the coordinator needs per-failure detail to make distinct, appropriate recovery decisions rather than treating every failure the same way."}, {"letter": "B", "text": "Schema B, since structured objects always take priority regardless of what information they actually contain", "correct": false, "explanation": "Structured objects aren't beneficial merely because they're structured; the reasoning must be about the specific recovery-relevant information they carry, not the format itself."}, {"letter": "C", "text": "Schema A, since reducing failures to just 'error' or 'ok' forces subagents to resolve problems before reporting up", "correct": false, "explanation": "A binary status doesn't force local resolution; it only removes the coordinator's visibility into what happened, which is the opposite of what's needed during an incident with varied causes."}, {"letter": "D", "text": "Schema A, since a simple string status is easier for the coordinator to parse quickly during a live incident", "correct": false, "explanation": "Ease of parsing a bare string doesn't help the coordinator decide what to actually do about each of the three distinct failures, since the string carries no differentiating detail."}], "correct": "A", "select": 1, "group": "A"}, {"id": "f5-082", "domain": 5, "task_id": "5.3", "objective": "Implement error propagation strategies across multi-agent systems", "situation": "A team is designing the error payload that subagents return to a coordinator when a tool call fails.", "question": "Which set of fields best supports intelligent coordinator recovery decisions?", "options": [{"letter": "A", "text": "A free-text log excerpt of the subagent's entire internal reasoning trace leading right up to the failure", "correct": false, "explanation": "A full reasoning trace is noisy and unstructured; the coordinator would have to parse free text to extract the same information that structured fields would provide directly."}, {"letter": "B", "text": "Failure type, the action that was attempted, any partial results gathered, and alternative approaches to try next", "correct": true, "explanation": "These four elements together let a coordinator judge whether the failure is worth retrying, whether existing partial results are usable, and what alternative path might succeed."}, {"letter": "C", "text": "A single numeric error code mapped to an internal lookup table that only the platform team is allowed to maintain", "correct": false, "explanation": "A bare numeric code requires the coordinator to look up meaning elsewhere and still lacks partial results or alternatives, so it cannot drive a recovery decision on its own."}, {"letter": "D", "text": "A retry count and a timestamp, since the coordinator can infer the rest from prior attempts alone", "correct": false, "explanation": "Retry count and timestamp say nothing about why the failure happened, what was attempted, or whether partial results exist, so the coordinator still cannot make an informed recovery choice."}], "correct": "B", "select": 1, "group": "A"}, {"id": "f5-083", "domain": 5, "task_id": "5.2", "objective": "Design effective escalation and ambiguity resolution patterns", "situation": "A rideshare customer asks for a fare adjustment based on stacking two promotional codes together, a combination the promotions policy documentation never addresses in either direction. The customer is calm and not asking for a human agent.", "question": "What should the agent do?", "options": [{"letter": "A", "text": "Escalate the request, since documented policy neither permits nor prohibits stacking these codes", "correct": true, "explanation": "- the policy doesn't address this specific combination either way, so the ambiguity should be escalated rather than resolved by the agent's own inference."}, {"letter": "B", "text": "Ask the customer to select only one promotional code, then apply that single code's discount", "correct": false, "explanation": "- unilaterally deciding to apply only one code sidesteps the customer's actual request rather than resolving the underlying ambiguity."}, {"letter": "C", "text": "Deny the fare adjustment, since undocumented combinations should default to being disallowed", "correct": false, "explanation": "- assuming denial from the absence of explicit permission also invents a policy stance instead of acknowledging the actual gap."}, {"letter": "D", "text": "Apply the fare adjustment, since neither code's terms explicitly forbid combining it with another", "correct": false, "explanation": "- assuming permission from the absence of an explicit prohibition invents a policy position that was never actually stated."}], "correct": "A", "select": 1, "group": "C"}, {"id": "f5-084", "domain": 5, "task_id": "5.3", "objective": "Implement error propagation strategies across multi-agent systems", "situation": "A document-search subagent hits an authentication error against a knowledge base and simply returns the string 'search unavailable' to the coordinator. The coordinator has no other information to act on.", "question": "What is the primary problem with this design, and what should replace it?", "options": [{"letter": "A", "text": "The generic status is fine since coordinators should never see subagent detail; log the auth error only in the subagent trace", "correct": false, "explanation": "Hiding failure detail from the coordinator is the problem being described, not an acceptable design; the coordinator needs that context to make recovery decisions."}, {"letter": "B", "text": "The generic status is too verbose for the coordinator to parse; shorten it to a bare success or failure boolean flag", "correct": false, "explanation": "A boolean flag carries even less information than the current string and would make the coordinator's recovery decision harder, not easier."}, {"letter": "C", "text": "The generic status hides recovery detail; the subagent should return failure type, what was queried, and any partial results", "correct": true, "explanation": "A bare status string like 'search unavailable' strips away exactly the information (why it failed, what was attempted, what alternatives exist) that lets a coordinator decide between retrying, rerouting, or proceeding with partial coverage."}, {"letter": "D", "text": "The generic status correctly shields the coordinator; have the subagent retry silently until the auth error clears", "correct": false, "explanation": "Silently retrying without ever surfacing the failure risks masking a persistent auth problem that the subagent cannot resolve on its own, leaving the coordinator unaware anything is wrong."}], "correct": "C", "select": 1, "group": "A"}, {"id": "f5-085", "domain": 5, "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "situation": "A SaaS support chat has a customer raise a billing dispute (invoice #4821, $89.00 overcharge) and, later in the same conversation, a service outage affecting their workspace (incident started 2:15 PM, status: unresolved). The agent's single running narrative summary blends both threads, and a later reply about \"the amount\" ambiguously could refer to either issue.", "question": "What should the agent do to avoid this ambiguity in future multi-issue sessions?", "options": [{"letter": "A", "text": "Maintain separate structured records per issue: log invoice number and amount for billing dispute, start time and status for outage.", "correct": true, "explanation": "Maintaining separate structured records per issue ensures that each issue's specific details—such as the invoice number and amount for billing, and the start time and status for the outage—remain distinct. This approach eliminates the ambiguity that arises when a single narrative mixes multiple topics and makes it clear which issue a later reference like \"the amount\" belongs to."}, {"letter": "B", "text": "Summarize only the most recent issue, such as the outage, and treat earlier issues like the billing dispute as resolved once a new topic arises.", "correct": false, "explanation": "Summarizing only the most recent issue and treating earlier topics like the billing dispute as resolved when a new issue arises risks losing critical context. The still‑unresolved billing dispute could be forgotten entirely, leading to poor service and potential escalation."}, {"letter": "C", "text": "Ask the customer to hold the outage issue until the billing dispute for invoice #4821 is fully resolved, to keep the session focused.", "correct": false, "explanation": "Asking the customer to hold one issue until another is completely resolved restricts normal support workflows and does not address the root cause of the context‑tracking problem. Even with sequential handling, the agent still needs a robust way to distinguish details between issues."}, {"letter": "D", "text": "Continue using one summary but add issue-type labels: \"billing\" for amounts, \"outage\" for times and statuses, to each mention of an amount or time.", "correct": false, "explanation": "Adding manual labels like \"billing\" or \"outage\" to each mention of an amount or time within a single blended summary is error‑prone and easy to miss. This method does not create a durable, reliable structure that the agent can later reference to resolve ambiguous references."}], "correct": "A", "select": 1, "group": "G"}, {"id": "f5-086", "domain": 5, "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "situation": "An architect is about to delegate a codebase investigation to a subagent and is deciding how to phrase the instruction.", "question": "Which instruction best supports getting a targeted, high-quality summary back without still burdening the main agent's context?", "options": [{"letter": "A", "text": "Give the subagent no instructions at all and let it infer the goal from the main conversation's git status.", "correct": false, "explanation": "Giving no instructions forces the subagent to guess at the goal, risking wasted exploration and an unfocused, less useful summary."}, {"letter": "B", "text": "'Report back the full, unsummarized contents of every single file you happen to open during the investigation.'", "correct": false, "explanation": "Asking for full unsummarized file contents defeats the purpose of delegation, since the subagent's report carries the same volume of data back into the main context."}, {"letter": "C", "text": "'Trace refund flow dependencies across payments, ledger, and notifications, and summarize what you find.'", "correct": true, "explanation": "A specific, bounded question directs the subagent's search and lets it return a targeted, distilled summary, which keeps the main agent's context free of unnecessary bulk output."}, {"letter": "D", "text": "'Look around the codebase for a while and report back anything you happen to find interesting.'", "correct": false, "explanation": "An open-ended instruction like this tends to produce broad, unfocused exploration and a diffuse report, undermining the value of isolating a specific investigation."}], "correct": "C", "select": 1, "group": "G"}, {"id": "f5-087", "domain": 5, "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "situation": "A travel-booking assistant developer wants to cut latency by sending only the user's latest message plus a one-sentence rolling summary on each API call, rather than the full prior turns. Ten turns into a session, the customer says \"book the same seat type I mentioned earlier,\" but the assistant no longer has access to that detail because it was dropped from the rolling summary.", "question": "What is the underlying issue?", "options": [{"letter": "A", "text": "The rolling summary approach fails only because it was one sentence long; a two- or three-sentence summary would retain the seat type", "correct": false, "explanation": "Lengthening the rolling summary might help this one case but does not address the root cause: any compression scheme can drop an arbitrary detail depending on wording and length limits."}, {"letter": "B", "text": "The model's built-in conversational memory should have retained the seat preference across calls without it being resent", "correct": false, "explanation": "There is no persistent memory between separate API calls; assuming otherwise is the mistake that caused the failure."}, {"letter": "C", "text": "The customer should have repeated the seat preference on every subsequent turn to guarantee it stays available", "correct": false, "explanation": "Requiring the customer to repeat themselves works around the symptom but blames the user for a context-management problem the application should have solved."}, {"letter": "D", "text": "The API treats each request as stateless, so any turn left out of the message history sent with it is unavailable to the model", "correct": true, "explanation": "— the API does not retain memory between separate requests; every fact the model needs, including details from turn one, must be present in the message history sent with the current request."}], "correct": "D", "select": 1, "group": "G"}, {"id": "f5-088", "domain": 5, "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "situation": "A market-research pipeline dispatches five research subagents, each of which reads several documents and returns a free-form paragraph summary to a coordinator agent. When the coordinator assembles the final report, reviewers find that several claims can no longer be traced to any specific source.", "question": "Which change to the subagent output contract would best prevent this loss of provenance?", "options": [{"letter": "A", "text": "Require each subagent to limit its paragraph to three sentences so the coordinator can quote it directly", "correct": false, "explanation": "Shortening the paragraph limits how much a subagent can say but does nothing to attach individual claims to their sources, so provenance can still be lost when paragraphs are combined."}, {"letter": "B", "text": "Require each subagent to pair every claim with its source URL or document name and a supporting excerpt", "correct": true, "explanation": "A structured claim-to-source mapping ties each individual assertion to its origin, so when the coordinator merges outputs from multiple subagents it can preserve attribution instead of losing it during re-summarization."}, {"letter": "C", "text": "Require each subagent to append a bibliography of all documents it consulted at the end of its free-form paragraph", "correct": false, "explanation": "A trailing bibliography lists what was consulted overall but does not link any single claim to the specific source it came from, so attribution is still lost once claims are merged."}, {"letter": "D", "text": "Require each subagent to raise its temperature setting so summaries retain more of the original document wording", "correct": false, "explanation": "Temperature affects the randomness of generated text, not whether structured source metadata is captured, so this would not restore per-claim provenance."}], "correct": "B", "select": 1, "group": "B"}, {"id": "f5-089", "domain": 5, "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "situation": "A due-diligence pipeline dispatches five subagents to review different regulatory filings for the same company. Four subagents independently report that the company's debt-to-equity ratio is stable, while one subagent, using a filing from a different fiscal quarter, reports a much higher ratio. The synthesis agent must decide how to characterize this in the final report.", "question": "What is the best approach?", "options": [{"letter": "A", "text": "Present the stable ratio as well established and the higher one as a distinct, dated data point, not an outlier to discard", "correct": true, "explanation": "Once the differing figure is recognized as coming from a different fiscal quarter, it should be presented as a distinct, dated data point alongside the well-established trend rather than discarded or blended, preserving both provenance and temporal context."}, {"letter": "B", "text": "Report the higher ratio as the current figure since it is presumed to reflect the most recently filed quarter", "correct": false, "explanation": "Assuming the higher ratio is simply 'more recent' without verifying the filing dates risks mischaracterizing which figure is actually current."}, {"letter": "C", "text": "Discard the higher ratio as noise since it disagrees with the majority of the subagent reports on this filing", "correct": false, "explanation": "Discarding the figure as noise ignores the fact that it comes from a different, valid fiscal quarter and simply reflects timing rather than an error."}, {"letter": "D", "text": "Report only the average of all five ratios as the company's representative debt-to-equity figure for this quarter", "correct": false, "explanation": "Averaging figures from different fiscal quarters produces a number that does not correspond to any actual reporting period and obscures the true trend."}], "correct": "A", "select": 1, "group": "B"}, {"id": "f5-090", "domain": 5, "task_id": "5.1", "objective": "Manage conversation context to preserve critical information across long interactions", "situation": "A finance team's orchestrator runs six analyst subagents, each producing several paragraphs of exploratory reasoning before stating a conclusion, to build a single executive summary. The summary-writing agent has a strict context budget and currently truncates the last two subagents' output entirely because the first four already filled its budget.", "question": "What change addresses this?", "options": [{"letter": "A", "text": "Have the orchestrator forward only the first four subagents' output and inform the summary-writing agent that the analysis is complete", "correct": false, "explanation": "Silently dropping two subagents' analyses while claiming completeness produces an executive summary that omits real findings without any indication that content was lost."}, {"letter": "B", "text": "Instruct the summary-writing agent to skip ahead and read the last two subagents' output first, then work backward", "correct": false, "explanation": "Reordering which subagents are read first only changes which ones get truncated; it does not increase the amount of content that fits overall."}, {"letter": "C", "text": "Increase the number of analyst subagents so each covers a smaller slice of the analysis in equally long reasoning narratives", "correct": false, "explanation": "More subagents with equally long reasoning narratives increases total volume rather than fitting more analyses into the same budget."}, {"letter": "D", "text": "Have each subagent output only its key facts, supporting citations, and a relevance score, dropping the exploratory narrative", "correct": true, "explanation": "— trimming each subagent's contribution to key facts, citations, and a relevance score reduces the total volume so all six analyses fit within the summary-writing agent's budget instead of only the first four."}], "correct": "D", "select": 1, "group": "G"}, {"id": "f5-091", "domain": 5, "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "situation": "An architect is designing a multi-agent system that explores a large codebase over a run that may take hours and could be interrupted by a process crash. They want the run to resume cleanly rather than restart from zero.", "question": "How should the system be structured to support this?", "options": [{"letter": "A", "text": "Rely on the conversation history staying in context indefinitely, since the coordinator will always remember every agent's progress unaided.", "correct": false, "explanation": "Conversation history does not survive a process crash, and even within one long session it is subject to the same context degradation the system is trying to avoid."}, {"letter": "B", "text": "Each agent exports its state to a known file location as it progresses, and the coordinator loads a manifest of agent states on resume.", "correct": true, "explanation": "This is structured state persistence: agents export progress to a known location, and the coordinator's manifest lets it load and inject the relevant state into prompts when resuming."}, {"letter": "C", "text": "Restart the entire exploration from the beginning any time the process is interrupted, since partial progress cannot be recovered.", "correct": false, "explanation": "Restarting from zero after every interruption wastes the work already completed by agents that finished before the crash and does not scale to hours-long runs."}, {"letter": "D", "text": "Have each agent keep its state only in its own memory during execution, without writing anything the coordinator can read later.", "correct": false, "explanation": "State kept only in an agent's own memory is lost the moment that agent's process ends, giving the coordinator nothing to load on resume."}], "correct": "B", "select": 1, "group": "G"}, {"id": "f5-092", "domain": 5, "task_id": "5.3", "objective": "Implement error propagation strategies across multi-agent systems", "situation": "A log-analysis subagent's query to a metrics store fails once with a transient 503, then succeeds on an internal retry a second later. A separate subagent's query to the same store fails repeatedly for two minutes because the store's credentials were rotated and never propagated to the subagent's environment.", "question": "How should each situation be handled?", "options": [{"letter": "A", "text": "Escalate both failures to the coordinator immediately, since subagents should never attempt any local retries at all", "correct": false, "explanation": "Escalating every transient blip defeats the purpose of local recovery and adds unnecessary coordinator overhead for a problem the subagent already resolved itself."}, {"letter": "B", "text": "Retry both failures locally and indefinitely by the subagent until one of them eventually succeeds on its own", "correct": false, "explanation": "Retrying the credential failure indefinitely wastes time and resources on a problem that local retries cannot fix, since the root cause requires external intervention."}, {"letter": "C", "text": "Resolve the transient 503 locally without the coordinator, but escalate the credential failure with what was attempted", "correct": true, "explanation": "A short-lived transient error is exactly the kind of failure a subagent can resolve on its own; a persistent credential problem cannot be fixed by more retries and should be propagated with details of the attempt so the coordinator can intervene."}, {"letter": "D", "text": "Escalate the transient 503 to the coordinator but retry the credential failure locally until the rotation eventually completes", "correct": false, "explanation": "This reverses the correct handling: the already-resolved transient error doesn't need escalation, while the unresolvable credential error does need to reach the coordinator."}], "correct": "C", "select": 1, "group": "A"}, {"id": "f5-093", "domain": 5, "task_id": "5.6", "objective": "Preserve information provenance and handle uncertainty in multi-source synthesis", "situation": "A synthesis agent combines a subagent finding that unemployment in a region is '5.1%' with another subagent finding that it is '6.3%'. The coordinator initially treats this as a contradiction requiring reconciliation, but on closer inspection the two figures come from reports published fourteen months apart.", "question": "What structural change to subagent output would prevent this kind of false contradiction going forward?", "options": [{"letter": "A", "text": "Require every subagent to include the publication or data-collection date alongside each figure it extracts", "correct": true, "explanation": "Attaching the publication or collection date to each figure lets the coordinator recognize that a difference reflects the passage of time rather than a genuine conflict between sources."}, {"letter": "B", "text": "Require every subagent to discard any figure that is more than six months old before passing it to the coordinator", "correct": false, "explanation": "Discarding older figures throws away real information and does not solve the underlying problem of missing dates on the figures that remain."}, {"letter": "C", "text": "Require every subagent to convert all statistics into a rolling twelve-month average before reporting them", "correct": false, "explanation": "Computing a rolling average requires more data points than a single subagent typically has access to, and it would blend the two figures rather than preserving each one with its own context."}, {"letter": "D", "text": "Require every subagent to phrase all extracted figures as approximate ranges rather than exact percentages", "correct": false, "explanation": "Converting exact figures into vague ranges reduces precision without addressing the actual cause of the false contradiction, which is the missing timestamp."}], "correct": "A", "select": 1, "group": "B"}, {"id": "f5-094", "domain": 5, "task_id": "5.4", "objective": "Manage context effectively in large codebase exploration", "situation": "An architect has been running a single Claude Code session for several hours exploring a large monorepo. Early on, the agent correctly identified that OrderService extends a custom TransactionalBase class with a distinctive retry mechanism. Hours later, asked about OrderService again, the agent describes it as 'typically extending a standard base controller with default error handling,' ignoring its earlier finding.", "question": "What should the architect do going forward to prevent this?", "options": [{"letter": "A", "text": "Have the agent record concrete findings, such as exact class names, in a scratchpad file and consult it before answering.", "correct": true, "explanation": "Persisting exact findings to a scratchpad file gives the agent a durable, specific anchor that survives context degradation, so later answers stay grounded in what was actually discovered."}, {"letter": "B", "text": "Increase the maximum output token limit for the session so the agent has more room to describe the class in full detail each time it is asked.", "correct": false, "explanation": "Output token limits control response length, not whether earlier findings remain salient in context, so raising this limit does not address degradation."}, {"letter": "C", "text": "Ask the agent to re-read the entire repository from scratch every time a question about a previously discovered class comes up.", "correct": false, "explanation": "Re-reading the whole repository for every follow-up is expensive and does not build a durable record the agent can reliably reference for the rest of the session."}, {"letter": "D", "text": "Switch to a larger model mid-session, since added parameter count restores recall of facts discovered earlier in the transcript.", "correct": false, "explanation": "Model size does not restore recall of facts diluted across a long context window; the problem is context management, not raw model capability."}], "correct": "A", "select": 1, "group": "G"}, {"id": "f5-095", "domain": 5, "task_id": "5.5", "objective": "Design human review workflows and confidence calibration", "situation": "A vendor pitches a document-processing tool that claims '99% accuracy, no human review needed.' Before adopting it for a regulated workflow, an architect wants to stress-test this claim using the same rigor the team applies to its own Claude-based pipelines.", "question": "Which question is most important to ask the vendor?", "options": [{"letter": "A", "text": "Can you confirm the tool uses the newest available model version, since model recency is the strongest predictor of whether an accuracy claim will hold up in production?", "correct": false, "explanation": "Model recency is not the strongest predictor of production accuracy. While newer models may have improved capabilities, accuracy in document processing depends heavily on prompt engineering, document quality, system architecture (e.g., use of a semantic layer or dedicated OCR pipelines), and specific fine-tuning or validation for the task. Official documentation does not support that using the latest model version alone guarantees a specific accuracy figure."}, {"letter": "B", "text": "Can you provide a written guarantee that the 99% figure will not decline over time, so the team has contractual recourse if accuracy drops after deployment?", "correct": false, "explanation": "A written guarantee addresses contractual risk but does not help stress-test the technical validity of the accuracy claim before adoption. The architect's goal is to evaluate the claim's rigor, not to secure post-deployment recourse. The research emphasizes that understanding the measurement methodology and accuracy breakdown is the most important step when assessing such claims."}, {"letter": "C", "text": "Can you provide the accuracy breakdown by document type and field, along with the labeled validation methodology used to produce the 99% figure, rather than just the aggregate number?", "correct": true, "explanation": "Asking for a granular accuracy breakdown by document type and field, along with the validation methodology, is the most rigorous approach. According to Anthropic documentation, the often-cited 99% accuracy for Claude models typically refers to the 'Needle In A Haystack' (NIAH) evaluation, which measures recall of specific information in long contexts, not general document processing accuracy across diverse document types and fields. Without a detailed breakdown and labeled methodology, a blanket 99% claim is not trustworthy for regulated workflows. Official recommendations emphasize structured prompting, high-resolution inputs, and controlled temperature settings, but no official documentation supports a universal 99% field-level accuracy across all documents."}, {"letter": "D", "text": "Can you confirm the 99% figure was measured on a sample of at least ten thousand documents, since larger sample sizes are the primary determinant of a metric's trustworthiness?", "correct": false, "explanation": "Sample size alone is not the primary determinant of trustworthiness. A large sample without a proper validation methodology, field-level accuracy breakdown, or stratification across document types can still produce misleading aggregate figures. The research indicates that the methodology and granularity of the accuracy assessment (e.g., per-field, per-document-type) are far more critical than sheer sample size."}], "correct": "C", "select": 1, "group": "C"}]