- 本文基于
anthropic(Python)1.12.1(commit50b78d1,2026-10-08)。1.x 版本要求 Python ≥ 3.10,底层 HTTP 库从httpx换成了 httpx2,从 0.x 升级前请先阅读 迁移指南12。 - 示例模型统一使用
claude-opus-5-5,官方建议大多数场景都从它开始3。文中代码已用 pyright 对照 1.12.1 做过类型检查;没有 API Key,因此没有实际运行。 - Tool Runner、compaction、server-side fallback 等功能目前是 beta。资料截至 2026-10-09,本文不随官方同步更新,请对照 Claude API 文档 和 Release notes 使用。
- Messages API 是无状态的:每次请求都要带上完整历史。响应内容是一组 content block,类型有
text、thinking、tool_use等,用stop_reason判断模型为什么停下45。 - 新一代模型使用 adaptive thinking,用
effort控制推理深度。claude-opus-5-5的 thinking 始终开启,默认 effort 是medium3。 - 工具分两类:客户端工具由你执行,服务端工具由 Anthropic 执行。循环可以手写,也可以交给 Tool Runner(beta)。
- 生产环境常用的能力:结构化输出(
messages.parse)、流式输出(messages.stream)、prompt caching、refusal 与 fallback、compaction。
1. 定位#
- 分层位置:
anthropic包属于 L1,Messages API 属于 L0/L2,Tool Runner 属于 L3。 - Tool Runner 和 Claude Agent SDK 是两个不同的东西。Tool Runner 在
anthropic包里,通过client.beta.messages.tool_runner调用,只负责循环你自己定义的工具,没有内置工具,也没有沙箱。Agent SDK 是另一个包,带有完整的 Claude Code harness,见 3.3 Claude Agent SDK67。
2. 安装与客户端配置#
pip install anthropicexport ANTHROPIC_API_KEY=sk-ant-...import anthropic
client = anthropic.Anthropic() # 默认从环境变量 ANTHROPIC_API_KEY 读取密钥# 异步版本:client = anthropic.AsyncAnthropic()
response = client.messages.create( model="claude-opus-5-5", max_tokens=16000, # 必填:本次最多输出多少 token system="你是一个资深后端工程师,回答要简洁。", messages=[{"role": "user", "content": "用一句话解释什么是幂等性。"}],)for block in response.content: # content 是一组 content block if block.type == "text": print(block.text)print(response.stop_reason, response.usage.input_tokens, response.usage.output_tokens)| 配置项 | 默认值 | 说明 |
|---|---|---|
| 重试 | 2 次 | 连接错误、408、409、429、≥500 会自动重试,使用指数退避;max_retries=0 可以关闭2 |
| 超时 | 10 分钟 | 可以传浮点数,或者 httpx2.Timeout;超时后同样会重试2 |
| 单次请求覆盖 | — | client.with_options(timeout=5.0, max_retries=5).messages.create(...) |
3. 请求与响应的结构#
messages中role的取值是 user 或 assistant。支持该特性的模型还可以在对话中途插入 system 消息4。ToolUseBlock.input已经是解析好的 dict,不需要再json.loads。
stop_reason 一览#
| 值 | 含义 | 你该做什么 |
|---|---|---|
end_turn | 模型自然结束 | 读取文本 |
tool_use | 模型想调用客户端工具 | 执行工具,然后回传 tool_result |
max_tokens | 达到 max_tokens 上限 | 增大上限,或者改用流式;被截断的工具调用不要执行 |
stop_sequence | 遇到了自定义的停止序列 | — |
pause_turn | 服务端工具的循环暂停了 | 把这一轮追加进历史后重新发送,让模型继续 |
refusal | 安全分类器拒绝了请求,HTTP 状态仍为 200 | 查看 stop_details,或者使用 fallback(见第 11 节) |
各值的含义见官方 Stop reasons 页面5,refusal 的细节见 Refusals and fallback8。
多轮对话:自己保存历史#
from anthropic import Anthropicfrom anthropic.types import MessageParam
client = Anthropic()history: list[MessageParam] = []
def chat(user_text: str) -> str: history.append({"role": "user", "content": user_text}) response = client.messages.create(model="claude-opus-5-5", max_tokens=16000, messages=history) # 把完整的 content(包括 thinking 块)原样追加进历史,而不只是文本 history.append({"role": "assistant", "content": response.content}) return "".join(b.text for b in response.content if b.type == "text")
print(chat("我叫小王。"))print(chat("我叫什么?")) # 模型能记住,是因为我们把历史都发了过去4. Adaptive Thinking 与 Effort#
from anthropic import Anthropic
client = Anthropic()response = client.messages.create( model="claude-opus-5-5", max_tokens=16000, # Opus 5.5 的 thinking 始终开启;display=summarized 表示返回可读的推理摘要 thinking={"type": "adaptive", "display": "summarized"}, output_config={"effort": "high"}, # 可选值:low / medium / high / xhigh / max messages=[{"role": "user", "content": "一步一步地证明 √2 是无理数。"}],)for block in response.content: if block.type == "thinking": print("[推理摘要]", block.thinking) elif block.type == "text": print("[回答]", block.text)| 要点 | 说明 |
|---|---|
| 自适应 | 模型自己决定什么时候思考、思考多少,用 effort 引导;budget_tokens 这种手动预算的写法已经淘汰39 |
| 各模型的默认 effort | Fable 5.1 为 high,Opus 5.5 为 medium,Sonnet 5.5 为 high,Haiku 5.5 为 medium3 |
| 传回规则 | 带着 thinking 块继续对话时,要原样传回,工具调用期间尤其如此9 |
| 缓存 | 连续请求保持相同的 thinking 配置和 effort,prompt cache 就不会失效9 |
5. 工具调用#
5.1 手写循环#
见 1.2 Tool Calling 与 Agent Loop 原理 › 4. 手写一个最小 Agent Loop:Anthropic 版。
5.2 Tool Runner(beta):不用手写循环#
from anthropic import Anthropic, beta_tool
client = Anthropic()
@beta_tooldef get_weather(city: str) -> str: """查询城市今天的天气。
Args: city: 城市名,例如 北京。 """ return {"北京": "晴,25°C", "上海": "小雨,22°C"}.get(city, "暂无数据")
@beta_tooldef calculate_sum(a: int, b: int) -> str: """计算两个整数之和。
Args: a: 第一个数 b: 第二个数 """ return str(a + b)
runner = client.beta.messages.tool_runner( model="claude-opus-5-5", max_tokens=16000, tools=[get_weather, calculate_sum], # 装饰器会根据函数签名和 docstring 生成 JSON Schema messages=[{"role": "user", "content": "北京天气如何?另外 15 + 27 等于多少?"}],)final_message = runner.until_done() # 自动循环,直到模型不再调用工具print("".join(b.text for b in final_message.content if b.type == "text"))官方说明 Tool Runner 会自动完成四件事:调用工具、处理请求与响应的往返、维护对话状态、做类型校验。它目前处于 beta,Python、TypeScript、C#、Go、Java、Ruby、PHP 各 SDK 都已提供7。
- 需要人工审批、自定义日志或条件执行时,官方建议改用手写循环7。
- Python 版 runner 遇到
pause_turn不会自动续跑(服务端工具可能触发这种情况),循环会直接结束,返回的结果看起来就像是截断了。如果用到服务端工具,请自己处理续跑,或者改用手写循环。这一点来自 Anthropic 官方的 claude-api 技能参考,笔者没有独立验证。
5.3 强制工具调用与 strict 模式#
在 Claude Opus 5.5、Sonnet 5.5、Fable 5.1 上,tool_choice 设为 {"type": "any"} 或 {"type": "tool", ...},也就是强制工具调用,会返回 400。官方给出的替代方案是:使用 auto,加上 strict tool use 保证参数符合 Schema,再在提示词里引导模型去调用工具10。
from anthropic import Anthropicfrom anthropic.types import ToolParam
client = Anthropic()book_flight: ToolParam = { "name": "book_flight", "description": "预订航班", "strict": True, # 保证 tool_use.input 严格符合下面的 Schema "input_schema": { "type": "object", "properties": { "destination": {"type": "string"}, "date": {"type": "string", "format": "date"}, "passengers": {"type": "integer", "enum": [1, 2, 3, 4, 5, 6]}, }, "required": ["destination", "date", "passengers"], "additionalProperties": False, },}response = client.messages.create( model="claude-opus-5-5", max_tokens=16000, tools=[book_flight], messages=[{"role": "user", "content": "帮我订 3 月 15 日去东京的机票,2 个人。请使用 book_flight 工具。"}],)calls = [b for b in response.content if b.type == "tool_use"]print(calls[0].input if calls else "模型没有调用工具,需要再引导一次")5.4 服务端工具:由 Anthropic 执行#
from anthropic import Anthropicfrom anthropic.types import MessageParam
client = Anthropic()messages: list[MessageParam] = [{"role": "user", "content": "火星车最近有什么新进展?"}]
def ask(msgs: list[MessageParam]): return client.messages.create( model="claude-opus-5-5", max_tokens=16000, tools=[{"type": "web_search_20260209", "name": "web_search", "max_uses": 5}], messages=msgs, )
response = ask(messages)for _ in range(5): # 给 pause_turn 续跑设一个上限 if response.stop_reason != "pause_turn": break # 服务端工具的循环暂停了:把这一轮追加进历史后重新发送,模型会接着做 messages.append({"role": "assistant", "content": response.content}) response = ask(messages)
for block in response.content: if block.type == "text": print(block.text)| 服务端工具 | type | 说明 |
|---|---|---|
| Web 搜索 | web_search_20260209 | 带动态过滤:在代码执行环境里运行搜索,API 会自动开通需要的代码执行,不必在 tools 里另外声明11 |
| Web 抓取 | web_fetch_20260209 | 为了降低数据外泄风险,只能抓取对话里之前出现过的 URL,抓不到只在 Claude 自己的输出里出现的 URL12 |
| 代码执行 | code_execution_20260521 | 在沙箱中运行 bash 和代码,可以搭配 Agent Skills |
| 工具搜索 | tool_search_tool_regex_20251119 / ..._bm25_... | 配合 defer_loading 按需加载工具 |
表中的类型名都已在官方 Tool reference 中核对过13。几点补充:
- 工具类型名带有日期后缀,同一个工具会有多个版本。例如 web search 和 web fetch 已经有更新的
_20260318版本,新增了控制响应内容的选项13。 - 带动态过滤的
_20260209及更新版本,内部依赖代码执行,因此默认不符合 ZDR(零数据保留)要求。需要满足 ZDR 时,可以设置allowed_callers: ["direct"]关闭动态过滤14。 - 服务端工具的公共机制(
server_tool_use块、pause_turn续跑、服务端与客户端工具混用)见 Server tools 页面14。
还有一类 Anthropic 定义的客户端工具,它们没有 input_schema,但执行的仍然是你:
bashtext_editormemory- computer use
声明方式类似 {"type": "bash_20250124", "name": "bash"}15。
5.5 并行工具调用#
Claude 默认可以在一次响应里请求多个工具。官方总结了两条格式规则16:
- 所有结果都放在同一条 user 消息里;
- 这条消息里,结果前面不能有文本。
只读、相互独立的工具适合并行执行;有副作用的工具,最好串行执行16。
6. 结构化输出#
from pydantic import BaseModel
from anthropic import Anthropic
client = Anthropic()
class ContactInfo(BaseModel): name: str email: str plan: str demo_requested: bool
response = client.messages.parse( model="claude-opus-5-5", max_tokens=16000, messages=[{"role": "user", "content": "抽取信息:Jane Doe(jane@co.com)想要 Enterprise 套餐,并希望看演示。"}], output_format=ContactInfo, # SDK 会把它转换成 API 的 output_config.format)contact = response.parsed_output # 已经校验过的 ContactInfo 实例print(contact)messages.parse()是官方推荐的写法。SDK 会把 Pydantic 模型的 schema 转换后,作为output_config.format发送出去,再校验响应,最后通过parsed_output返回解析好的对象17。- API 层的旧参数
output_format已经迁移到output_config.format,并标注为废弃17。注意区分两件事:Python SDK 里parse()的形参名仍然叫output_format,但它是 SDK 层的参数。
7. 流式输出#
from anthropic import Anthropic
client = Anthropic()
with client.messages.stream( model="claude-opus-5-5", max_tokens=64000, # 流式请求可以设更大的上限,不用担心 HTTP 超时 messages=[{"role": "user", "content": "写一个 300 字的科幻小故事。"}],) as stream: for text in stream.text_stream: # 只关心文本增量 print(text, end="", flush=True) final = stream.get_final_message() # 流结束后拿到完整的 Message
print("\n", final.stop_reason, final.usage.output_tokens)SSE 事件的顺序是固定的18:
message_start;- 每个内容块依次出现
content_block_start、若干个content_block_delta、content_block_stop; - 一个或多个
message_delta,其中的 usage 是累计值; message_stop。
delta 的类型有 text_delta、input_json_delta(工具参数)、thinking_delta,以及 signature_delta18。
8. Prompt Caching#
from anthropic import Anthropic
client = Anthropic()LONG_DOC = "(这里是一份很长的产品手册……)" * 500 # 示意:被缓存的前缀需要足够长
for question in ["退货政策是什么?", "保修多久?"]: response = client.messages.create( model="claude-opus-5-5", max_tokens=16000, system=[ { "type": "text", "text": f"你是客服助手。以下是产品手册:\n{LONG_DOC}", "cache_control": {"type": "ephemeral"}, # 在这里打一个缓存断点 } ], messages=[{"role": "user", "content": question}], ) u = response.usage print(question, "写入缓存:", u.cache_creation_input_tokens, "命中缓存:", u.cache_read_input_tokens)| 要点 | 说明 |
|---|---|
| 前缀匹配 | 缓存按 tools → system → messages 的顺序渲染,前缀中任何一个字节变化,都会让它后面的缓存失效19 |
| 典型的“隐形失效” | 在 system prompt 里放时间戳或 UUID、json.dumps 时 key 的顺序不固定、工具列表每次都不一样 |
| 价格 | 缓存读取通常是基础输入价格的 10%;在 Opus 5.5 和 Sonnet 5.5 上是 5%3 |
| 最小可缓存长度 | 前缀太短时不会被缓存,也不会报错,具体阈值与模型有关19 |
9. 长对话:Compaction 与 Context Editing(beta)#
- Compaction(压缩):上下文接近触发阈值时,服务端自动把较早的历史总结成摘要。使用时必须把
response.content整个追加进历史,因为 compaction 块需要在下一次请求中传回去20。 - Context editing(上下文编辑):清理掉旧的工具结果或 thinking 块。它是“清除”,不是“总结”21。
10. 错误处理#
import anthropic
client = anthropic.Anthropic()
try: response = client.messages.create( model="claude-opus-5-5", max_tokens=1024, messages=[{"role": "user", "content": "hi"}] ) print(response._request_id) # 排查问题时把这个 ID 提供给 Anthropicexcept anthropic.BadRequestError as e: # 400:参数有误,不要重试 print("请求参数错误:", e.message)except anthropic.RateLimitError as e: # 429:退避后重试 print("限流,建议等待:", e.response.headers.get("retry-after"))except anthropic.APIStatusError as e: # 其他非 2xx 状态码 print("API 错误:", e.status_code)except anthropic.APIConnectionError: # 网络问题 print("网络连接失败")按照“从具体到宽泛”的顺序逐级捕获异常,就能区分哪些错误可以重试(429、5xx、网络问题),哪些不能重试(400、404)。
11. 生产模板:流式输出、refusal 处理与服务端 fallback#
Claude Fable 5.1、Fable 5、Opus 5.5、Opus 5、Sonnet 5.5、Haiku 5.5 都带有安全分类器。请求被拒绝时,API 返回一个正常的响应,而不是错误:stop_reason 为 "refusal",stop_details.category 标明拒绝所属的策略类别8。
在 Claude API 上,可以开启 beta 的服务端 fallback:设置 fallbacks="default",API 会把被拒绝的请求转给 Anthropic 针对该类别推荐的备用模型重试;如果某个类别没有推荐的备用模型,拒绝结果保持不变8。
from anthropic import Anthropic
client = Anthropic()
response = client.beta.messages.create( model="claude-opus-5-5", max_tokens=16000, messages=[{"role": "user", "content": "给我一份 Linux 服务器加固清单。"}], fallbacks="default", # 被拒绝时,由服务端自动换成推荐的备用模型重试 betas=["server-side-fallback-2026-07-01"],)
if response.stop_reason == "refusal": # 备用模型也拒绝了 print("请求被拒绝:", response.stop_details)else: print("实际作答的模型:", response.model) print("".join(b.text for b in response.content if b.type == "text"))服务端 fallback 目前只在 Claude API 上可用。在 Bedrock、Vertex、Foundry 上,官方提供客户端 SDK 的 middleware 方案。计费规则(拒绝本身怎么计费、备用模型怎么计费)请查阅官方页面8。
12. 其他端点与平台#
| 能力 | 入口 | 说明 |
|---|---|---|
| Message Batches | client.messages.batches.* | 异步批处理,价格为 50%3 |
| Files API | client.files.* | 上传一次,多次引用 |
| Models API | client.models.list() / retrieve() | 返回 max_input_tokens、max_tokens、capabilities 等字段3 |
| Token 计数 | client.messages.count_tokens(...) | 请求前估算输入 token |
| 其他平台 | Bedrock、Vertex AI、Foundry、Claude Platform on AWS | 各平台使用专用的客户端类,接口与 messages.create 一致 |
小结#
- 一个端点,加上 content block 和 stop_reason,构成了 Messages API 的全部心智模型。
- 2026 年的写法要点:
- 使用 adaptive thinking 加 effort;
- 不能强制工具调用,改用
auto加strict; - 结构化输出用
output_config.format,Python SDK 里写messages.parse; - 处理
refusal,并考虑开启 fallback。
- Tool Runner 省掉了手写循环;如果想要带内置工具的完整 harness,请看 3.3 Claude Agent SDK。
相关笔记#
- 3.1 Claude 开发者生态全景 · 3.3 Claude Agent SDK · 3.4 Claude Managed Agents
- 对照:2.2 OpenAI 客户端 SDK 与 Responses API
参考资料#
注释与出处#
-
anthropics/anthropic-sdk-python,
README.md(Requirements:Python 3.10+;0.x 升级请参阅 MIGRATION.md),https://github.com/anthropics/anthropic-sdk-python/blob/50b78d17a8a73bef97c3884102310344ac00f056/README.md ↩ -
Anthropic,Python SDK(Retries、Timeouts 等节),https://platform.claude.com/docs/en/cli-sdks-libraries/sdks/python ↩ ↩2 ↩3
-
Anthropic,Models overview,https://platform.claude.com/docs/en/about-claude/models/overview ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7
-
Anthropic,Working with the Messages API,https://platform.claude.com/docs/en/build-with-claude/working-with-messages ↩ ↩2
-
Anthropic,Stop reasons and fallback,https://platform.claude.com/docs/en/build-with-claude/handling-stop-reasons ↩ ↩2
-
Anthropic,Agent SDK overview,https://code.claude.com/docs/en/agent-sdk/overview ↩
-
Anthropic,Tool runner (SDK),https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-runner ↩ ↩2 ↩3
-
Anthropic,Refusals and fallback,https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback ↩ ↩2 ↩3 ↩4
-
Anthropic,Adaptive thinking,https://platform.claude.com/docs/en/build-with-claude/adaptive-thinking ↩ ↩2 ↩3
-
Anthropic,Define tools(“Not every model and setting supports forced tool use” 一节中的表格),https://platform.claude.com/docs/en/agents-and-tools/tool-use/define-tools ↩
-
Anthropic,Web search tool(Dynamic filtering 一节),https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool ↩
-
Anthropic,Web fetch tool(URL 来源限制),https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-fetch-tool ↩
-
Anthropic,Tool reference(各工具 type 名及其版本),https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-reference ↩ ↩2
-
Anthropic,Server tools(pause_turn 续跑;ZDR 与动态过滤),https://platform.claude.com/docs/en/agents-and-tools/tool-use/server-tools ↩ ↩2
-
Anthropic,Tool use with Claude,https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview ↩
-
Anthropic,Parallel tool use,https://platform.claude.com/docs/en/agents-and-tools/tool-use/parallel-tool-use ↩ ↩2
-
Anthropic,Structured outputs(推荐使用
client.messages.parse();API 层的output_format已迁移到output_config.format并废弃),https://platform.claude.com/docs/en/build-with-claude/structured-outputs ↩ ↩2 -
Anthropic,Streaming Messages(Event types 一节),https://platform.claude.com/docs/en/build-with-claude/streaming ↩ ↩2
-
Anthropic,Prompt caching,https://platform.claude.com/docs/en/build-with-claude/prompt-caching ↩ ↩2
-
Anthropic,Compaction,https://platform.claude.com/docs/en/build-with-claude/compaction ↩
-
Anthropic,Release notes(2025-09-29 条目:context editing beta),https://platform.claude.com/docs/en/release-notes/overview ↩