![[LangChain智能体本质论-05]结构化输出的两种实现方式](http://pic.xiahunao.cn/yaotu/[LangChain智能体本质论-05]结构化输出的两种实现方式)
数据只有具有希望的结构才可能被有效地处理我们在调用create_agent函数创建Agent的时候可以利用response_format控制输出的格式。我们可以定义一个Pydantic模型表示期望输出的结构。如果将此类型表示成ResponseT我们可以直接将此类型作为response_format参数的值也可以将此参数赋值为ResponseFormat[ResponseT]。除此之外我们也可以创建一个表示JSON Schema的字典作为response_format参数的值。ResponseFormat类型是针对ToolStrategy、ProviderStrategy和AutoStrategy三个类型的联合。ToolStrategy和ProviderStrategy代表了两种实现结构化输出的不同技术路径。defcreate_agent(...response_format:ResponseFormat[ResponseT]|type[ResponseT]|dict[str,Any]|NoneNone,...)ResponseFormatToolStrategy[SchemaT]|ProviderStrategy[SchemaT]|AutoStrategy[SchemaT]1. ToolStrategyToolStrategy是目前针对结构化输出最稳定、最通用的实现方式因为它利用的是工具调用这个所有LLM都具备的能力。LangChain会针对指定的Pydantic类型或者JSON Schema额外注册一个工具来完成格式化。这样LLM就会在回复消息的tool_calls中生成针对格式化工具的调用。Agent拦截这个调用将其解析为结构化数据并自动终止整个执行流程不在继续调用模型。下面的代码演示了针对ToolStrategy的使用。我们定义如下这个名为WeatherResponse的Pydantc模型类描述针对天气查询的响应。工具函数get_weather以字符串文本的形式返回指定城市的天气信息。我们调用create_agent函数创建的Agent时将模型指定为一个ChatOpenAI对象使用模型gpt-5.2-chatresponse_format参数设置为针对WeatherResponse类型创建的ToolStrategy对象。fromtypingimportAnnotated,Anyfromlangchain_openaiimportChatOpenAIfromlangchain_core.toolsimporttoolfrompydanticimportBaseModel,Fieldfromlangchain.agentsimportcreate_agentfromlangchain.agents.structured_outputimportToolStrategyfromdotenvimportload_dotenvimportjson load_dotenv()classWeatherResponse(BaseModel):A structured response format for weather information.city:strField(descriptionCity for which the weather is being reported)temperature:floatField(descriptionCurrent temperature in Celsius)summary:strField(descriptionBrief summary of the weather conditions)suggestion:strField(descriptionClothing suggestion based on the weather)defget_weather(city:str):Get the current weather for a given city.returnIt is sunny today, and the temperature is about 25.0 outsideagentcreate_agent(modelChatOpenAI(modelgpt-5.2-chat),tools[get_weather],response_formatToolStrategy(WeatherResponse),)inputs{messages:[(user,What is whether like in Suzhou, and what kind of closing is sugguested?)]}result:dict[str,Any]agent.invoke(inputs)# type: ignoreresponse:WeatherResponseresult.get(structured_response)# type: ignoreprint(fCity:{response.city})print(fTemperature:{response.temperature}°C)print(fSummary:{response.summary})print(fClothing Suggestion:{response.suggestion})通过前面的内容我们知道作为Pregel对象的Agent它默认具有三个输出通道其中包括一个表示结构化输出的structured_response。我们在完成Agent调用后从结果中提取此数据成员得到的就是一个WeatherResponse对象它承载的信息会以如下形式输出City: Suzhou Temperature: 25.0°C Summary: Sunny Clothing Suggestion: Light, breathable clothing such as a T-shirt or blouse with jeans or light trousers. Bring a light jacket if you stay out in the evening.如果我们拦截针对OpenAI API的调用会得到两轮请求/响应。如下所示的第一次调用OpenAI API的请求可以看出在请求中提供了针对两个工具的描述其中一个是我们注册的get_weather,另一个名为WeatherResponse的工具就是LangChain根据注册的结构化输出类型自行创建的。第二个请求携带工具get_weather执行的结果也会包含相同的可用工具列表。{messages:[{content:What is whether like in Suzhou, and what kind of closing is sugguested?,role:user}],model:gpt-5.2-chat,stream:false,tool_choice:required,tools:[{type:function,function:{name:get_weather,description:Get the current weather for a given city.,parameters:{properties:{city:{type:string}},required:[city],type:object}}},{type:function,function:{name:WeatherResponse,description:A structured response format for weather information.,parameters:{properties:{city:{description:City for which the weather is being reported,type:string},temperature:{description:Current temperature in Celsius,type:number},summary:{description:Brief summary of the weather conditions,type:string},suggestion:{description:Clothing suggestion based on the weather,type:string}},required:[city,temperature,summary,suggestion],type:object}}}]}如下所示的是第二次调用OpenAI API得到的响应此时LLM已经得到了get_weather执行的结果纯文本形式然后它利用自身的推理能力根据格式化工具WeatherResponse的描述生成对应的工具调用具体内容体现在choices-message-tool_calls节点中。{choices:[{content_filter_results:{},finish_reason:tool_calls,index:0,logprobs:null,message:{annotations:[],content:null,refusal:null,role:assistant,tool_calls:[{function:{arguments:{\city\:\Suzhou\,\temperature\:25,\summary\:\Sunny\,\suggestion\:\Light, breathable clothing such as a T-shirt or blouse with jeans or light trousers. Bring a light jacket if you stay out in the evening.\},name:WeatherResponse},id:call_z9H4ZDeVJx9FwGgJQEkMncea,type:function}]}}],created:1772802482,id:chatcmpl-DGPCMdPFQc9YwdYo5UmSXoteV1wyQ,model:gpt-5.2-chat-2025-12-11,object:chat.completion,prompt_filter_results:[...],system_fingerprint:null,usage:{...}}ToolStrategy类型定义如下从它的构造函数可知除了指定Schema类型之外我们还可以指定其他参数。虽然用于格式化输出的工具是LangChain自行生成的但是针对它的调用也应该生成一个ToolMessagetool_message_content参数用于指定消息内容。如果模型输出的JSON与Schema不匹配我们可以利用handle_errors参数设置相应的错误处理策略。该参数支持多种形式包括布尔值、字符串、异常类型或自定义函数。如果设置为True则会将错误信息发回给模型让它重试修复JSON。如果设置成字符串系统会向消息历史中追加一条以此为内容的ToolMessage。dataclass(initFalse)classToolStrategy(Generic[SchemaT]):schema:type[SchemaT]|UnionType|dict[str,Any]schema_specs:list[_SchemaSpec[Any]]tool_message_content:str|Nonehandle_errors:(bool|str|type[Exception]|tuple[type[Exception],...]|Callable[[Exception],str])def__init__(self,schema:type[SchemaT]|UnionType|dict[str,Any],*,tool_message_content:str|NoneNone,handle_errors:bool|str|type[Exception]|tuple[type[Exception],...]|Callable[[Exception],str]True,)-None2. ProviderStrategy我们之所以称ToolStrategy是目前最稳定、最通用的实现方式是因为这种格式化输出是采用工具调用方式实现的这是所有LLM都具备的能力。但是很多LLM如OpenAI和Anthropic等自身就具有结构化输出的能力。如果确认使用的模型具有此种能力我们可以利用ProviderStrategy直接利用LLM对输出进行格式化这无疑使更加高效的方式。agentcreate_agent(modelllm,tools[get_weather],response_formatProviderStrategy(WeatherResponse),)对于上面演示的例子我们只需要在调用create_agent函数的时候将response_format参数设置为ProviderStrategy对象同样可以得到几乎一致的结构化输出。但是调用OpenAI API的请求和响应会有所不同。如下所示的使第一次调用的请求可以看出希望输出的格式以JSON Schema的形式被置于请求的response_format节点此节点将会包含在后续的所有请求中。可用工具列表中也不再有用于格式化的工具。{messages:[{content:What is whether like in Suzhou, and what kind of closing is sugguested?,role:user}],model:gpt-5.2-chat,response_format:{type:json_schema,json_schema:{name:WeatherResponse,description:A structured response format for weather information.,strict:false,schema:{properties:{city:{description:City for which the weather is being reported,title:City,type:string},temperature:{description:Current temperature in Celsius,title:Temperature,type:number},summary:{description:Brief summary of the weather conditions,title:Summary,type:string},suggestion:{description:Clothing suggestion based on the weather,title:Suggestion,type:string}},required:[city,temperature,summary,suggestion],type:object}}},stream:false,tools:[{type:function,function:{name:get_weather,description:Get the current weather for a given city.,parameters:{properties:{city:{type:string}},required:[city],type:object,additionalProperties:false},strict:true}}]}如下所示的第二次调用OpenAI API的响应,可以看出它提供的已经是与Schema匹配的结构化JSON内容了。{choices:[{content_filter_results:{...},finish_reason:stop,index:0,logprobs:null,message:{annotations:[],content:{\city\:\Suzhou\,\temperature\:25,\summary\:\Sunny and pleasant\,\suggestion\:\Light clothing such as a T-shirt or blouse with thin pants or a skirt is suitable. You may also want a light jacket for the morning or evening.\},refusal:null,role:assistant}}],created:1772805154,id:chatcmpl-DGPtSmO45cggzuhJvfd9id0MZNNyQ,model:gpt-5.2-chat-2025-12-11,object:chat.completion,prompt_filter_results:[...],system_fingerprint:null,usage:{...}}ProviderStrategy定义如下其构造函数除了表示Schema类型的参数之外还有一个布尔类型的参数strict。目前它主要针对的是OpenAI的Strict Mode开关。如果设为TrueOpenAI会保证输出100%与指定的Schema匹配。但是带来的代价生成的第一个Token会有明显的延迟预处理Schema且Schema中不能使用一些高级校验如自定义正则表达式。dataclass(initFalse)classProviderStrategy(Generic[SchemaT]):schema:type[SchemaT]|dict[str,Any]schema_spec:_SchemaSpec[SchemaT]def__init__(self,schema:type[SchemaT]|dict[str,Any],*,strict:bool|NoneNone,)-Nonedefto_model_kwargs(self)-dict[str,Any]3. AutoStrategy如果LLM自身就具有结构化输出的能力那么使用ProviderStrategy一般是更好的选择。倘若我们不清楚使用的LLM是否支持结构化输出或者会随时切换LLM我们可以使用AutoStrategy。它会自动检测当前LLM的能力并选择最优的结构化输出路径。如果在调用create_agent函数时将response_format参数直接指定为Pydantic类型或者表示JSON Schema的字典相当于选择了AutoStrategy。classAutoStrategy(Generic[SchemaT]):schema:type[SchemaT]|dict[str,Any]def__init__(self,schema:type[SchemaT]|dict[str,Any],)-NoneAutoStrategy背后的逻辑是优先选择ProviderStrategy如果当前使用的模型如OpenAI、Anthropic的最新版本支持原生的结构化输出功能AutoStrategy会直接利用LLM产生希望的结构化输出这种方式通常最稳定、解析率最高回退至ToolStrategy如果模型不支持原生结构化输出但支持工具调用它会将结构化模式转换为一个内部的隐形工具引导模型通过调用工具的方式输出符合要求的结构化数据。使用AutoStrategy能够带来如下的好处跨模型兼容性开发者无需为OpenAI编写一套逻辑再为本地模型如Ollama运行的Llama编写另一套逻辑。AutoStrategy屏蔽了不同模型供应商在结构化数据实现上的差异简化代码你只需定义好Pydantic数据结构直接传给response_format即可系统会自动完成复杂的绑定工作动态支持在中间件中我们可以利用AutoStrategy根据运行时的上下文动态调整输出模式。