
原文New in llama.cpp: Decision Models发布日期2026 年 10 月 2 日作者Xuan-Son Nguyen (ngxson)、Victor Mustar、ggml-orgllama.cpp 服务器现在通过/v1/systemone端点支持决策模型。你发送一个状态文本、JSON、截图和类型化问题模型在单次前向传播中返回每个选项的概率。API 遵循 TypeSafe 的 Jev 模型引入的 System One 格式因此现有客户端只需更换 base URL 即可。实现细节见 PR #29818。什么是决策模型决策模型通过为你给定的选项打分来回答问题而不是生成文本。聊天模型每个输出 token 需要一次前向传播且输出仍需解析。决策模型只读取一次输入其答案始终是你给出的选项之一并附带概率。典型用途包括路由请求、内容审核、检查 Agent 步骤是否成功、或选择下一步动作。支持的模型模型大小基于语言图片许可证速度*Julia-1144MmmBERT-small50否Apache 2.03 msLaya421MModernBERT-large英语否Apache 2.05 msKev-4B4BQwen3.5-4B-Base英语否Apache 2.012 mslev4BQwen3.5-4B英语否Apache 2.036 msOpenJev27BQwen3.8-27Ben, de, fr, hi, zh, ja是CC BY-NC 4.043 msClef27BQwen3.8-27B英语是Apache 2.0—* 在单块 NVIDIA RTX PRO 6000 上回答一个问题的中位时间。在 Decision models 合集 中查找这些模型更多模型即将推出。社区 Decision Index 展示了它们的对比。快速开始从 llama.app 获取最新版 llama.cpp或运行llama update然后启动模型llama serve-hfggml-org/Kev-4B-GGUF一个请求包含一个状态和一个或多个问题。问题有三种类型类型你发送你得到choice选项含可选描述最佳选项 每个选项的概率score2 到 10 个等级从低到高期望等级可以落在两个等级之间noul是/否问题是yes的概率发送包含状态和问题的请求curlhttp://localhost:8080/v1/systemone\-HContent-Type: application/json\-d{ state: Customer message: I was charged twice for my order last week and nobody has replied., questions: { route: { type: choice, instructions: Which team should handle this?, criteria: { billing: payments, charges, refunds, invoices, shipping: delivery, tracking, lost or late parcels, technical: bugs, errors, login problems } }, angry: { type: noul, instructions: Is the customer angry? }, urgency: { type: score, instructions: How urgent is this?, criteria: [can wait, this week, today, right now] } } }响应数值已四舍五入{model:ggml-org/Kev-4B-GGUF,answers:{route:{type:choice,choice:billing,probabilities:{billing:0.9049,shipping:0.0275,technical:0.0676},confidence:0.8574},angry:{type:noul,noul:0.8208},urgency:{type:score,score:2.2821,legend:{0:can wait,1:this week,2:today,3:right now},probabilities:{0:0.036,1:0.1937,2:0.2225,3:0.5478},confidence:0.2821}},usage:{input_tokens:130,output_tokens:0}}完整参考见 服务器文档。图片某些模型目前为 OpenJev还可以读取图片如文档或截图。视觉投影器会自动下载llama serve-hfggml-org/OpenJev-GGUF例如对上传的文档进行分类importbase64importrequestswithopen(document.png,rb)asf:imagedata:image/png;base64,base64.b64encode(f.read()).decode()responserequests.post(http://localhost:8080/v1/systemone,json{state:A file uploaded by a customer.,images:[image],questions:{kind:{type:choice,instructions:What kind of document is this?,criteria:{invoice:None,receipt:None,contract:None,other:None},},},})print(response.json()[answers][kind][choice])# invoicestate也可以是聊天消息列表。任何image_url部分data URL都会被当作图片读取与聊天补全相同。多模型单服务器在路由模式下模型按需加载你可以在每个请求中选择一个llama servecurlhttp://localhost:8080/v1/systemone\-HContent-Type: application/json\-d{model: ggml-org/Julia-1-GGUF:Q8_0, state: ..., questions: {...}}/v1/models列出所有 id。当只加载了一个模型时model字段会被忽略。使用技巧尝试不同大小的模型。小模型更快大模型知识更丰富。Decision Index 对它们进行了比较。描述你的选项。Julia-1 在仅有标签时将我被扣了两次钱路由到shipping而在每个选项都有描述时路由到billing0.99。为每个模型选择置信度阈值。常见的模式是对有信心的答案采取行动其余的发送给人工。正确的阈值取决于模型一个模糊的工单“Hi, quick question about my account”在 Julia-1 上得分 0.25而在 Kev-4B 上得分 0.80。在选择阈值之前请在你自己的示例上进行测试。批量处理问题。问题是独立回答的Kev-4B、lev 和 OpenJev 只处理一次状态。尝试不同的量化。与任何 GGUF 一样这些模型有多种精度可选例如llama serve -hf ggml-org/Kev-4B-GGUF:Q8_0。