
模型评测人工智能大模型AI 评测【免费下载链接】opencompassOpenCompass is an LLM evaluation platform, supporting a wide range of models from OpenAI, Anthropic, Gemini, Qwen, GLM, DeepSeek, etc, across 100 datasets covering knowledge, reasoning, coding, science, language, long-context, and safety.项目地址https://gitcode.com/gh_mirrors/op/opencompass点击查看免费下载导读本文围绕 OpenCompass 新增的RawPromptTemplate模板解析类展开说明如何告别PromptTemplate中begin/round/end与role/fallback_role的映射转换直接以 OpenAI 兼容的messages列表定义输入、配置数据集推理并附加模型侧信息。读完本文你将掌握RawPromptTemplate的语法与校验规则、在客观数据集评测中的实战迁移方法、Few-shot 与动态展开能力以及在 API 模型配置中通过meta_template注入额外系统提示的实现路径。一、为什么要引入 RawPromptTemplateOpenCompass 此前在推理配置中主要使用PromptTemplate类opencompass/openicl/icl_prompt_template.py。它通过template字典中的begin、round、end三个分区来描述对话结构每个分区内的条目再由role、fallback_role、prompt等字段构成最后经PromptList与APITemplateParseropencompass/models/base_api.py多层转换才能得到最终喂给模型的 message 序列。这种间接映射虽然灵活但在面向 OpenAI 对话格式[{role: ..., content: ...}, ...]的场景下显得繁琐且不够直观。为此OpenCompass 新增了RawPromptTemplate模板解析类opencompass/openicl/icl_raw_prompt_template.py其核心设计目标有三点输入配置与最终 message 一一对应你在配置里写出的每一条{role: ..., content: ...}就是最终发往模型的同构消息不再存在映射转换更贴合 OpenAI 对话格式原生支持system、user、assistant三种角色天然适配 GPT、Claude、Qwen、GLM、DeepSeek 等 API 对话模型支持在模型配置中追加额外 Prompt 内容配合 API 模型类的meta_template可在推理输入之外附加额外的系统/用户提示。从源码看RawPromptTemplate通过ICL_PROMPT_TEMPLATES.register_module()注册到模板注册表prompt_type被标记为raw_messages底层检索器会依据该标记走独立的 ICE 生成路径见 opencompass/openicl/icl_retriever/icl_base_retriever.py直接把生成的 messages 列表原样返回这正是“绕过多层转换、直接透传”的实现基础。二、从 PromptTemplate 迁移修改数据集的推理配置2.1 迁移前PromptTemplate 格式基于PromptTemplate的原推理配置格式示例如下infer_cfg dict( prompt_templatedict( typePromptTemplate, templatedict( begin[ dict( roleSYSTEM, fallback_roleHUMAN, promptYou are a helpful assistant., ) ], round[ dict( roleHUMAN, prompt{problem}\nRemember to put your final answer within \\boxed{}., ), ], ), ), retrieverdict(typeZeroRetriever), inferencerdict(typeGenInferencer), )可以看到这里需要显式声明begin系统开场白与round对话轮次并区分role与fallback_role角色不可用时降级到 HUMAN配置经过编码后才会形成带section/pos标记的PromptList对应PromptTemplate._encode_template的实现见 opencompass/openicl/icl_prompt_template.py。2.2 迁移后RawPromptTemplate 格式进行如下修改即可调整为RawPromptTemplate格式infer_cfg dict( prompt_templatedict( typeRawPromptTemplate, messages [ {role: system, content: You are a helpful assistant.}, {role: user, content: {problem}\nRemember to put your final answer within \\boxed{}.}, ], ), retrieverdict(typeZeroRetriever), inferencerdict(typeGenInferencer), )对比可见begin分区的系统提示被写成一条rolesystem的 messageround分区中的提问被写成roleuser的 message{problem}等数据集字段占位符依然保留在content中并在生成阶段被替换。配置写法与最终发给模型的输入在结构上完全一致。三、RawPromptTemplate 的输入校验与能力边界RawPromptTemplate.__init__的完整签名来自 opencompass/openicl/icl_raw_prompt_template.py为def __init__( self, messages: List[MessageElementType], format_variables: bool True, ice_token: str /E, ) - None:其中关键参数说明如下参数类型默认值说明messagesList[Union[Dict, str]]必填OpenAI 格式消息列表支持三种元素类型见下表format_variablesboolTrue是否对content中的{变量}执行safe_format替换ice_tokenstr/EICE 占位符标识messages中与其相等的字符串元素会被替换为 ICE 内容messages支持的三种元素类型源码 docstring 明确给出Dict 标准消息如{role: user, content: ...}是最基本的类型Dict 带expand_column字段动态展开占位符从数据集某列读取List[Dict]并展开为多条消息例如{expand_column: history}会把entry[history]中的多轮对话记录原样插入可用于多轮对话场景str 字符串元素作为 ICE 占位符即ice_token例如/E在 Few-shot 推理时被替换为检索到的示例消息列表。3.1 严格校验规则_validate_messagesopencompass/openicl/icl_raw_prompt_template.py在初始化时即执行严格校验messages必须是list否则抛出TypeError: messages must be a list, got ...每个元素必须是dict或str否则抛出TypeError: messages[i] must be a dict or strDict 元素若包含expand_column字段则跳过后续校验动态展开消息无需role/content其余 Dict 元素必须同时包含role与content两个键缺失任一键会分别抛出ValueError: messages[i] missing role key与ValueError: messages[i] missing content keyrole必须是system、user、assistant三者之一否则抛出ValueError: messages[i] has invalid role。上述规则均有对应单元测试覆盖见 tests/openicl/test_raw_prompt_template.pytest_validation_not_list、test_validation_missing_role、test_validation_invalid_role等用例。3.2 变量替换与接口兼容当format_variablesTrue默认时generate_item会对每条 Dict 消息的content执行safe_format替换把{input}、{problem}等数据集字段替换为实际取值当format_variablesFalse时占位符原样保留对应测试test_generate_item_no_formatgenerate_item内部对所有消息做copy.deepcopy因此不会修改原始messages配置对应测试test_generate_item_does_not_modify_originalgenerate_ice_item与generate_label_prompt_item作为兼容接口存在其中字符串 ICE 占位符在生成 ICE 时被忽略若messages中存在与ice_token相等的字符串元素且ice_field_replace_token传入的是消息列表则会将 ICE 内容展开插入到该位置对应generate_item中的result.extend(ice_field_replace_token)分支。四、仓库内已集成的 rawprompt 数据集配置官方已为常用客观数据集添加了文件名包含rawprompt的新配置可直接参考或复用AIME2026opencompass/configs/datasets/aime2026/aime2026_cascade_eval_rawprompt_gen_0970dd.py推理侧使用RawPromptTemplate构造仅含 user 消息的提问并用CascadeEvaluator规则 LLM 评判完成答案判定MMLU Proopencompass/configs/datasets/mmlu_pro/mmlu_pro_0shot_nocot_genericllmeval_rawprompt_gen_0321fb.py在循环构造各 category 数据集时推理模板只保留roleuser的QUERY_TEMPLATE评判模板则由system评判者设定与user评分指令两条消息组成。4.1 AIME2026 配置拆解aime2026_reader_cfg dict(input_columns[problem], output_columnanswer) aime2026_infer_cfg dict( prompt_templatedict( typeRawPromptTemplate, messages [ {role: user, content: {problem}\nRemember to put your final answer within \\boxed{}.}, ], ), retrieverdict(typeZeroRetriever), inferencerdict(typeGenInferencer), )该配置没有 system 消息直接以 user 消息提问{problem}由reader_cfg中的input_columns[problem]提供。评判阶段则复用RawPromptTemplate构造裁判模板cascade_evaluator dict( typeCascadeEvaluator, rule_evaluatordict(typeMATHVerifyEvaluator), llm_evaluatordict( typeGenericLLMEvaluator, prompt_templatedict( typeRawPromptTemplate, messages[ {role: system, content: You are a helpful assistant who evaluates the correctness and quality of models outputs.}, {role: user, content: GRADER_TEMPLATE}, ], ), ... ), )这一示例说明RawPromptTemplate不仅服务于待测模型的推理输入也同样被GenericLLMEvaluator这类 LLM 评判器用于构造裁判系统提示与评分指令覆盖“生成—评判”全链路。4.2 MMLU Pro 配置拆解mmlu_pro_infer_cfg dict( prompt_templatedict( typeRawPromptTemplate, messages[ {role: user, content: QUERY_TEMPLATE}, ], ), retrieverdict(typeZeroRetriever), inferencerdict(typeGenInferencer), )QUERY_TEMPLATE中通过{question}、{options_str}等占位符引入数据集字段且无需预先拆分成多轮消息一条 user 消息即可承载完整指令与题目。此外仓库中还有大量同样命名规约的配置例如 opencompass/configs/datasets/aime2025/aime2025_cascade_eval_rawprompt_gen_2f2c96.py、opencompass/configs/datasets/gpqa/gpqa_cascade_eval_rawprompt_gen_706039.py、opencompass/configs/datasets/gsm8k/gsm8k_cascade_eval_rawprompt_gen_36fce7.py 等覆盖数学、科学、代码、长文本等类别主观及多轮对话相关数据集如 opencompass/configs/datasets/subjective/ 下多个*_rawprompt.py评判配置也已在陆续接入。五、Few-shot 与多轮历史动态展开实践RawPromptTemplate在 Few-shot 场景下的典型写法是把ice_token默认/E作为字符串元素插入消息列表配合带ice_token的PromptTemplate一并使用。其底层逻辑位于 opencompass/openicl/icl_retriever/icl_base_retriever.py当检测到prompt_type raw_messages时检索器会对每个检索到的索引调用ice_template.generate_ice_item(...)生成消息并拼接generated_ice.extend(ice_messages)最终以消息列表形式注入。多轮对话历史则推荐使用expand_column动态展开。例如prompt_templatedict( typeRawPromptTemplate, messages[ {role: system, content: You are a helpful assistant.}, {expand_column: history}, # 从数据集 history 列读取多轮对话并展开 {role: user, content: {question}}, {role: assistant, content: }, ], )generate_item对expand_column的处理逻辑是若entry[col]存在且为列表则通过copy.deepcopy展开插入否则跳过见 opencompass/openicl/icl_raw_prompt_template.py。六、在模型配置侧附加额外信息meta_template 的列表用法6.1 配置格式使用OpenAI、OpenAISDK、OpenAISDKStreaming等 API 对话模型类进行评测时可通过模型配置中的meta_template来附加所需的额外信息格式如下dict( abbrYOUR_MODEL, typeOpenAISDK, pathYOUR_MODEL, keyYOUR_API_KEY, openai_api_baseYOUR_API_BASE, meta_template[ {content: Your extra system prompt here., role: system}, {content: Your extra user prompt here., role: user}, ], query_per_second1, batch_size8, temperature1.0, max_out_len32768, max_seq_len32768, ... )注意这里的meta_template是列表形式每条为{role: ..., content: ...}与旧式dict(round[...])的元模板结构不同——列表形式与RawPromptTemplate的 messages 风格保持一致。6.2 底层合并机制OpenAISDKopencompass/models/openai_api.py继承自OpenAI基类其meta_template最终交给APITemplateParser处理。当输入消息列表与列表形式的meta_template同时存在时opencompass/models/base_api.py 中的合并逻辑为若两条消息在相同位置、相同角色则合并meta_template的内容拼接在当前 prompt 内容之前meta[content] prompt[content]若meta_template中的角色在剩余 prompt 中不存在则将该 meta 消息原样插入若角色在后续消息中才出现则先保留当前 prompt 消息继续向后推进所有 meta 消息处理完后剩余 prompt 消息追加到末尾。因此meta_template[{content: Your extra system prompt here., role: system}]会把额外系统提示合并进最终的 system 消息roleuser的额外内容则按上述规则插入或合并到对应位置。这种机制让用户可以在不修改数据集推理配置的前提下为 API 评测统一注入系统指令如安全约束、输出格式要求、身份设定等。同时APITemplateParser.__init__对列表形式有校验要求每个元素都是含role的 dict且role属于system/user/assistantopencompass/models/base_api.py与RawPromptTemplate的角色白名单一致。七、实测验证与调试建议单元测试入口tests/openicl/test_raw_prompt_template.py 覆盖了初始化、校验、变量替换、format_variablesFalse、原始消息不被修改、ICE/多轮消息生成、__repr__等 12 个用例可作为迁移配置后的自测参照。运行方式在仓库根目录python -m pytest tests/openicl/test_raw_prompt_template.py -v调试建议先用ZeroRetriever 单条 user 消息跑通最小推理链路再逐步增加 system 消息、Few-shot 与expand_column若出现ValueError: messages[i] has invalid role检查是否使用了HUMAN/SYSTEM等旧式角色名——RawPromptTemplate只接受小写的system/user/assistant若{占位符}未被替换确认format_variables未设为False且占位符对应的字段确实存在于数据集的input_columns或数据条目中配置文件名建议遵循仓库既有规约如xxx_rawprompt_gen_hash.py便于在opencompass/configs/datasets/下统一检索与复用。八、小结RawPromptTemplate让 OpenCompass 的 Prompt 构建路径更贴近现代 API 模型的对话习惯数据集侧用messages列表直接定义输入模型侧用列表形式的meta_template注入额外信息两层配置共同构成最终请求中间不再有begin/round/end与fallback_role的隐式映射。其严格校验、expand_column动态展开、Few-shot ICE 透传等特性使它在客观题、LLM 评判、多轮对话与长上下文等评测场景中都有直接可用的落地配置仓库内大量rawprompt命名的数据集配置与 tests/openicl/test_raw_prompt_template.py 测试即为最佳实践参照。赞分享模型评测人工智能大模型AI 评测【免费下载链接】opencompassOpenCompass is an LLM evaluation platform, supporting a wide range of models from OpenAI, Anthropic, Gemini, Qwen, GLM, DeepSeek, etc, across 100 datasets covering knowledge, reasoning, coding, science, language, long-context, and safety.项目地址https://gitcode.com/gh_mirrors/op/opencompass点击查看免费下载相关推荐OpenCompass 提示模板构建进阶RawPromptTemplate 直通 OpenAI Chat 消息格式实战指南OpenCompass 提示模板构建进阶RawPromptTemplate 直通 OpenAI Chat 消息格式实战指南 导读 在 OpenCompass模型评测人工智能大模型AI 评测Verified-Smart-Contracts项目进阶教程如何快速使用eDSL编写EVM级规范Verified Smart Contracts项目进阶教程如何快速使用eDSL编写EVM级规范 Verified Smart Contracts是一个专注于Open Interpreter消息转换OpenAI格式兼容指南Open Interpreter消息转换OpenAI格式兼容指南 引言跨平台AI交互的兼容性挑战 在AI驱动的应用开发中不同大语言模型LLM提供商采用人工智能大模型AI Agent代码智能体AI 应用CLI上一篇开发者必备Chrome DevTools高级功能与隐藏技巧大全下一篇终极指南5步彻底解决Wireshark编译缓存问题创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考