系统全解析:以 LLM 为控制器编排 Hugging Face 专家模型)
人工智能大模型AI Agent工具调用模型评测【免费下载链接】JARVISJARVIS, a system to connect LLMs with ML community. Paper: https://arxiv.org/pdf/2303.17580.pdf项目地址https://gitcode.com/gh_mirrors/jarvis3/JARVIS点击查看免费下载本文以 JARVIS 项目主仓库的 README.md 为骨架系统讲解其“LLM 控制器 专家模型执行器”的协作架构、四阶段工作流、四种运行模式Server / Web / Gradio / CLI、Web API 接口与核心配置项并结合仓库源码awesome_chat.py、models_server.py、config.default.yaml给出可复制、可运行的实操方案。读完本文你将掌握从零部署 JARVIS、按需选择本地 / HuggingFace / 混合推理模式、调用其 REST API以及在嵌入式设备NVIDIA Jetson上运行该系统的完整方法。项目定位与核心思想JARVIS 的使命是探索通用人工智能AGI并把前沿研究成果带给整个社区。其论文《HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in HuggingFace》提出了一种协作式系统范式LLM 作为控制器Controller负责理解用户意图、拆解任务、选择模型、汇总结果大量专家模型作为协作执行器Collaborative Executors来自 HuggingFace Hub 的各类预训练模型负责具体执行文本、图像、音频、视频等子任务。其核心设计理念可概括为一句话语言是 LLM 连接海量 AI 模型的接口通过语言把复杂 AI 任务拆解给擅长各自领域的专家模型再由 LLM 统一编排。系统级架构如下图所示四阶段工作流JARVIS 的工作流由四个阶段构成这也是整个系统最核心的运行机制阶段名称职责Stage #1Task Planning任务规划用 ChatGPT 分析用户请求、理解意图并将其拆解为可求解的任务列表Stage #2Model Selection模型选择根据任务描述从 Hugging Face 上选择合适的专家模型Stage #3Task Execution任务执行调用并执行每个被选中的模型把结果返回给 ChatGPTStage #4Response Generation响应生成由 ChatGPT 整合所有模型的预测结果生成最终回答源码中的四阶段印证从源码结构看四个阶段在 awesome_chat.py 中都有对应的独立函数与提示词模板阶段 1 对应parse_task()第 317 行起使用config[tprompt][parse_task]作为系统提示词要求 LLM 输出严格 JSON 格式的任务数组每个任务包含task、id、dep依赖任务 id 列表、args只能是text/image/audio三种键若无法解析则返回空 JSON[]阶段 2 对应choose_model()第 350 行起基于任务元信息metas让 LLM 输出{id: ..., reason: ...}的严格 JSON阶段 4 对应response_results()第 377 行起要求 LLM 描述执行过程与推理结果三个阶段之间由chat_huggingface()第 891 行起串联先解析任务再通过unfold()/fix_dep()处理任务依赖与展开然后用多线程并发执行无依赖关系的任务最后汇总结果生成回复。此外get_avaliable_models()第 678 行起会在模型选择前并发探测候选模型在本地端点与 HuggingFace 端点上的可用状态并按num_candidate_models限制候选数量。系统要求与两种部署形态默认形态推荐完整本地部署对应 config.default.yaml需要完整下载并本地部署专家模型Ubuntu 16.04 LTSVRAM ≥ 24GBRAM12GBminimal、16GBstandard、80GBfull磁盘 284GB其中42GBdamo-vilab/text-to-video-ms-1.7b文生视频126GBControlNet各类 control 模型66GBstable-diffusion-v1-550GB其余模型轻量形态Lite零本地模型对应 config.lite.yamlUbuntu 16.04 LTS除此之外不需要任何额外硬件要求Lite 配置不会下载、部署任何本地专家模型JARVIS 将完全依赖 HuggingFace Inference Endpoints远程推理端点上稳定运行的那些模型。代价是可用模型集合会受远程端点稳定性限制。快速开始环境与密钥准备开始前需要替换 config.default.yaml 中的两项密钥openai.api_key替换为你的个人 OpenAI Keyhuggingface.token替换为你的 Hugging Face Token在 Hugging Face 账户 Settings → Tokens 处生成以hf_开头。也可以不修改配置文件直接设置环境变量export OPENAI_API_KEYsk-... export HUGGINGFACE_ACCESS_TOKENhf_...从源码看awesome_chat.py密钥校验逻辑为配置文件中的 key 优先其次读取环境变量两者都无效则直接抛错退出。同时代码还支持azure配置块见 config.azure.yaml含api_key、base_url、deployment_name、api_versionAPI 类型优先级为dev本地 LLM 端点 azure openai。环境搭建conda pipcd hugginggpt/server conda create -n jarvis python3.8 conda activate jarvis conda install pytorch torchvision torchaudio pytorch-cuda11.7 -c pytorch -c nvidia pip install -r requirements.txt下载模型仅 local / hybrid 模式需要cd hugginggpt/server/models bash download.sh # 需要先安装 git-lfsdownload.sh 会逐个git clone --recurse-submodules或git pull git lfs pull拉取近 30 个模型与Matthijs/cmu-arctic-xvectors数据集包括图像描述nlpconnect/vit-gpt2-image-captioning、ControlNet 系列、stable-diffusion-v1-5、文生视频damo-vilab/text-to-video-ms-1.7b、语音whisper-base、speecht5_*、espnet/kan-bayashi_ljspeech_vits、检测分割detr-resnet-*、maskformer-*、owlvit-base-patch32等。Windows 用户可改用同目录下的 download.ps1。四种运行模式1. Server 模式部署 Web APIcd hugginggpt/server python models_server.py --config configs/config.default.yaml # 仅在 inference_mode 为 local 或 hybrid 时需要 python awesome_chat.py --config configs/config.default.yaml --mode server # LLM 默认 text-davinci-003启动后即可通过 Web API 访问 JARVIS 服务默认监听0.0.0.0:8004见 awesome_chat.py 的server()函数接口方法返回内容/hugginggptPOST完整服务任务规划 模型选择 执行 响应生成/tasksPOSTStage #1 的中间结果任务规划/resultsPOSTStage #1-3 的中间结果任务规划 模型选择与执行结果一个请求示例基于/examples/d.jpg的姿态与/examples/e.jpg的内容生成新图像curl --location http://localhost:8004/tasks \ --header Content-Type: application/json \ --data { messages: [ { role: user, content: based on pose of /examples/d.jpg and content of /examples/e.jpg, please show me a new image } ] }返回的任务规划 JSON[{args:{image:/examples/d.jpg},dep:[-1],id:0,task:openpose-control},{args:{image:/examples/e.jpg},dep:[-1],id:1,task:image-to-text},{args:{image:GENERATED-0,text:GENERATED-1},dep:[1,0],id:2,task:openpose-text-to-image}]这段返回清晰展示了任务依赖机制任务 0、1 无依赖dep: [-1]任务 2 依赖任务 1图像转文本与任务 0姿态检测的输出通过GENERATED-dep_id占位符引用上游生成资源。示例图片位于 hugginggpt/server/public/examples。从源码看三个路由都从请求 JSON 的messages中取对话内容并支持在请求体中覆盖api_key/api_type/api_endpoint未提供时回退到配置文件服务由 waitress 承载。2. Web 模式浏览器交互界面仓库提供了用户友好的 Web 前端Vue Vite位于 web 目录。在 Server 模式启动后运行以下命令即可在浏览器中与 JARVIS 对话cd web npm install npm run dev注意事项需要先安装nodejs与npm重要如果 Web 客户端运行在另一台机器上需要把http://{服务器局域网IP}:{端口}/设置到 web/src/config/index.ts 的HUGGINGGPT_BASE_URL若要使用视频生成功能需要手动编译带 H.264 编码的 ffmpeg并用如下命令验证安装成功# 可选安装 ffmpeg此命令需无报错执行 LD_LIBRARY_PATH/usr/local/lib /usr/local/bin/ffmpeg -i input.mp4 -vcodec libx264 output.mp4双击设置图标可以切换到 ChatGPT 直接对话。3. Gradio 模式交互式 DemoGradio demo 已托管在 Hugging Face Space 上使用inference_modehybrid、local_deploymentstandard构建。本地启动方式python models_server.py --config configs/config.gradio.yaml python run_gradio_demo.py --config configs/config.gradio.yaml也可以把 HF Space 作为 Docker 镜像运行docker run -it -p 7860:7860 --platformlinux/amd64 registry.hf.space/microsoft-hugginggpt:latest python app.py4. CLI 模式终端直接对话CLI 模式无需启动模型服务器是最轻量的体验方式cd hugginggpt/server python awesome_chat.py --config configs/config.default.yaml --mode cli启动后即可输入请求输入exit退出。下面是从仓库 README 原样保留的三组典型对话实录示例 1跨图像的目标计数[ User ]: Given a collection of image A: /examples/a.jpg, B: /examples/b.jpg, C: /examples/c.jpg, please tell me how many zebras in these picture? [ Jarvis ]: Based on your request, I have used several models to process the images you provided. For the image A: /examples/a.jpg, I used the image-to-text model nlpconnect/vit-gpt2-image-captioning to generate the text a cat sitting on a window sill looking out. Then I used the object-detection model facebook/detr-resnet-50 to generate the image with predicted box /images/a3dc.jpg, which contains two objects with labels potted plant and cat. For the image B: /examples/b.jpg, I used the image-to-text model nlpconnect/vit-gpt2-image-captioning to generate the text two zebras standing in a field of dry grass. Then I used the object-detection model facebook/detr-resnet-50 to generate the image with predicted box /images/f5df.jpg, which contains three objects with labels zebra. For the image C: /examples/c.jpg, I used the image-to-text model nlpconnect/vit-gpt2-image-captioning to generate the text a zebra standing in a field of tall grass. Then I used the object-detection model facebook/detr-resnet-50 to generate the image with predicted box /images/69c7.jpg, which contains one object with label zebra. Finally, I used the canny-control model canny-control to generate the image /images/7e18.png based on the image /examples/f.jpg. Therefore, there are four zebras in these pictures. Is there anything else I can help you with?示例 2命名实体识别 问答[ User ]: Please answer all the named entities in the sentence: Iron Man is a superhero appearing in American comic books published by Marvel Comics. The character was co-created by writer and editor Stan Lee, developed by scripter Larry Lieber, and designed by artists Don Heck and Jack Kirby. [ Jarvis ]: Sure, I understand your request. Based on the inference results of the models, I will answer your request step by step. Firstly, I used the token-classification model dslim/bert-base-NER to recognize the named entities in the sentence. The inference result is that the named entities are Iron Man, Stan Lee, Larry Lieber, Don Heck and Jack Kirby. Then, I used the question-answering model deepset/roberta-base-squad2 to answer your request. The inference result is that there is no output since the context cannot be empty. Therefore, I cannot make it. I hope this answer your request. Is there anything else I can help you with?示例 3Canny 图生成[ User ]: Please generate a canny image based on /examples/f.jpg [ Jarvis ]: Sure. I have generated a canny image based on /examples/f.jpg. To do this, I first used the image-to-text model nlpconnect/vit-gpt2-image-captioning to generate the text description of the image, which is a herd of giraffes and zebras grazing in a field. Then I used the canny-control model to generate the canny image of the original image, which is stored in the path /images/1ea2.png. Finally, I used the canny-text-to-image model lllyasviel/sd-controlnet-canny to generate the canny image based on the text description and the canny image, which is stored in the path /images/ba66.png. Do you need anything else?CLI 模式对应的实现位于 awesome_chat.py 的cli()函数维护消息历史循环读取用户输入并调用chat_huggingface()。仓库同时内置了--mode test模式test()函数内置了单轮与多轮的测试用例便于快速验证部署。配置详解config.default.yaml服务端配置文件为 config.default.yaml核心参数如下参数可选值 / 默认值说明modeltext-davinci-003当前支持text-davinci-003、gpt-4LLM 控制器仓库表示后续会支持更多开源 LLMuse_completiontrue/falsetrue使用/v1/completionsfalse使用/v1/chat/completionsinference_modelocal/huggingface/hybrid推理端点模式仅本地、仅 HuggingFace免本地推理端点、两者混合推荐hybridlocal_deploymentminimal/standard/full本地部署模型规模仅在local/hybrid下生效详见下文devicecuda:0/cpu模型运行设备num_candidate_models默认5模型选择阶段的最大候选数量max_description_length默认100模型描述截断长度proxy可选如http://ip:port代理服务器用于访问外网 APIhttp_listenhost: 0.0.0.0、port: 8004awesome_chat Web 服务监听地址若用 Web 客户端需将http://{服务器局域网IP}:{port}/配置到web/src/config/index.tslocal_inference_endpointhost: localhost、port: 8005本地模型服务器监听地址logit_biasparse_task: 0.1、choose_model: 5对任务解析与模型选择的关键 token 施加的 logit 偏置log_filelogs/debug.log调试日志路径debug: true时生效local_deployment三档规模minimalRAM 12GB仅部署 ControlNetstandardRAM 16GBControlNet 标准流水线标准流水线指 models_server.py 中standard_pipes注册的模型如whisper-base、detr-resnet-101、owlvit-base-patch32、vilt-b32-finetuned-vqa、layoutlm-document-qa、vit-gpt2-coco-en、dpt-large等fullRAM 42GB全部注册模型other_pipes中的文生视频、stable-diffusion-v1-5、maskformer-swin-base-coco、dpt-hybrid-midas、语音相关模型等。从 models_server.py 的load_pipes()源码可以看到三档规模的装载逻辑minimal只加载 ControlNet 及其 control 模型standard在此基础上叠加标准流水线full再叠加其余模型最终合并为pipes字典对外提供推理。个人笔记本推荐配置在个人笔记本上README 推荐inference_mode: hybridlocal_deployment: minimal的组合。需要说明的是该组合下可用的模型集合可能受限因为远程 HuggingFace Inference Endpoints 存在不稳定性。提示词模板tprompt/prompt/demos_or_presteps配置文件中还包含三组可调优的提示词与示例tprompt.parse_task任务规划阶段的系统提示词明确规定了任务的 JSON schema、GENERATED-dep_id依赖占位符规则以及 36 种可选任务类型覆盖文本分类、翻译、问答、目标检测、文生图/文生视频、ControlNet 系列、语音系列等并要求“解析出的任务越少越好、注意依赖关系、无法解析则返回空 JSON[]”tprompt.choose_model模型选择阶段提示词强调聚焦模型描述、并优先选择有本地推理端点的模型更快更稳tprompt.response_results响应生成阶段提示词要求 LLM 基于推理结果直接作答并详细描述工作流、所用模型、推理结果文件路径demos_or_presteps对应三个阶段的 few-shot 示例文件即 demo_parse_task.json、demo_choose_model.json、demo_response_results.jsonprompt.*各阶段拼入用户消息的动态提示词模板{{input}}、{{context}}、{{metas}}、{{task}}、{{processes}}等占位符由replace_slot()在运行时替换。NVIDIA Jetson 嵌入式设备支持仓库根目录提供了 Dockerfile.jetson为 NVIDIA Jetson 嵌入式设备提供实验性支持。该镜像内置了加速版 ffmpeg、pytorch、torchaudio 与 torchvision 依赖。构建前需确认 Docker 默认运行时已设置为nvidia。构建命令docker build --pull --rm -f Dockerfile.jetson -t toolboc/nv-jarvis:r35.2.1受内存限制JARVIS 要求运行在 Jetson AGX Orin 系列设备上优先 64G 板载内存版本且配置需要设为inference_mode: locallocal_deployment: standard模型与配置文件推荐通过卷挂载从宿主机传入容器如示例中的-v ~/jarvis/configs:/app/server/configs与-v ~/src/JARVIS/server/models:/app/server/models。也可以取消注释 Dockerfile 中# Download local models一节构建一个内置模型的镜像。在 Jetson Orin AGX 上启动模型服务器、awesomechat 与 Web 应用# 运行容器容器会自动启动模型服务器 docker run --name jarvis --nethost --gpus all -v ~/jarvis/configs:/app/server/configs -v ~/src/JARVIS/server/models:/app/server/models toolboc/nv-jarvis:r35.2.1 # 等待模型服务器完成初始化 # 启动 awesome_chat.py docker exec jarvis python3 awesome_chat.py --config configs/config.default.yaml --mode server # 启动 Web 应用应用将可通过 http://localhost:9999 访问 docker exec jarvis npm run dev --prefix/app/web衍生研究项目仓库 README 的 Whats New 记录了三个相互关联的研究成果EasyTool2024.01 发布用于更简便的工具使用代码与数据集位于 easytool 目录对应论文《EasyTool: Enhancing LLM-based Agents with Concise Tool Instruction》包含funcQA、restbench、toolbench三类工具的指令数据与评测代码TaskBench2023.11 发布用于评测 LLM 任务自动化能力的基准代码与数据集位于 taskbench 目录对应论文《TaskBench: Benchmarking Large Language Models for Task Automation》包含图生成、数据引擎、批量评测脚本等HuggingGPT 相关演进2023.04 支持 Azure OpenAI 服务与 GPT-4 模型新增 Gradio demo 与/tasks、/resultsWeb API2023.04.03 新增 CLI 模式并提供本地端点规模配置参数python awesome_chat.py --config configs/config.lite.yaml即可轻量体验2023.07 起进入评测与项目重建规划阶段。结语JARVISHuggingGPT展示了“LLM 规划 专家模型执行”这一 Agent 范式的完整工程落地四阶段工作流、任务依赖解析、多线程并发执行、本地 / 远程混合推理以及 CLI / Server / Web / Gradio / 嵌入式多形态部署。从 awesome_chat.py 与 models_server.py 的源码中可以看到整个系统在提示词、候选模型筛选、logit 偏置与依赖调度上都做了精细化设计。若希望深入验证论文结论或在此基础上做二次开发可直接运行仓库内置的--mode test用例并结合 config.default.yaml 与 config.lite.yaml 调整部署形态。赞分享人工智能大模型AI Agent工具调用模型评测【免费下载链接】JARVISJARVIS, a system to connect LLMs with ML community. Paper: https://arxiv.org/pdf/2303.17580.pdf项目地址https://gitcode.com/gh_mirrors/jarvis3/JARVIS点击查看免费下载相关推荐终极指南如何让老旧设备焕发新生的完整方案终极指南如何让老旧设备焕发新生的完整方案 OpenCore Legacy PatcherOCLP是一款专为老旧设备升级设计的开源工具它通过创新的系统兼容操作系统固件驱动开发最完整LLM模型转换教程将Hugging Face模型转为GGML格式最完整LLM模型转换教程将Hugging Face模型转为GGML格式 痛点与解决方案 你是否曾遇到这些问题Hugging Face模型文件体积庞大难以部署人工智能大模型本地部署Model-Optimizer 自定义 Hugging Face 模型量化插件开发以 DBRX MoE 为例实现 TensorRT-LLM 部署Model Optimizer 自定义 Hugging Face 模型量化插件开发以 DBRX MoE 为例实现 TensorRT LLM 部署 导读 本文聚人工智能大模型模型优化模型量化模型压缩上一篇React Native Boilerplate 10个常见问题终极解决方案 下一篇别再迷信 Revit 破解版5 个更安全、更划算的正规 BIM 方案创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考