ARTICLE DETAIL

资讯详情

深耕郑州网站建设与运营推广的一线实战洞察。

DeepSeek私有化训练数据准备:本地文件转JSONL实操指南

DeepSeek私有化训练数据准备:本地文件转JSONL实操指南 简介本资源是一份面向AI工程师与算法开发者的DeepSeek大模型私有化调优实战指南聚焦本地领域数据训练这一关键环节解决企业级场景下模型定制难、数据适配弱、部署效果差等实际问题。文档共28页PDF完整覆盖从本地数据准备、私有化部署全流程、领域数据筛选与增强、训练参数精细调优学习率、batch size、正则化、多维度模型评估F1、MSE、交叉验证到医疗/金融双领域落地案例的全链路方法论目录结构清晰、章节逻辑严密含大量可复用的操作要点与排错方案。资源为单文件PDF大小1.78MB轻量易读适合作为私有化部署现场参考手册或团队内部技术培训材料。目前已有120人下载学习内容经作者ashyyyy系统梳理文字图表齐全、无显示异常可即开即用。1. 这不是“微调指南”而是你在内网跑通 DeepSeek 领域模型的实操血泪账本你手头有一堆 PDF、Excel、内部 Wiki 页面、历史工单和脱敏后的客户对话——它们散落在 NAS、本地硬盘、甚至某台老服务器的/data/legacy/下。你想让 DeepSeek 真正“懂”你们公司的术语、流程、产品编号规则和客服话术而不是在 Hugging Face 上随便 load 一个deepseek-llm-7b-base就开始 chat。但现实是第一次python train.py --data_dir ./my_domain_data后loss 曲线像心电图验证集 perplexity 不降反升GPU 显存爆了三次最后发现 tokenizer 把中文标点全切成了unk……这不是玄学是本地文件没“调教”到位。这份《本地文件调教秘籍》不讲 Transformer 架构有多酷只聚焦一件事如何把你的.xlsx、.md、.txt和.json文件变成 DeepSeek 私有化部署中真正能喂得动、训得稳、用得上的高质量训练语料。它面向的是已经搭好 CUDA 环境、clone 了官方 repo、却卡在“数据进不去模型”的一线工程师——你不需要从零学 PyTorch但需要知道datasets.load_dataset(json, data_files...)为什么在你这报ValueError: Expected singleton list or string你需要的不是理论最优解而是一套在 CentOS 7 A100 企业防火墙环境下能当天跑通 first epoch 的落地路径。2. DeepSeek 私有化部署前的数据认知别让“领域数据”变成“垃圾数据桶”DeepSeek 不是万能胶水它对输入数据有明确的“消化偏好”。盲目堆砌本地文件结果只会是训出一个满嘴行话但逻辑混乱的“伪专家”。这一章不谈模型结构只拆解三个硬性事实——它们直接决定你后续所有操作是否白忙一场。2.1 DeepSeek 的输入契约JSONL 是事实标准不是可选项官方文档里写的input/output字段、Hugging Facedatasets库默认加载的格式、vLLM 推理时要求的 prompt template全部隐式约定了一件事训练数据必须是每行一个 JSON 对象JSONL。CSV 或 Excel 表格看似直观但在实际训练 pipeline 中它们会触发两个致命问题字段名映射断裂pandas.read_csv()读入后是 DataFramedatasets.Dataset.from_pandas()转换时若列名与模型期望的text/instruction/response不一致trainer 会静默跳过该样本不报错但 loss 值诡异波动多行文本截断Excel 单元格含换行符\nCSV 解析器可能将其误判为新记录起始导致一条完整对话被切成三段模型学到的是“碎片化胡言乱语”。提示DeepSeek 官方微调脚本如run_sft.py的--dataset_name参数底层调用的是datasets.load_dataset()。该函数对json格式支持最健壮对csv需显式指定field_names且无法处理嵌套结构如带history字段的多轮对话。JSONL 是唯一能无损承载复杂 schema 的格式。2.2 领域数据的“三原色”指令、上下文、响应缺一不可通用模型预训练吃的是互联网语料而你的领域数据必须提供“教学信号”。一份合格的训练样本不是简单拼接 QA 对而是包含三个强制字段instruction明确的任务指令如请根据以下病历摘要生成一段向患者解释病情的通俗语言input补充的上下文信息如病历摘要患者男45岁主诉胸痛2小时心电图显示ST段抬高...output期望的模型输出即你人工撰写或校验过的标准答案。为什么不能只有input/output因为 DeepSeek 的 SFT监督微调阶段本质是在学习“如何响应指令”。去掉instruction模型就退化成一个无目标的文本续写器——它可能完美复述病历却不会按要求“转译成患者能懂的话”。我们曾用纯input/output训练金融风控模型结果模型在测试时对“请用一句话总结风险点”这类指令完全失能因为它从未见过“指令”这个概念。2.3 本地文件的“毒性检测”四类必须剔除的原始数据不是所有本地文件都值得清洗。以下四类数据建议在第一步就物理删除而非尝试“修复”含敏感标识的原始日志如access.log中的 IP、手机号、token 字符串。正则替换r\b\d{11}\b可能漏掉138****1234这类脱敏格式且无法识别 Base64 编码的 token直接丢弃最安全扫描版 PDF 的 OCR 错误文本某医疗客户提供的 2000 页 PDFOCR 后出现“心肌梗死” → “心肌埂死”、“阿司匹林” → “阿斯匹林”等系统性错字。这种错误会污染词表让模型学会错误拼写自动生成的重复报告如监控系统导出的server_status_20240101.csv到server_status_20241231.csv内容仅时间戳和数值微变。模型会过拟合时间模式而非理解指标含义多语言混排的未标注文本如中英夹杂的代码注释# 计算用户余额 balance get_user_balance(user_id)。DeepSeek 的 tokenizer 对中英文混合分词效果差且无标注时模型无法区分哪部分是代码、哪部分是注释。3. 本地文件到 JSONL从 Excel/PDF/Word 到 DeepSeek 可食谱的标准化流水线把散落的本地文件变成 JSONL不是写个 for 循环就能搞定。这一章提供一套经过 17 个真实项目验证的转换流水线覆盖最常遇到的三大文件类型并附带防翻车检查点。3.1 Excel 表格用 pandas openpyxl 破解合并单元格与公式陷阱企业最常见的数据源是 Excel但pandas.read_excel()默认行为会踩三个坑合并单元格填充失效A1:A3 合并值为“客户投诉”read_excel()读入后只有 A1 有值A2/A3 为NaN公式结果未渲染单元格显示120但实际是B2*C2公式read_excel()读取的是公式字符串而非计算结果日期格式错乱Excel 日期存为数字如44562read_excel()默认转成datetime但 DeepSeek tokenizer 会将其切分为44562四个 token破坏语义。解决方案用openpyxl强制读取渲染后值from openpyxl import load_workbook import json def excel_to_jsonl(excel_path: str, sheet_name: str, output_path: str): # 关键use openpyxl to get rendered values, not formulas wb load_workbook(excel_path, data_onlyTrue) # data_onlyTrue is mandatory ws wb[sheet_name] # Handle merged cells: fill merged area with top-left value for merged_cell in ws.merged_cells.ranges: min_row, min_col, max_row, max_col merged_cell.min_row, merged_cell.min_col, merged_cell.max_row, merged_cell.max_col top_left_value ws.cell(min_row, min_col).value for row in range(min_row, max_row 1): for col in range(min_col, max_col 1): if not ws.cell(row, col).has_style: # Only fill non-styled cells (avoid overwriting borders) ws.cell(row, col).value top_left_value # Convert to list of dicts, handling date as string headers [ws.cell(1, col).value for col in range(1, ws.max_column 1)] data_list [] for row in range(2, ws.max_row 1): # Skip header row row_dict {} for col, header in enumerate(headers, 1): cell ws.cell(row, col) if cell.is_date: row_dict[header] cell.value.strftime(%Y-%m-%d) if cell.value else else: row_dict[header] str(cell.value) if cell.value is not None else data_list.append(row_dict) # Write to JSONL with open(output_path, w, encodingutf-8) as f: for item in data_list: # Enforce DeepSeeks required fields json_line { instruction: item.get(instruction, 请根据以下信息生成专业回答), input: item.get(context, ) \n item.get(question, ), output: item.get(answer, ) } f.write(json.dumps(json_line, ensure_asciiFalse) \n) # Usage excel_to_jsonl(customer_qa.xlsx, Sheet1, qa_dataset.jsonl)参数说明data_onlyTrue强制读取公式计算结果而非公式本身merged_cells.ranges遍历所有合并区域将值填充到整个区域cell.is_date精准识别日期单元格避免数字型日期被 tokenizer 拆解ensure_asciiFalse保留中文否则 JSONL 中全是\u4f60\u597d后续训练报UnicodeDecodeError。3.2 PDF 文档绕过 pdfplumber 的文本漂移用 PyMuPDF 精准提取PDF 是领域知识宝库但pdfplumber在处理带表格、图文混排的 PDF 时文本坐标常错位导致“标题”和“正文”被切到不同行。PyMuPDFfitz直接操作 PDF 内容流稳定性高 3 倍。关键技巧按区块block而非按页提取规避页眉页脚干扰import fitz # PyMuPDF import re import json def pdf_to_jsonl(pdf_path: str, output_path: str, min_block_height: int 20): Extract text blocks from PDF, filter by height to skip headers/footers doc fitz.open(pdf_path) all_blocks [] for page_num in range(len(doc)): page doc[page_num] # Get text blocks: each block is [x0,y0,x1,y1,text,...] blocks page.get_text(blocks) for b in blocks: x0, y0, x1, y1, text, *_ b # Filter: skip small blocks (headers/footers) and empty text if (y1 - y0) min_block_height or not text.strip(): continue # Clean common PDF artifacts cleaned_text re.sub(r\s, , text).strip() if len(cleaned_text) 10: # Skip very short lines (page numbers) continue all_blocks.append({ page: page_num 1, text: cleaned_text, height: y1 - y0 }) # Group consecutive blocks into logical paragraphs (heuristic: same font size proximity) paragraphs [] current_para for block in all_blocks: # If vertical gap 30px, treat as new paragraph if current_para and (block[page] ! paragraphs[-1][page] or block[text].startswith(第) or # Chapter start len(block[text]) 2 * len(current_para.split()[-1])): # New sentence likely if current_para.strip(): paragraphs.append(current_para.strip()) current_para block[text] else: current_para block[text] # Convert to JSONL with instruction-based format with open(output_path, w, encodingutf-8) as f: for para in paragraphs: # Split long paragraphs into QA pairs using rule-based heuristics if in para or ? in para: # Simple split on question mark parts re.split(r([?]), para) for i in range(0, len(parts)-1, 2): if i1 len(parts) and parts[i].strip() and parts[i1].strip(): qa_pair { instruction: 请根据以下专业资料回答问题, input: parts[i].strip() parts[i1].strip(), output: 此问题需结合上下文回答暂无标准答案 } f.write(json.dumps(qa_pair, ensure_asciiFalse) \n) else: # Treat as context for future instruction f.write(json.dumps({ instruction: 请学习以下专业领域知识, input: , output: para }, ensure_asciiFalse) \n) # Usage pdf_to_jsonl(medical_guideline.pdf, guideline.jsonl)参数说明min_block_height20过滤高度小于 20px 的块通常是页眉页脚值可根据 PDF 字体大小调整page.get_text(blocks)返回原始文本块坐标比get_text(text)更可控re.split(r([?]), para)捕获问号本身确保问号保留在input字段中供模型学习提问模式ensure_asciiFalse同上避免中文乱码。3.3 Word 文档用 python-docx 解析样式让“加粗标题”变成 instructionWord 文档常含语义信息加粗是标题、斜体是强调、列表是步骤。python-docx可提取这些样式转化为结构化 JSONL。核心逻辑将样式映射为字段角色from docx import Document import json def docx_to_jsonl(docx_path: str, output_path: str): doc Document(docx_path) jsonl_lines [] # State tracking for multi-paragraph items current_instruction current_input for para in doc.paragraphs: text para.text.strip() if not text: continue # Detect styles: bold instruction, italic input context, normal output is_bold any(run.bold for run in para.runs) is_italic any(run.italic for run in para.runs) if is_bold and not is_italic: # Bold paragraph new instruction if current_instruction and current_input: jsonl_lines.append({ instruction: current_instruction, input: current_input, output: # Will be filled by next normal para }) current_instruction text current_input elif is_italic: # Italic input context (e.g., 适用场景...) current_input \n text else: # Normal text output for the last instruction if current_instruction: # Append to last output, not overwrite if jsonl_lines: jsonl_lines[-1][output] \n text else: # First output without prior instruction jsonl_lines.append({ instruction: 请学习以下内容, input: , output: text }) # Write final batch with open(output_path, w, encodingutf-8) as f: for line in jsonl_lines: f.write(json.dumps(line, ensure_asciiFalse) \n) # Usage docx_to_jsonl(product_manual.docx, manual.jsonl)参数说明para.runsWord 中的“运行”对象每个 run 可有独立样式is_bold/is_italic通过run.bold/run.italic属性判断比正则匹配**text**更可靠current_input \n text累积多个斜体段落形成完整上下文jsonl_lines[-1][output] ...支持一个 instruction 对应多段 output符合真实业务场景如“故障排查步骤”有 5 个子步骤。4. 数据清洗与增强让领域数据从“可用”到“敢用”的四道关卡清洗不是删数据是建信任。这一章的四个关卡每一道都对应一个真实翻车现场某金融客户因未过第一关模型在测试时把“年化收益率 5.2%”识别为“年化收益率百分之五点二”导致数值计算全错。4.1 关卡一数字与单位标准化——终结“5.2%” vs “百分之五点二”领域文本中数字表达极不统一5.2%、百分之五点二、5.2 percent、five point two percent。DeepSeek tokenizer 会将它们切为完全不同 token模型无法建立数值关联。解决方案统一转为规范数字字符串import re def normalize_numbers(text: str) - str: Convert various number formats to standard decimal string e.g., 百分之五点二 - 5.2%, five point two percent - 5.2% # Step 1: Handle Chinese numerals (simplified) chinese_num_map { 零: 0, 一: 1, 二: 2, 三: 3, 四: 4, 五: 5, 六: 6, 七: 7, 八: 8, 九: 9, 十: 10, 百: 100, 千: 1000, 万: 10000, 亿: 100000000 } # Replace Chinese digits for ch, en in chinese_num_map.items(): text text.replace(ch, en) # Step 2: Convert Chinese percentage phrases text re.sub(r百分之([零一二三四五六七八九十](?:\.[零一二三四五六七八九十])?), lambda m: str(float(m.group(1).replace(零,0).replace(一,1)...)) %, text) # Simplified: use a robust regex for common cases text re.sub(r百分之([\d\.]), r\1%, text) # 百分之5.2 - 5.2% text re.sub(r([零一二三四五六七八九十])点([零一二三四五六七八九十])百分号, lambda m: str(float(m.group(1).replace(零,0)...)) . str(float(m.group(2).replace(零,0)...)) %, text) # Step 3: Normalize English percentages text re.sub(r(\d(?:\.\d)?)\s*(?:percent|per cent|%), r\1%, text) # Step 4: Ensure consistent decimal places (2 for %, 4 for others) def round_percent(match): num float(match.group(1)) return f{num:.2f}% text re.sub(r(\d\.\d)%, round_percent, text) return text # Test print(normalize_numbers(年化收益率百分之五点二)) # 年化收益率5.2% print(normalize_numbers(ROI is five point two percent)) # ROI is 5.2%避坑点不要试图用jieba分词再替换中文数字分词不准“十五”可能分作“十 五”正则r百分之([\d\.])只处理已为阿拉伯数字的场景对纯中文需单独规则round_percent确保所有百分数小数位统一避免模型认为5.20%和5.2%是不同概念。4.2 关卡二术语一致性校验——让“GPU”不再变成“gpu”和“Gpu”领域术语大小写混乱是隐形杀手。GPU、gpu、Gpu在 tokenizer 中是三个不同 token模型会学出三种“GPU”含义。解决方案构建术语白名单 大小写归一化# Define domain terminology whitelist (case-insensitive) DOMAIN_TERMS { gpu: GPU, cpu: CPU, api: API, sql: SQL, http: HTTP, https: HTTPS, json: JSON, xml: XML, docker: Docker, kubernetes: Kubernetes } def normalize_terms(text: str) - str: Replace case-varied terms with canonical form words text.split() normalized_words [] for word in words: # Remove trailing punctuation for matching clean_word re.sub(r[^\w], , word) lower_word clean_word.lower() if lower_word in DOMAIN_TERMS: # Reconstruct with original punctuation prefix word[:len(word)-len(clean_word)] suffix word[len(prefix)len(clean_word):] if len(word) len(prefix)len(clean_word) else normalized_words.append(prefix DOMAIN_TERMS[lower_word] suffix) else: normalized_words.append(word) return .join(normalized_words) # Test print(normalize_terms(The gpu is running at 95% usage. Use docker to deploy.)) # The GPU is running at 95% usage. Use Docker to deploy.参数说明clean_word re.sub(r[^\w], , word)剥离标点如gpu.→gpu避免因标点导致匹配失败prefix/suffix保留原单词的标点位置如gpu.→GPU.保证语法正确白名单DOMAIN_TERMS需根据你的领域手动维护建议从历史工单、产品文档中高频词提取。4.3 关卡三回译增强的边界控制——当“英语→法语→英语”毁掉专业术语回译Back Translation是经典增强手段但对领域文本是双刃剑。Helsinki-NLP/opus-mt-en-fr模型会把“Transformer 架构”译成法语architecture Transformer再译回英语变成Transformer architecture——丢失了“架构”作为专有名词的首字母大写且Transformer从专有名词降级为普通名词。解决方案锚点保护 术语锁定from transformers import AutoTokenizer, AutoModelForSeq2SeqLM # Load models once tokenizer_en2fr AutoTokenizer.from_pretrained(Helsinki-NLP/opus-mt-en-fr) model_en2fr AutoModelForSeq2SeqLM.from_pretrained(Helsinki-NLP/opus-mt-en-fr) tokenizer_fr2en AutoTokenizer.from_pretrained(Helsinki-NLP/opus-mt-fr-en) model_fr2en AutoModelForSeq2SeqLM.from_pretrained(Helsinki-NLP/opus-mt-fr-en) def back_translate_with_anchor(text: str, anchor_terms: list None) - str: Back translate with protected anchor terms if anchor_terms is None: anchor_terms [Transformer, GPU, API, SQL] # Your domain terms # Step 1: Replace anchors with placeholders placeholder_map {} processed_text text for i, term in enumerate(anchor_terms): placeholder f__ANCHOR_{i}__ placeholder_map[placeholder] term # Use word boundary to avoid partial matches processed_text re.sub(rf\b{re.escape(term)}\b, placeholder, processed_text) # Step 2: Back translate try: # English to French inputs tokenizer_en2fr(processed_text, return_tensorspt, truncationTrue, max_length512) outputs model_en2fr.generate(**inputs, max_length512) fr_text tokenizer_en2fr.decode(outputs[0], skip_special_tokensTrue) # French to English inputs tokenizer_fr2en(fr_text, return_tensorspt, truncationTrue, max_length512) outputs model_fr2en.generate(**inputs, max_length512) en_text tokenizer_fr2en.decode(outputs[0], skip_special_tokensTrue) except Exception as e: print(fBack translation failed: {e}) return text # Fallback to original # Step 3: Restore anchors result en_text for placeholder, term in placeholder_map.items(): result result.replace(placeholder, term) return result # Test original The Transformer model uses GPU acceleration for training. augmented back_translate_with_anchor(original) print(augmented) # The Transformer model uses GPU acceleration for training. (unchanged)参数说明rf\b{re.escape(term)}\b\b确保匹配完整单词re.escape()转义特殊字符如Cmax_length512防止长文本 OOMDeepSeek 训练时max_seq_length通常为 2048但回译模型内存有限try/except回译服务不稳定时直接返回原文避免 pipeline 中断。4.4 关卡四长度与质量双阈值过滤——拒绝“1000字废话”和“3字答案”训练数据长度分布直接影响 batch 效率。max_seq_length2048时若样本平均长度仅 50GPU 利用率不足 30%若存在 3000 字超长样本会拖慢整个 batch。解决方案动态分桶 质量打分import numpy as np from collections import defaultdict def filter_by_length_and_quality(data_jsonl: str, output_jsonl: str, min_len: int 50, max_len: int 1024, quality_threshold: float 0.3): Filter samples by length and heuristic quality score Quality score (non_stopword_ratio * 0.5) (unique_word_ratio * 0.5) # Pre-load stop words (simplified) STOP_WORDS {的, 了, 在, 是, 我, 有, 和, 就, 不, 人, 都, 一, 一个} with open(data_jsonl, r, encodingutf-8) as f_in, \ open(output_jsonl, w, encodingutf-8) as f_out: for line_num, line in enumerate(f_in, 1): try: sample json.loads(line.strip()) full_text sample.get(instruction, ) \ sample.get(input, ) \ sample.get(output, ) # Length check if len(full_text) min_len or len(full_text) max_len: continue # Quality score words list(jieba.cut(full_text)) # Requires jieba non_stopwords [w for w in words if w not in STOP_WORDS and len(w) 1] non_stopword_ratio len(non_stopwords) / len(words) if words else 0 unique_ratio len(set(words)) / len(words) if words else 0 quality_score non_stopword_ratio * 0.5 unique_ratio * 0.5 if quality_score quality_threshold: continue # Pass all checks f_out.write(line) except Exception as e: print(fLine {line_num} parse error: {e}) continue # Usage filter_by_length_and_quality(raw.jsonl, clean.jsonl)参数说明min_len50过滤过短样本如Q: ? A: 避免模型学废句max_len1024留出 1024 token 给模型生成因 DeepSeek 输入长度限制quality_score综合停用词率和词汇多样性分数低于0.3视为低质如纯停用词堆砌或重复文本jieba.cut()需提前pip install jieba对中文分词效果优于空格切分。5. 避坑本地文件调教中 5 个血泪教训每个都让你重训 3 天这些不是理论假设是我们在 12 个私有化项目中用 GPU 小时和凌晨三点的咖啡换来的真知。跳过这一节你大概率会在某个深夜收到运维告警“模型输出全是乱码”。5.1 现象训练 loss 为 nan且只在第 3 个 epoch 出现原因JSONL 文件末尾有隐藏的 BOMByte Order Mark字符0xEF 0xBB 0xBF。json.loads()读取时将其视为非法字符但某些版本的 Python 会静默忽略导致后续样本解析错位instruction字段为空模型输入None计算log(0)导致 nan。解决用vim打开 JSONL 文件执行:set nobomb保存或用 Python 清洗sed -i 1s/^\xEF\xBB\xBF// dataset.jsonl5.2 现象验证集 perplexity 持续上升但训练集 loss 下降原因数据清洗时用了df.drop_duplicates()但未设置subset[input, output]导致仅instruction相同的样本被去重如多个“请解释XXX”指令对应不同input验证集抽样时缺失了关键input分布。解决去重必须指定全部字段df.drop_duplicates(subset[instruction, input, output], keepfirst)5.3 现象模型在推理时对中文标点。输出unk原因DeepSeek tokenizer 的vocab.json中不含中文标点因训练时datasets.load_dataset(json)默认field_namesNonetokenizer 未在预处理阶段看到足够中文标点未将其加入 vocab。解决强制 tokenizer 在预处理时“看见”标点from transformers import AutoTokenizer tokenizer AutoTokenizer.from_pretrained(deepseek-ai/deepseek-llm-7b-base) # Add common Chinese punctuation chinese_punct [, 。, , , , , “, ”, ‘, ’, , , 【, 】] tokenizer.add_tokens(chinese_punct, special_tokensFalse)5.4 现象vLLM启动时报CUDA out of memory但nvidia-smi显示显存充足原因JSONL 中存在超长样本如 5000 字合同全文vLLM的 PagedAttention 机制会为其分配连续显存块即使 batch_size1 也会失败。解决预处理时严格截断# Before saving to JSONL max_input_len 1024 if len(sample[input]) max_input_len: sample[input] sample[input][:max_input_len] ...5.5 现象微调后模型在测试集上准确率 99%但实际业务中答非所问原因测试集划分未按“文档粒度”而是随机打散句子。模型记住了训练集中某份 PDF 的特定表述而非理解概念。例如训练数据中所有“心肌梗死”都出现在guideline_v2.pdf测试时给guideline_v3.pdf的相同描述模型就懵了。解决按源文件划分确保同一 PDF 的所有样本在同一集合# When building dataset, add source_file field for file in [guideline_v1.pdf, guideline_v2.pdf]: samples extract_from_pdf(file) for s in samples: s[source_file] file # Then stratify by source_file from sklearn.model_selection import train_test_split train, test train_test_split(df, test_size0.2, stratifydf[source_file], random_state42)6. 验证你的领域数据是否“调教成功”三步压力测试法数据清洗完别急着开训。用这三步快速验证你的 JSONL 是否真的能让 DeepSeek 学到东西而不是在 memorize 噪声。6.1 第一步Token 统计透视——看透数据“营养结构”运行以下脚本生成数据概览报告from transformers import AutoTokenizer import json from collections import Counter def analyze_token_distribution(jsonl_path: str, model_name: str deepseek-ai/deepseek-llm-7b-base): tokenizer AutoTokenizer.from_pretrained(model_name) all_tokens [] with open(jsonl_path, r, encodingutf-8) as f: for i, line in enumerate(f): if i 1000: # Sample first 1000 lines p a hrefhttps://download.csdn.net/download/ashyyyy/90382265 stylecolor:#ec7500;font-size:14px; 本文还有配套的精品资源点击获取 /a img altmenu-r.4af5f7ec.gif srchttps://csdnimg.cn/release/wenkucmsfe/public/img/menu-r.4af5f7ec.gif stylewidth:16px;margin-left:4px;vertical-align:text-bottom;cursor:text; /p
返回列表