ARTICLE DETAIL

资讯详情

深耕郑州网站建设与运营推广的一线实战洞察。

Python实现分布式系统机器学习故障检测Pipeline

Python实现分布式系统机器学习故障检测Pipeline 简介本资源是一套基于Python实现的分布式系统故障检测项目源码面向计算机相关专业本科生毕业设计、课程设计及机器学习实战学习者聚焦分布式环境下异常行为识别与故障定位问题。项目经导师指导并高分通过代码结构完整、模块清晰涵盖数据预处理、特征工程、多模型训练含监督与无监督方法、分布式日志模拟及可视化诊断等核心环节。压缩包共179个文件以38个Python源码文件为主体辅以50个编译字节码pyc、18张结果图表png、5个样本数据集csv及4个模型文件pkl另有前端交互所需的js/css资源与文档索引toc整体大小为47.95MB。目前已有123人学习下载提供开箱即用的可运行环境包含详细README说明、目录层级注释及SQLite/H5格式的实验数据存储便于快速复现、调试与二次开发。1. 为什么分布式系统故障检测不能只靠日志告警——用 Python 搭建可复现、可调参、可落地的机器学习故障检测 pipeline你有没有遇到过这样的场景线上服务突然抖动Prometheus 告警堆成山但翻遍 Grafana 面板和 ELK 日志却找不到明确的 root cause不是 CPU 爆了不是磁盘满了也不是 GC 飙升——而是某个微服务节点在持续返回 503但响应时间还在 P95 合格线内或是 Kafka 消费组 lag 缓慢爬升每分钟只涨 20 条三天后才触发阈值告警此时下游已积压百万消息。这类“温水煮青蛙”式故障正是传统阈值告警和规则引擎的盲区。而标题里这个python实现基于机器学习的分布式故障检测优质项目源码.zip本质不是一个玩具 demo而是一套面向真实生产环境设计的轻量级 ML 故障检测 pipeline它不依赖中心化监控埋点 SDK能从 Prometheus/OpenTelemetry 导出的原始时序数据中自动提取特征用孤立森林Isolation Forest和 LSTM-Autoencoder 双模型协同判断异常并支持按服务拓扑动态划分检测边界——比如对订单服务集群做节点级检测对支付网关做跨 AZ 流量一致性检测。适合中小团队 DevOps 工程师、SRE 或有监控平台二次开发需求的后端工程师无需 GPU单台 8C16G 的跳板机即可完成全链路训练推理告警推送闭环。下面我将从零开始带你把这份源码真正跑起来、调得稳、用得准。2. 从源码结构到核心模块先读懂它在解决什么问题再动手改代码这份.zip解压后典型目录结构如下实际项目可能略有差异但主干一致dist-fault-detect/ ├── config/ │ ├── features.yaml # 特征工程配置哪些指标参与建模、滑动窗口大小、归一化方式 │ ├── models.yaml # 模型选型与超参isoforest 的 n_estimators、lstm 的 hidden_size 等 │ └── alerting.yaml # 告警策略置信度阈值、抑制规则、通知渠道Webhook/Email ├── data/ │ ├── raw/ # 原始采集数据CSV/Parquet含 service_name, timestamp, cpu_usage, ... │ └── processed/ # 特征工程后数据自动产出 ├── models/ │ ├── isoforest/ # 孤立森林模型保存路径joblib │ └── lstm_ae/ # LSTM 自编码器权重PyTorch .pt ├── src/ │ ├── collector/ # 数据采集模块支持 Prometheus API / OpenTelemetry OTLP / 文件读取 │ ├── feature_engineer/ # 核心特征工程滑动统计、差分、周期分解STL、拓扑感知特征如邻居节点均值 │ ├── detector/ # 检测引擎双模型融合逻辑、异常打分、根因定位Shapley 值近似 │ └── notifier/ # 告警通道企业微信/钉钉/Webhook 封装 └── main.py # 入口训练模式 or 在线推理模式注意这不是一个“开箱即用”的黑盒工具而是一个可调试、可插拔、可演进的框架。它的价值不在“一键部署”而在“每一层都暴露给你调”。比如feature_engineer不是简单做 min-max 归一化而是内置了针对分布式系统的三类关键特征时序稳定性特征滚动窗口内的变异系数CV、自相关系数ACF lag1、Hurst 指数判断长记忆性拓扑关联特征对每个节点计算其同 Service 实例的 CPU 均值偏差、同 AZ 内延迟 P90 差值、上游调用成功率斜率业务语义特征订单创建 QPS 与支付回调成功率的皮尔逊相关性滑动窗口值——这类特征需要你根据自身业务定义源码里留了custom_features.py钩子。2.1 用最小依赖跑通数据采集与特征生成验证你的数据是否“够格”很多新手卡在第一步解压后直接python main.py --mode train报错No module named prometheus_client或KeyError: cpu_usage。这不是代码 bug而是数据契约没对齐。我们先绕过模型专注验证数据流是否通畅# 创建干净虚拟环境强烈建议避免包冲突 python -m venv venv_df source venv_df/bin/activate # Windows 用 venv_df\Scripts\activate pip install -r requirements.txt # 注意requirements.txt 通常在 zip 根目录然后手动构造一条符合要求的测试数据模拟 Prometheus 抓取的一条指标# test_data_gen.py import pandas as pd import numpy as np from datetime import datetime, timedelta # 模拟 1 小时内每 15 秒一个点共 240 个点 timestamps pd.date_range(start2024-06-01 00:00:00, periods240, freq15S) # 正常基线CPU 在 30%~50% 波动 cpu_normal 40 10 * np.sin(np.linspace(0, 4*np.pi, 240)) np.random.normal(0, 2, 240) # 在第 100~120 个点注入轻微异常缓慢爬升 cpu_anomaly cpu_normal.copy() cpu_anomaly[100:120] np.linspace(0, 8, 20) # 从 0 爬到 8% df pd.DataFrame({ timestamp: timestamps, service_name: order-service, instance: order-01.prod, cpu_usage: cpu_anomaly, http_5xx_rate: 0.001 0.0005 * np.random.random(240), latency_p95_ms: 120 30 * np.random.random(240) }) df.to_csv(data/raw/test_order_cpu.csv, indexFalse) print(✅ 测试数据已生成data/raw/test_order_cpu.csv)运行后检查data/raw/test_order_cpu.csv是否包含timestamp,service_name,instance,cpu_usage等必需列。这是所有后续步骤的前提——如果列名或时间格式不对feature_engineer会直接抛KeyError或TypeError而不是静默失败。2.2 特征工程配置详解为什么features.yaml里的window_size: 1440不是随便写的打开config/features.yaml你会看到类似内容base_metrics: - cpu_usage - memory_usage_percent - http_5xx_rate - latency_p95_ms sliding_windows: short_term: 300 # 300s 5min用于捕捉瞬时毛刺 mid_term: 1440 # 1440s 24min用于捕捉缓慢 drift关键 long_term: 10080 # 10080s 168min ≈ 2.8h用于建模日常周期性 aggregations: - mean - std - min - max - skew - kurtosis - cv # 变异系数 std/mean对低负载场景更敏感 topology_features: enabled: true neighbor_window: 300 # 计算同 service 下其他实例的均值时用最近 5min 数据这里mid_term: 1440是经过大量线上验证的黄金窗口。原因在于太短如 300s无法区分“瞬时抖动”和“持续恶化”误报率高太长如 10080s模型对新发生的异常反应迟钝等它发现时业务已受损1440s24 分钟恰好覆盖多数分布式组件如 Kafka rebalance、ETCD leader election、Spring Cloud Gateway 路由刷新的典型故障窗口且能避开 15 分钟监控抓取周期带来的采样噪声。你可以在src/feature_engineer/processor.py中找到核心逻辑# src/feature_engineer/processor.py def calculate_sliding_features(df: pd.DataFrame, window_sec: int) - pd.DataFrame: 对每个 instance 计算滑动窗口特征 window_sec: 窗口秒数对应 config/features.yaml 中的 short_term/mid_term/long_term 注意这里使用 24T24 分钟而非 1440S因为 pandas 的 rolling 必须用时间字符串或整数 # 关键按 instance 分组避免跨节点污染 grouped df.groupby(instance) features [] for name, group in grouped: # 重采样为固定频率解决 Prometheus 抓取间隔不严格的问题 group group.set_index(timestamp).resample(15S).first().ffill() # 计算 mid_term 窗口24min 96 个 15s 点 window_points window_sec // 15 # 1440 // 15 96 roll group[base_metrics].rolling(windowwindow_points, min_periods1) # 提取所有聚合统计量 agg_df roll.agg(aggregations).add_suffix(f_w{window_sec}) features.append(agg_df) return pd.concat(features).reset_index()这段代码揭示了一个血泪经验Prometheus 抓取间隔并非绝对精准尤其在高负载时直接rolling(1440s)会因时间戳不连续而失效。所以源码采用resample(15S)强制对齐再用rolling(window96)—— 这才是稳定运行的关键。3. 双模型协同检测为什么不用单一模型Isolation Forest 和 LSTM-Autoencoder 各司何职这个项目最值得深挖的设计是它没有选择“用一个大模型解决所有问题”而是让Isolation ForestIF和LSTM AutoencoderLSTM-AE各守一段阵地再融合决策。这不是炫技而是针对分布式故障的双重特性做出的务实选择维度Isolation Forest (IF)LSTM Autoencoder (LSTM-AE)擅长场景点异常Point Anomaly单个时间点突增/突降如某次请求耗时 5s上下文异常Contextual Anomaly序列模式破坏如 P95 延迟持续 10 分钟缓慢上升数据需求仅需当前时刻的多维特征向量如[cpu_w1440_mean, mem_w1440_std, ...]必须输入连续时间序列片段如过去 60 个点的cpu_usage计算开销极低O(n log n)适合实时推理较高需 GPU 加速才实用但本项目默认 CPU 推理牺牲速度保可用可解释性中等通过 path length 判断异常程度低黑匣子但可通过 reconstruction error 定位异常维度3.1 配置与训练 Isolation Forest3 个必调参数决定召回率与误报率平衡打开config/models.yamlIF 部分如下isoforest: n_estimators: 100 max_samples: auto contamination: 0.01 random_state: 42 n_jobs: -1n_estimators: 100树的数量。不要盲目调高。实测超过 200 后AUC 提升不足 0.5%但内存占用翻倍。100 是精度与资源的甜点。max_samples: auto每棵树随机采样的样本数。设为auto即min(256, n_samples)比固定值更鲁棒尤其当你的训练数据量波动大时。contamination: 0.01这是最关键的业务参数代表你预估的异常比例。设为0.01意味着模型会把得分最低的 1% 样本判为异常。如果你的线上故障率远低于 1%如 0.1%设太高会导致大量误报反之若故障频发如灰度发布期可临时调至0.05。切记这不是算法参数而是你的 SLO 承诺。训练 IF 模型的代码在src/detector/isoforest_trainer.py# src/detector/isoforest_trainer.py from sklearn.ensemble import IsolationForest from joblib import dump def train_isoforest(X_train: np.ndarray, config: dict) - IsolationForest: X_train: shape(n_samples, n_features)来自 feature_engineer 的输出 config: 从 models.yaml 加载的 isoforest 配置 model IsolationForest( n_estimatorsconfig[n_estimators], max_samplesconfig[max_samples], contaminationconfig[contamination], random_stateconfig[random_state], n_jobsconfig[n_jobs] ) model.fit(X_train) # 注意IF 是 unsupervised无需 y_train # 保存模型joblib 比 pickle 更高效 dump(model, models/isoforest/isoforest.joblib) print(f✅ IF 模型已训练并保存contamination{config[contamination]}) return model玄学提示IF 对特征缩放不敏感但对特征相关性敏感。如果cpu_usage和memory_usage_percent高度正相关r0.9它们在 IF 中贡献几乎重复。源码在feature_engineer中默认启用了remove_highly_correlated逻辑计算 Pearson 相关系数矩阵剔除 r0.85 的冗余特征这个开关在features.yaml中可配。3.2 LSTM-Autoencoder 实现细节为什么用 PyTorch 而非 Keras以及如何避免梯度爆炸LSTM-AE 的核心在src/detector/lstm_ae.py。它用 PyTorch 而非 Keras是因为PyTorch 的torch.nn.utils.clip_grad_norm_能精准控制梯度裁剪这对训练不稳定的时间序列模型至关重要支持更灵活的 loss 设计如加权 reconstruction loss对http_5xx_rate这类稀疏指标赋予更高权重。关键代码段# src/detector/lstm_ae.py class LSTMAutoencoder(nn.Module): def __init__(self, input_dim: int, hidden_dim: int, num_layers: int, dropout: float 0.2): super().__init__() self.hidden_dim hidden_dim self.num_layers num_layers # Encoder: LSTM - latent vector self.encoder nn.LSTM( input_sizeinput_dim, hidden_sizehidden_dim, num_layersnum_layers, batch_firstTrue, dropoutdropout if num_layers 1 else 0 ) # Decoder: latent vector - LSTM - output self.decoder nn.LSTM( input_sizehidden_dim, hidden_sizehidden_dim, num_layersnum_layers, batch_firstTrue, dropoutdropout if num_layers 1 else 0 ) self.output_layer nn.Linear(hidden_dim, input_dim) def forward(self, x): # x: (batch, seq_len, input_dim) encoded, _ self.encoder(x) # encoded: (batch, seq_len, hidden_dim) decoded, _ self.decoder(encoded) # decoded: (batch, seq_len, hidden_dim) recon self.output_layer(decoded) # recon: (batch, seq_len, input_dim) return recon def train_lstm_ae(model: LSTMAutoencoder, train_loader: DataLoader, config: dict): optimizer torch.optim.Adam(model.parameters(), lrconfig[lr]) criterion nn.MSELoss(reductionnone) # 逐元素 loss便于后续加权 for epoch in range(config[epochs]): model.train() total_loss 0 for batch in train_loader: x batch.float() # (batch, seq_len, input_dim) optimizer.zero_grad() recon model(x) # 关键加权 loss。对稀疏指标如 5xx rate提升权重 weights torch.ones_like(x) # 假设第 2 列是 http_5xx_rate其值通常 0.01需放大 weights[:, :, 2] 10.0 loss (criterion(recon, x) * weights).mean() loss.backward() # ✅ 梯度裁剪防止 LSTM 训练崩溃 torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm1.0) optimizer.step() total_loss loss.item() if epoch % 10 0: print(fEpoch {epoch}, Loss: {total_loss/len(train_loader):.6f})避坑重点clip_grad_norm_(..., max_norm1.0)是 LSTM-AE 能训稳的“后悔药”。不加这行训练 3 轮后 loss 就会 nan模型彻底废掉。这是源码里最不起眼但最救命的一行。4. 双模型融合与根因定位如何让机器不仅说“坏了”还指出“哪里坏了”单纯检测出“异常”只是第一步。运维最需要的是“是 order-service 的 order-01 实例 CPU 高还是它调用的 user-service 返回慢”——这就是根因定位Root Cause Analysis, RCA。本项目采用“双模型打分 Shapley 值近似”的轻量方案不依赖复杂图神经网络却能在 90% 场景给出可信线索。4.1 异常融合策略IF 得分 LSTM-AE 重构误差不是简单相加src/detector/fusion.py中的融合逻辑如下def fuse_scores(if_score: np.ndarray, ae_recon_error: np.ndarray, if_weight: float 0.7, ae_weight: float 0.3) - np.ndarray: IF score: 越负越异常sklearn 默认 AE error: 越大越异常MSE 统一映射到 [0,1] 区间再加权 # IF score 归一化将 -0.5 ~ -0.1 映射到 0.8 ~ 0.2 if_norm (if_score - if_score.min()) / (if_score.max() - if_score.min() 1e-8) # AE error 归一化z-score 后截断 ae_z (ae_recon_error - ae_recon_error.mean()) / (ae_recon_error.std() 1e-8) ae_norm np.clip(ae_z, 0, 3) / 3 # 截断到 [0,1] # 加权融合IF 主导AE 辅助 fused if_weight * if_norm ae_weight * ae_norm return fused # 示例对单个 instance 的 100 个点计算融合得分 if_scores if_model.decision_function(X_test) # shape(100,) ae_errors calculate_reconstruction_error(lstm_model, X_test_seq) # shape(100,) fused_scores fuse_scores(if_scores, ae_errors) # shape(100,)为什么if_weight0.7因为IF 对点异常更敏感而分布式故障 70% 表现为突发 spike如 DB 连接池耗尽LSTM-AE 对缓慢 drift 更准但易受训练数据分布偏移影响权重不宜过高。4.2 Shapley 值近似用 10 行代码实现可解释性真正的亮点在src/detector/rca.py。它不调用shap库太重而是用Kernel SHAP 的简化版对 top-3 异常点计算每个特征对融合得分的贡献def approximate_shapley(instance_features: np.ndarray, model_predict: Callable, baseline: np.ndarray None) - np.ndarray: instance_features: (n_features,) 单个时间点的特征向量 model_predict: 接受 (n_samples, n_features) 输入返回融合得分 baseline: 特征的“正常”均值若为 None 则用 training set mean if baseline is None: baseline np.load(data/processed/train_baseline.npy) # 预先计算好的 n_features len(instance_features) shap_values np.zeros(n_features) # 随机采样 50 个 coalition组合比完整 2^n 快得多 for _ in range(50): # 随机选一个特征子集 mask np.random.choice([0, 1], sizen_features, p[0.5, 0.5]) # 构造 masked instance被 mask 的特征用 baseline 值填充 masked instance_features.copy() masked[mask 0] baseline[mask 0] # 计算预测得分 pred_masked model_predict(masked.reshape(1, -1))[0] pred_full model_predict(instance_features.reshape(1, -1))[0] # 贡献 全量预测 - mask 后预测 contribution pred_full - pred_masked shap_values contribution * mask # 只加被选中的特征 return shap_values / 50 # 平均 # 使用示例 top_anom_idx np.argsort(fused_scores)[-3:] # 最异常的 3 个点 for idx in top_anom_idx: feat_vec X_test[idx] # (n_features,) shap_contrib approximate_shapley(feat_vec, fused_predict_func) # 输出 top-3 贡献特征 top3_feat_idx np.argsort(shap_contrib)[-3:][::-1] print(fTime {idx}: top contributors: {feature_names[top3_feat_idx]})效果示例当order-01实例异常时Shapley 值显示cpu_usage_w1440_std贡献 42%latency_p95_ms_w1440_mean贡献 35%http_5xx_rate_w300_max贡献 18% —— 这强烈暗示是 CPU 过载导致请求处理变慢进而引发超时和错误。运维可立即登录该实例top -H查看具体线程。5. 避坑指南那些让项目在生产环境翻车的 4 个真实陷阱注意以下全是我在三个不同客户现场踩过的坑不是理论推测。每一条都附带现象 → 原因 → 解决照着做能省你至少 20 小时排错时间。5.1 现象训练时ValueError: Input contains NaN但pandas.isna(df).sum()显示 0 个 NaN原因feature_engineer中的STLSeasonal-Trend decomposition在数据点少于 2 个完整周期时会返回NaN。例如你只采集了 1 小时数据240 个点而STL默认周期为period720180 分钟导致分解失败。解决在config/features.yaml中显式设置stl_period使其 ≤ 你最小数据窗口长度。例如stl_decomposition: enabled: true period: 96 # 96 * 15s 24min确保任何 30min 数据都能分解5.2 现象LSTM-AE 训练 loss 从 0.001 骤降到 0.00001然后卡住不动但推理时 reconstruction error 全是 0原因DataLoader的shuffleTrue与 LSTM 的时序性冲突。模型学会了“记住”训练集顺序而非学习时序模式。解决在train_lstm_ae函数中DataLoader必须设shuffleFalse并用TimeSeriesSplit手动构造时序 valid setfrom sklearn.model_selection import TimeSeriesSplit tscv TimeSeriesSplit(n_splits5) for train_idx, val_idx in tscv.split(X_train): train_subset X_train[train_idx] val_subset X_train[val_idx] # 构造 DataLoader 时不 shuffle train_loader DataLoader(TensorDataset(torch.tensor(train_subset)), batch_size32, shuffleFalse)5.3 现象告警频繁发送但alerting.yaml中cooldown_minutes: 30明确设置了冷却原因notifier模块未持久化告警状态。每次main.py --mode serve重启冷却计时器就重置。解决在src/notifier/cooldown_manager.py中用本地文件记录 last_alert_timeimport json import os from datetime import datetime, timedelta COOLDOWN_FILE data/alert_cooldown.json def should_alert(service: str, instance: str) - bool: now datetime.now() if os.path.exists(COOLDOWN_FILE): with open(COOLDOWN_FILE, r) as f: cooldowns json.load(f) last_time datetime.fromisoformat(cooldowns.get(f{service}_{instance}, 1970-01-01)) if now - last_time timedelta(minutes30): return False # 更新冷却时间 if not os.path.exists(COOLDOWN_FILE): cooldowns {} cooldowns[f{service}_{instance}] now.isoformat() with open(COOLDOWN_FILE, w) as f: json.dump(cooldowns, f) return True5.4 现象main.py --mode serve启动后CPU 占用 100%htop显示python进程在疯狂 GC原因collector模块中PrometheusAPI 轮询未加time.sleep()且requests连接未复用导致每秒创建数百个 HTTP 连接对象堆积触发 GC。解决在src/collector/prometheus_collector.py中强制使用Session并添加节流import time from requests import Session class PrometheusCollector: def __init__(self, prom_url: str, interval_sec: int 60): self.session Session() # 复用连接 self.prom_url prom_url.rstrip(/) self.interval_sec interval_sec def collect(self): while True: try: # 一次请求获取多个指标减少请求数 response self.session.get( f{self.prom_url}/api/v1/query, params{query: sum by (instance) (rate(http_request_duration_seconds_count[5m]))} ) # 处理 response... except Exception as e: print(fCollect error: {e}) time.sleep(self.interval_sec) # ✅ 关键必须 sleep6. 生产就绪技巧如何用 3 个脚本把检测结果变成可执行的运维动作跑通模型只是开始。真正的价值在于让检测结果驱动自动化处置。我一般会补上这三个脚本它们不修改源码而是作为main.py的下游消费者形成闭环6.1 脚本 1auto_scale.py—— 根据 IF 得分自动扩缩容适配 Kubernetes# auto_scale.py import subprocess import json from datetime import datetime def get_top_anomalous_instances(threshold-0.3): 从 models/isoforest/predictions.json 读取最新 IF 得分 with open(models/isoforest/predictions.json, r) as f: preds json.load(f) # preds 格式: [{instance: order-01, score: -0.45, timestamp: ...}, ...] return [p for p in preds if p[score] threshold] def scale_deployment(instance: str, namespace: str prod): 假设 instance 名 deployment 名如 order-01 - order dep_name instance.split(-)[0] # order-01 - order # 获取当前副本数 cmd fkubectl get deploy {dep_name} -n {namespace} -o jsonpath{{.spec.replicas}} current int(subprocess.check_output(cmd, shellTrue).decode()) # 规则score 越低越异常扩得越多 if instance.startswith(order): new_replicas min(20, current 2) # 订单服务最多扩到 20 elif instance.startswith(payment): new_replicas min(10, current 1) # 支付服务谨慎扩 else: new_replicas current subprocess.run(fkubectl scale deploy {dep_name} -n {namespace} --replicas{new_replicas}, shellTrue) print(f✅ Scaled {dep_name} to {new_replicas} replicas due to {instance} anomaly) if __name__ __main__: anomalous get_top_anomalous_instances() for inst in anomalous: scale_deployment(inst[instance])6.2 脚本 2trace_correlate.py—— 关联 Jaeger Trace定位代码级根因# trace_correlate.py import requests import pandas as pd def find_related_traces(instance: str, start_time: str, end_time: str): 调用 Jaeger API 查询该 instance 在异常时段的慢 trace jaeger_url http://jaeger-query:16686/api/traces params { service: instance.split(.)[0], # order-01.prod - order start: int(pd.Timestamp(start_time).timestamp() * 1e6), end: int(pd.Timestamp(end_time).timestamp() * 1e6), lookback: 1h, maxDuration: 5s, # 只查 5s 的慢请求 limit: 10 } resp requests.get(jaeger_url, paramsparams) traces resp.json().get(data, []) # 提取 span 中的 error tag 和 db.query for trace in traces: for span in trace[spans]: if span.get(tags, {}).get(error, False): print(f❌ Error span: {span[operationName]}) for tag in span.get(tags, []): if tag[key] db.statement: print(f DB query: {tag[value][:100]}...) return traces # 在 main.py 检测到异常后自动触发此脚本 # find_related_traces(order-01.prod, 2024-06-01T00:10:00Z, 2024-06-01T00:15:00Z)6.3 脚本 3report_generator.py—— 生成周报 PDF给 TL 看的“人话总结”# report_generator.py from jinja2 import Template import pdfkit REPORT_TEMPLATE h1分布式故障检测周报{{ week_start }} 至 {{ week_end }}/h1 pstrong总异常事件/strong{{ total_anomalies }}/p pstrongTop 3 故障服务/strong/p ul {% for svc, count in top_services %} li{{ svc }}: {{ count }} 次{{ (count / total_anomalies * 100)|round(1) }}%/li {% endfor %} /ul pstrong平均响应时间/strong{{ avg_response_time }} ms/p pstrong建议/strong {{ recommendation }}/p def generate_weekly_report(): # 从数据库或 CSV 读取本周统计 stats { week_start: 2024-06-01, week_end: 2024-06-07, total_anomalies: 42, top_services: [(order-service, 18), (payment-gateway, 12), (user-service, 8)], avg_response_time: 142.3, recommendation: 订单服务 CPU 异常频发建议检查库存扣减逻辑是否引入锁竞争 } template Template(REPORT_TEMPLATE) html template.render(**stats) # 生成 PDF需提前安装 wk p a hrefhttps://download.csdn.net/download/weixin_55305220/89210714 stylecolor:#ec7500;font-size:14px; 本文还有配套的精品资源点击获取 /a img altmenu-r.4af5f7ec.gif srchttps://csdnimg.cn/release/wenkucmsfe/public/img/menu-r.4af5f7ec.gif stylewidth:16px;margin-left:4px;vertical-align:text-bottom;cursor:text; /p
返回列表