ARTICLE DETAIL

资讯详情

深耕郑州网站建设与运营推广的一线实战洞察。

Python后端爬虫专题26:测试不是“请求一下看看”——Fixture、RESPX与端到端旅程

Python后端爬虫专题26:测试不是“请求一下看看”——Fixture、RESPX与端到端旅程 Python后端爬虫专题26测试不是“请求一下看看”——Fixture、RESPX与端到端旅程上一篇练习完整答案target.test.evil.example的 hostname 不在精确 allowlist入口即拒绝target.testevil.example中 target.test 是 userinfo真实 hostname 为 evil.example同样在 URL 策略拒绝127.0.0.1 是字面回环地址即使错误加入 host allowlistIP 分类仍拒绝允许域 302 到内网时初始请求可以发送但解析 Location 后必须重新做 host、DNS 地址检查在第二跳前停止。超大响应的完整行为已由tests/test_http_client.py::test_fetcher_stops_reading_when_response_exceeds_budget覆盖MockTransport 返回超过预算的 bodyHttpFetcher 流式计数并抛 ResponseTooLarge调用者没有机会进入 parse_detail。三层出站设计应用仅允许精确 TargetLabCompose 网络只让 Worker 访问 targetlab/postgres/redis/minio 所需端口不暴露任意出口云/宿主防火墙拒绝元数据和内网段。allowed_private_hoststargetlab只因课程服务位于 Docker 私网不能复制到用户可控 host。一条失败测试应告诉你哪一层坏了解析器测试使用固定 HTML Fixture不走网络HTTP 客户端测试用 MockTransport/RESPX 控制状态码、头和重定向Repository 用临时 SQLiteAPI 用 ASGITransport端到端旅程把真实 TargetLab、HttpFetcher、Crawler、快照与数据库连起来。层次越低越快、定位越准层次越高越能发现接口拼接错误。只有端到端测试失败时很难知道是 CSS、重试还是数据库只有单元测试各模块都绿却可能没有正确接线。课程两者都保留。Fixture应该像真实输入不像实现抄写jobs-page-1.html包含两个详情、跟踪参数和下一页详情 Fixture 有标题、公司、城市、薪资、日期、多段描述、技能坏详情专门缺 title。测试期望使用手工确定的规范 URL 与字段不调用 normalize_url 生成 expected否则实现和期望可能一起错。页面改版时不要立刻改掉旧 Fixture。先把线上失败快照脱敏后加入回归样本写出旧/新结构的产品决策兼容两种还是明确切换。Fixture 文件名可含来源、场景和日期避免test1.html。RESPX/MockTransport负责协议场景真实公网不能稳定地产生“第一次503、第二次200”“429带Retry-After”“302到恶意host”“Content-Length撒谎”。受控 transport 可以精确安排响应并记录请求。测试应断言最终结果、等待秒数、请求头或发送次数等真实边界而不是断言 mock 对象存在。对 FastAPI 与 TargetLab 使用 HTTPX ASGITransport路由、中间件、序列化是真实的只省去 TCP 端口这比旧式 TestClient 更接近项目使用的异步 HTTPX。数据库仍是真实 SQLAlchemy schema。本篇检查点故意放一个坏详情cd project.\.venv\Scripts\python.exe-m pytest tests\test_pipeline.py::test_crawler_persists_valid_details_and_reports_broken_page-qMapFetcher 按 URL 返回固定列表/详情响应一条有效、一条缺 title。最终断言有效职位确实写入数据库、坏页进入 errors、任务报告 partial 所需计数正确、快照仍保存。若 Crawler 因一条坏数据回滚全部成功或者吞掉错误装作 completed该测试会失败。端到端旅程为什么跑三次第一次验证 created3第二次验证 ETag 导致 not_modified3修改 TargetLabStore 后第三次验证 updated1、not_modified2数据库仍三行。它捕获的不是单个函数而是 validators 从 Repository 到 HttpFetcher、304 从 FetchResult 到 Crawler 的跨模块链路。Compose 旅程再多一层真实 Redis、Celery、PostgreSQL、MinIO 和 TCP。它更慢不应替代每次几秒完成的单测在交付和部署变更时运行。外部商业网站不进入默认测试避免网络和内容变化造成假失败也避免未授权流量。如何测试重试而不让测试真的等HttpFetcher 注入 sleeper测试用 recorder 记录 1 秒而不真的睡生产默认 asyncio.sleep。这样仍验证计算出的 Retry-After 值而测试快速。时间、随机数、DNS、外部传输都是适合注入边界的依赖但不要为了测试把每个纯函数都包成接口。覆盖率不是验收标准100% 行覆盖仍可能没有断言租户隔离、重复运行和错误副作用。更好的问题是把 304 分支改成解析空 body哪条测试失败删掉 DNS 检查哪条失败删除唯一约束哪条失败每个现实变异都应被至少一个行为测试抓住。本篇完整流水线模块这次阅读聚焦异常隔离列表失败、详情下载失败、304、解析/写库失败分别怎样改变 report、快照和事务。它正是各测试层最终汇合的应用服务。把下载、快照、解析和幂等写入组织成一次可报告的采集运行。fromdataclassesimportdataclass,fieldfromtypingimportProtocolfrom.concurrencyimportbounded_mapfrom.http_clientimportFetchResultfrom.parsingimportparse_detail,parse_listingfrom.repositoryimportJobRepositoryfrom.snapshotsimportFileSnapshotStoreclassFetcher(Protocol):asyncdeffetch(self,url:str,*,etag:str|NoneNone,last_modified:str|NoneNone,)-FetchResult:...dataclass(frozenTrue)classCrawlFailure:url:strmessage:strdataclassclassCrawlReport:list_pages:int0discovered:int0created:int0updated:int0unchanged:int0not_modified:int0failed:int0errors:list[CrawlFailure]field(default_factorylist)classCrawler:一次任务的应用服务单个详情失败不会抹掉其他成功结果。def__init__(self,fetcher:Fetcher,repository:JobRepository,snapshots:FileSnapshotStore,*,detail_concurrency:int4,)-None:self._fetcherfetcher self._repositoryrepository self._snapshotssnapshotsifdetail_concurrency1:raiseValueError(detail_concurrency must be positive)self._detail_concurrencydetail_concurrencyasyncdefrun(self,seed_url:str,*,tenant_id:str,max_pages:int10)-CrawlReport:ifmax_pages1:raiseValueError(max_pages must be at least 1)reportCrawlReport()pending[seed_url]seen_list_pages:set[str]set()detail_urls:list[str][]seen_details:set[str]set()whilependingandreport.list_pagesmax_pages:page_urlpending.pop(0)ifpage_urlinseen_list_pages:continueseen_list_pages.add(page_url)try:responseawaitself._fetcher.fetch(page_url)self._snapshots.save(response.url,response.body,response.headers)listingparse_listing(response.text,response.url)report.list_pages1fordetail_urlinlisting.detail_urls:ifdetail_urlnotinseen_details:seen_details.add(detail_url)detail_urls.append(detail_url)iflisting.next_urlandlisting.next_urlnotinseen_list_pages:pending.append(listing.next_url)exceptExceptionasexc:report.failed1report.errors.append(CrawlFailure(page_url,str(exc)))report.discoveredlen(detail_urls)requests:list[tuple[str,str|None,str|None]][]fordetail_urlindetail_urls:etag,last_modifiedself._repository.get_validators(tenant_id,detail_url)requests.append((detail_url,etag,last_modified))asyncdefdownload(request:tuple[str,str|None,str|None])-tuple[str,FetchResult|None,Exception|None]:detail_url,etag,last_modifiedrequesttry:responseawaitself._fetcher.fetch(detail_url,etagetag,last_modifiedlast_modified)returndetail_url,response,NoneexceptExceptionasexc:returndetail_url,None,exc downloadedawaitbounded_map(requests,download,limitself._detail_concurrency)fordetail_url,response,download_errorindownloaded:ifdownload_errorisnotNone:report.failed1report.errors.append(CrawlFailure(detail_url,str(download_error)))continueassertresponseisnotNonetry:ifresponse.not_modified:report.not_modified1continuesnapshotself._snapshots.save(response.url,response.body,response.headers)itemparse_detail(response.text,response.url)resultself._repository.upsert(tenant_id,item,etagresponse.etag,last_modifiedresponse.last_modified,snapshot_idsnapshot.snapshot_id,)self._repository.commit()setattr(report,result.action,getattr(report,result.action)1)exceptExceptionasexc:self._repository.rollback()report.failed1report.errors.append(CrawlFailure(detail_url,str(exc)))returnreport本篇课后练习为解析器、HTTP、Repository、API、进程内旅程、Compose 旅程各写一个“最适合它发现的错误”不能重复。修改坏详情测试为三条一条正常、一条下载503、一条缺title先手算 CrawlReport 再写断言。运行全量 pytest记录通过数和耗时再说明为什么这个数字不能替代 Compose 实际旅程。下一篇会启动六个服务并进行持久化恢复演练。
返回列表