Files
intl_news/extractor/models.py
T
simon c72a5ed13a feat: AI 模型按场景独立配置(llm_scenes)+ 去重多来源合并
- system.yaml 新增 llm_scenes(translation / daily_report,含用途/方法/模型要求说明)
- load_llm_config(scene=...) 场景覆盖;日报摘要 temperature 0.3 硬编码 → 配置
- M3 去重:唯一篇记录 source_ids(跨源重复合并,首个来源为 source_id)
- source_ids 经翻译透传至 events,日报事件 source 多来源拼接展示(≤3 个)
- 已部署 pi5:merge 实证 10 源合并;日报 report_id=224 正常入库
2026-08-12 09:56:55 +08:00

28 lines
932 B
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""正文提取数据模型"""
from pydantic import BaseModel, Field
class ProcessedArticle(BaseModel):
"""经过正文提取清洗后的文章"""
source_id: str
source_name: str
url: str
url_hash: str
title: str = ""
content: str = "" # 清洗后的正文(英文)
publish_time: str = "" # ISO 8601
author: str = ""
word_count: int = 0
md_path: str = "" # 原始 Markdown 路径
html_path: str = "" # 原始 HTML 路径
status: str = "success" # success | no_content | failed
extractor: str = "trafilatura" # trafilatura | crawl4ai_md | none
error: str = ""
source_ids: list[str] = Field(
default_factory=list,
description="去重合并后的所有来源 ID(M3 去重时跨源命中重复会追加;"
"首个来源始终为 source_id)",
)