feat(harness): 评测闭环 —— 评测结果落库 + 分级 + 告警 + 可查(测温计→恒温器)
此前 eval 只打日志、不闭环。现在: - 分级:evalLevel 据综合分+忠实度 → ok(≥0.75) / warn(≥0.5 或忠实<0.6) / poor(<0.5);poor 出 slog.Warn 告警。 - 落库:dispatcher 评完经 NATS(SubjectEval) 广播 EvalEvent → 网关订阅写 PG(新表 sundynix_eval, 按 task_id upsert)。沿用任务状态回写那套(dispatcher 无 DB,经 bus→gateway 落库)。 - 可查:GET /api/v1/tasks/:id/eval 返回 overall/rule/llm/faithful/level/flags/reason/sources。 - 契约 EvalEvent + EvalOK/Warn/Poor;bus PublishEval/SubscribeEval;dispatcher EvalSink(NewOrchestrator 第9参)。 验证:三模块 build+vet+test 全绿;live RAG 任务评测落库,端点返回 overall~1.0 / level=ok / faithful=1 / sources=1。 剩:桌面端质量面板、低分自动重试(P3)。project_analysis 勾掉该项。 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -22,3 +22,19 @@ type Task struct {
|
||||
Status string `gorm:"size:32"` // submitted / running / done / failed / timeout
|
||||
Detail string `gorm:"type:text"` // 失败/超时原因等(状态机回写)
|
||||
}
|
||||
|
||||
// Eval 是一次任务的自动化评测结果(dispatcher 评完经 NATS 回写,每任务一条,按 task_id upsert)。
|
||||
type Eval struct {
|
||||
BaseModel
|
||||
TaskID string `gorm:"uniqueIndex;size:64"`
|
||||
Overall float64 // 综合分 [0,1]
|
||||
Rule float64 // 规则分
|
||||
LLM float64 // LLM 质量分
|
||||
Faithful float64 // RAG 忠实度分(0=无来源未评)
|
||||
Level string `gorm:"size:16"` // ok / warn / poor
|
||||
Flags string `gorm:"type:text"` // 命中问题(JSON 数组字符串)
|
||||
Reason string `gorm:"type:text"` // 评语
|
||||
Sources int // 检索来源数
|
||||
}
|
||||
|
||||
func (Eval) TableName() string { return "sundynix_eval" }
|
||||
|
||||
Reference in New Issue
Block a user