fix(dispatcher,mcp-go): 配置拉取改为后台重试,根治启动竞态

此前 dispatcher(chat)/mcp-go(embedding) 启动时一次性请求控制面配置,3s 扑空即
降级,且只能干等热更新广播——若消费方早于 gateway 启动,会全程降级(LLM 跑桩、
RAG 无向量),必须手动重启才恢复。

改为:先订阅热更新,再后台 RequestConfigWithRetry(重试至拿到配置,容忍 gateway
晚启)。新增 shared/bus.RequestConfigWithRetry + dispatcher Subscriber 包装。

验收:故意先起 dispatcher/mcp-go、后起 gateway,二者自动重试拿到 chat/embedding
配置,无需手动重启;make test-go 全绿。

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Blizzard
2026-06-24 09:57:06 +08:00
parent 56eee90705
commit 5cfed55e91
4 changed files with 32 additions and 19 deletions
+3 -8
View File
@@ -29,17 +29,12 @@ func main() {
sub := dnats.MustConnect(natsURL)
defer sub.Close()
// 配置控制面:启动时取激活模型配置,并订阅热更新。
cctx, ccancel := context.WithTimeout(context.Background(), 3*time.Second)
if cfg, _ := sub.RequestModelConfig(cctx); cfg != nil {
pool.SetConfig(cfg)
} else {
log.Println("[dispatcher] 未取到在线模型配置,降级桩运行(控制台配置后将热更新)")
}
ccancel()
// 配置控制面:先订阅热更新,再后台重试拉初始配置(容忍 gateway 晚于本服务启动,
// 避免一次性请求扑空后只能干等热更新 → 降级桩跑全程)。
if _, err := sub.SubscribeModelConfigUpdated(pool.SetConfig); err != nil {
log.Printf("[dispatcher] subscribe model config: %v", err)
}
go sub.FetchModelConfigWithRetry(context.Background(), pool.SetConfig)
// sub 同时作为 Token 回流(TokenSink)、MCP 工具调用(ToolCaller)、执行事件(ExecSink)与任务状态回写(StatusSink)出口。
orch, err := eino.NewOrchestrator(pool, breaker, eval, sub, sub, sub, sub)