Case Study / Version Intelligence
派蒙三千问
Paimon Asks Everything
MAIN · EFDABA8 · 2026.07.27
一个面向《原神》国际化发行场景的双语版本理解 Agent:先按地区与旅行者身份组织预热,再用受控检索和原生 Web Search 回答玩家,把匿名信号整理成制作动作,最后用增量实验判断“给谁看什么、哪些渠道真正带来变化”。五条工作流被统一进一套游戏内情报界面。
Paimon Asks Everything
Paimon Sanqianwen
MAIN · EFDABA8 · 2026.07.27
A bilingual release-intelligence agent for Genshin publishing: personalize preheat, answer through governed retrieval and native Web Search, turn anonymous signals into reviewable production actions, then test who should see what and which channels create incremental change—all inside one game-native interface.
Overview
它不只回答一个问题,而是连接一次版本发布的两端
玩家一端需要无剧透地接回复杂世界观,制作组一端需要知道大家究竟在哪里没看懂、哪项动作真的带来增量。PAIMON 把这些问题做成同一条回路:个性化预热建立语境,证据问答解决理解断点,关系图帮助继续探索,匿名信号形成发行建议,再由随机实验与 Uplift 模型验证目标人群和渠道贡献。
Interface System
把工具箱收进一套像“游戏里真的会出现”的情报界面
main@efdaba8 将首页、预热、问答、洞察和增量实验收进同一条任务导航。可折叠身份栏、路由编号、至冬观测状态、居中的语言控件与移动端底栏共享一套界面语法;深海蓝档案层负责沉浸,象牙白内容层保证长文本、矩阵和置信区间仍然清楚。
- 01 五站任务导航
- 02 路由级情报状态
- 03 桌面与移动端一致回退
Problem
版本越大,真正稀缺的不是内容,而是被正确理解的路径
新玩家、回归玩家和剧情党需要不同的背景深度;统一导览既容易剧透,也容易漏掉关键上下文。
任务文本、官方前瞻、Wiki 与社区讨论混在一起时,流畅回答很容易把推测写成设定。
发行侧看到点击与问题数量,却不一定知道该补哪段时间线、谁最受影响、两周内能做什么。
Live System
把个性化、检索与实验边界做成明确参数
地区导览
从蒙德到至冬,每个地区都有独立阵营、剧情脉络与后续引子。
旅行者身份
新玩家、回归玩家、剧情党看到真正不同的内容深度与证据层级。
检索路径
自研受控检索与 Anthropic 原生 Web Search 可按次切换并有界回退。
增量实验场景
PV × 玩家估计分群增量,达人 × 展会拆解独立、协同与联合贡献。
Product Journey
五站任务线,把“看懂版本”推进到“验证增量”
这不是五个彼此独立的 Demo 页面。它们从建立上下文开始,经过查证、探索与制作决策,再用受控实验验证动作是否真的有效,形成一条可回溯的信息路线。
先认识旅行者,再决定讲到哪里
地区与身份不只改变标题,而是决定阵营入门、剧情回顾、时间线、证据和关系图的开放程度。
- 八地区独立导览
- 三套角色化视图
- Level 3 剧透二次确认
先查证,再让派蒙开口
问题先经过意图识别、别名扩展与剧透门控,再进入可切换的两条搜索链路;来源去重、分级并通过引用校验后才生成回答。
- 自研检索 / 原生 Web Search
- 真实 URL 与来源分级
- 失败时一次有界回退
在至冬图鉴里追踪席位、状态与伏笔
图谱被重构为一张可操作的皇家地理志:星图视图强调关系,席位视图强调角色与状态;档案侧栏、关系筛选和文本线索把每次探索重新接回证据。
- 星图 / 席位双视图
- 角色状态与档案侧栏
- 双节点关系追踪
把匿名信号变成下一步动作
决策台不止汇总次数,而是给出玩家需求切片、机会与风险分数,以及带负责人、排期和验收标准的制作动作。
- AI 中文制作组简报
- 确定性机会 / 风险评分
- 证据越界自动回退规则版
先估增量,再决定给谁看什么
增量实验室把发行建议变成可检验决策:多处理 T-Learner 比较剧情、角色、玩法 PV 与不触达控制,2×2 随机实验再拆分达人、展会与协同贡献;置信区间不越过业务阈值时,系统明确建议不触达。
- 分群 CATE 与 95% 置信区间
- 1 pp 阈值 + 10% Holdout
- 2×2 归因与 Shapley 分摊
Evidence Chain
每个结论都要沿原路走回证据
默认路径从用户偏好开始,到证据回答、制作建议和增量验证结束;只有最小匿名信号与实验分流进入发行侧。模型不可用、引用越界、搜索不完整或实验区间不确定时,系统会显式降级,而不是用更自信的语气遮住失败。
- 01 身份与地区
- 02 主题编排
- 03 受控检索
- 04 来源治理
- 05 引用校验
- 06 匿名决策信号
- 07 随机分流
- 08 增量验证
Reliability
可信度不是一句承诺,而是失败路径也经过设计
- 官方、可信 Wiki、社区与未知网页分层处理,社区推测不会被包装成官方设定。
- 无最终正文、来源无关、请求超时或截断时,每次请求最多回退一次,并清理无法映射的引用编号。
- 原始问题默认不保存;没有模型密钥时仍可运行确定性证据回答,没有 Supabase 时切换到本地事件存储。
- 固定种子合成实验只验证方法链路,不宣称真实业务效果;Vitest 与 Python unittest 同时约束引用治理、特征泄漏、置信区间、Holdout 和 2×2 归因算术。
Current State
已经上线,也把边界写在产品里
截至 main@efdaba8,个性化预热、双搜索链路、引用治理、发行决策台、双视图至冬图鉴与发行增量实验室都已进入公开 Demo。增量实验室目前使用固定种子合成随机数据,只能证明方法和产品闭环可运行;真实留存提升仍需线上随机 Holdout 或地域实验验证。Vercel 冷启动、站外趋势未接入与社区 Wiki 覆盖边界也会被明确展示。
Overview
It connects both ends of a release, not just one question.
Players need to re-enter a dense world without being spoiled. Release teams need to know where understanding breaks—and which action actually creates lift. PAIMON connects both ends: personalized preheat establishes context, evidence Q&A resolves confusion, the graph supports exploration, anonymous signals become release recommendations, and controlled experiments validate target audiences and channel contribution.
Interface System
Turn a toolkit into an intelligence surface that could live inside the game.
At main@efdaba8, Home, Preheat, Ask, Insights, and the Incrementality Lab sit inside one mission rail. Collapsible identity controls, route codes, Snezhnaya archive status, centered language controls, and mobile navigation share one grammar: deep-navy archive layers create atmosphere while porcelain surfaces keep evidence, matrices, and confidence intervals readable.
- 01 Five-stop mission rail
- 02 Route-level intel status
- 03 Consistent desktop/mobile fallback
Problem
As releases grow, the scarce resource is a path to correct understanding.
New, returning, and lore-focused players need different context depth; one briefing either spoils too much or explains too little.
Quest text, official previews, wikis, and community threads blur together, so fluent answers can quietly turn speculation into canon.
Counts alone do not tell a release team which timeline to clarify, who is affected, or what can ship in the next two weeks.
Live System
Personalization, retrieval, and experiment boundaries are explicit parameters.
Region guides
Every region from Mondstadt to Snezhnaya has its own factions, story path, and forward hook.
Traveler profiles
New, returning, and lore-focused players receive genuinely different depth and evidence views.
Search routes
A controlled in-house route and Anthropic native Web Search switch per request with bounded fallback.
Incrementality scenarios
PV × player estimates segment lift; influencer × expo separates direct, interaction, and joint contribution.
Product Journey
Five mission stops move from “understand the update” to “validate the lift.”
These are not five disconnected demo pages. They establish context, verify claims, support exploration, form production decisions, and then test whether those actions create incremental value.
Know the traveler before choosing what to reveal.
Region and profile determine which faction primer, recap, timeline, evidence, and graph layers are actually shown.
- Eight independent region guides
- Three role-specific views
- Second confirmation for level-three spoilers
Verify first. Let Paimon speak second.
Intent, alias expansion, and spoiler gates precede two switchable search routes. Sources are deduplicated, graded, and citation-checked before generation.
- Controlled search / native Web Search
- Real URLs and source tiers
- One bounded fallback on failure
Track seats, status, and unresolved clues inside the Snezhnaya Atlas.
The graph is now an operable royal atlas: Constellation view prioritizes relationships, Seats view prioritizes characters and status, while the dossier panel, relation filters, and text trails reconnect exploration to evidence.
- Constellation / Seats views
- Status-aware dossier panel
- Two-node relationship tracing
Turn anonymous signals into the next production action.
The decision center goes beyond counts: player-demand slices, opportunity and risk scores, plus actions with owners, schedules, and acceptance criteria.
- Chinese AI production brief
- Deterministic opportunity / risk scores
- Rule fallback when citations leave the whitelist
Estimate lift before deciding who should see what.
The lab turns recommendations into testable decisions: a multi-treatment T-Learner compares story, character, and gameplay PVs against no contact, while a 2×2 randomized design separates influencer, expo, and interaction effects. If uncertainty does not clear the business threshold, the decision is explicitly hold.
- Segment CATE with 95% intervals
- 1 pp threshold + 10% holdout
- 2×2 attribution with Shapley allocation
Evidence Chain
Every conclusion has to walk back to evidence.
The path starts with player context and continues through evidence answers, production recommendations, and incrementality validation; only minimal anonymous signals and randomized assignments reach the release side. Model failure, invalid citations, incomplete search, or uncertain intervals trigger explicit degradation.
- 01 Profile + region
- 02 Topic orchestration
- 03 Controlled retrieval
- 04 Source governance
- 05 Citation validation
- 06 Anonymous decision signal
- 07 Random assignment
- 08 Incremental validation
Reliability
Trust is not a promise. Failure paths are designed too.
- Official, trusted-wiki, community, and unknown-web sources are separated, so community theories do not become official canon.
- Missing final text, irrelevant sources, timeouts, and truncation trigger at most one fallback; unmapped citation numbers are removed.
- Raw questions are not stored by default; deterministic evidence answers still run without a model key, and the event layer falls back locally without Supabase.
- Fixed-seed synthetic experiments validate the method—not real business lift. Vitest and Python unittest cover citation governance, leakage checks, confidence intervals, holdouts, and 2×2 attribution arithmetic.
Current State
Live today, with its boundaries visible.
As of main@efdaba8, personalized preheat, both search routes, citation governance, release decisions, the dual-view Snezhnaya Atlas, and the Release Incrementality Lab are live in the public demo. The lab currently uses fixed-seed synthetic randomized data, so it proves the method and product loop—not real retention lift, which still requires online holdouts or geo experiments. Vercel cold starts, missing external trends, and community-wiki coverage limits remain visible.