PostHog的Jeeves是一个会先推理再决策的Jev式分类器,带有diffusion drafter,用SFT和CISPO训练。
它是一个9B的Jev式模型(Qwen3.5-9B,LoRA,指针头),在给出判断之前先想一遍,配有block-4 diffusion drafter,以及完整的训练代码和train/dev/test数据。
![]()
在从未训练过的测试数据上,它超过Kev-9B和Jev(0.889 对 0.822 和 0.857);在JevBench的公开档位上,也超过Jev(0.935 对 0.866)。
它通过一套兼容Jev的API,在同一个请求里同时支持yes/no(noul)、多选(choice)和打分(score)三类问题。
在单张H100上,不思考约0.3秒一次请求,思考状态下中位数3.3秒。可以通过截断思维链长度来提速。
运行在CUDA上(FP8内核需要Hopper架构)。
要解决的问题
Jev式模型能给出校准过的决策概率,但准确率偏低。很多流水线因此依赖一个推理模型作为兜底。Jeeves训练了一个Jev式的Qwen3.5-9B(LoRA加指针头),用CISPO让它先推理再决策。
这带来了更好的域外任务表现,并在JevBench hard(公开档位)上超过Jev。
带思考、贪心、2560 token上限下的准确率。Kev-9B和Jev两列是Kev公布的数字。
* Kev-9B没有公开的JevBench成绩。这里是Kev-8B(Qwen3)。
所有JevBench数字都是在公开的easy、standard和hard档位上(231题)。封存的judge档位不在统计范围内,Jev和Kev的数字也限定在同一批公开题目上。
同一个checkpoint不带思考在自有测试集(2962题)上是0.804,带上思考是0.840。
跑起来需要什么
环境要求是Python 3.12加一块CUDA GPU。
pip install -r requirements.txt
下载已发布的权重并启动服务:
hf download PostHog/jeeves --local-dir jeeves-weights
python -m inference.serve --model jeeves-weights --drafter jeeves-weights/drafter_k4.safetensors --port 8009
或者把自己训练的checkpoint融合成独立模型,再配drafter服务:
python export.py runs/cispo/final --out runs/fused
python -m inference.serve --model runs/fused --drafter runs/drafter_k4/drafter.safetensors --port 8009
然后按Jev的格式发送请求:
curl -s localhost:8009/v1/systemone -H 'content-type: application/json' -d '{
"state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges on my card.",
"questions": {
"department": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"returns": "Exchanges, refunds, wrong or damaged items",
"shipping": "Delivery status, delays, lost packages",
"billing": "Charges, invoices, payment problems"}},
"escalate": {"type": "noul", "instructions": "Does this need urgent human attention?"},
"frustration": {"type": "score", "instructions": "How frustrated is the customer?",
"criteria": ["Calm", "Frustrated", "Very angry"]}
},
"options": {"max_think": 512}}'
单张H100(FP8)上的响应,三个问题并行思考:
{
"model": "jeeves-latest",
"answers": {
"department": {
"type": "choice",
"choice": "billing",
"confidence": 0.19,
"probabilities": { "returns": 0.4, "shipping": 0.14, "billing": 0.46 }
},
"escalate": { "type": "noul", "noul": 0.72 },
"frustration": {
"type": "score",
"score": 1.5,
"legend": { "0": "Calm", "1": "Frustrated", "2": "Very angry" },
"probabilities": { "0": 0.04, "1": 0.43, "2": 0.54 },
"confidence": 0.75
}
},
"usage": { "input_tokens": 129, "output_tokens": 160, "reasoning_tokens": 1536 },
"latency_ms": 8141.6
}
SDK是直接替换
sdk/ 是Jev Python SDK(typesafe-sdk)的drop-in replacement:
pip install ./sdk
from jeeves_sdk import Choice, Noul, Score, TypeSafeClient
with TypeSafeClient() as client:
result = client.system_one(
state="I was charged twice. Please help.",
questions={
"billing": Noul(instructions="Is this about billing?"),
"tone": Choice(instructions="What is the tone?", criteria={"calm": None, "angry": None}),
"urgency": Score(instructions="How urgent is this?", criteria=["can wait", "this week", "today"]),
},
max_think=768,
return_reasoning=True,
)
print(result.nouls["billing"].noul, result.choices["tone"].choice, result.scores["urgency"].score)
print(result.reasoning["tone"].text)
客户端默认连 http://127.0.0.1:8009(或用JEEVES_BASE_URL覆盖),不需要API key,等待上限120秒。
options是可选的,不发送它的Jev客户端会忽略它。服务端全局默认值通过对应的serve参数设置。
在325道dev题上:
问题、状态和答案会像这样加载进Qwen chat template:
…state…
特别声明:以上内容(如有图片或视频亦包括在内)为自媒体平台“网易号”用户上传并发布,本平台仅提供信息存储服务。
Notice: The content above (including the pictures and videos if any) is uploaded and posted by a user of NetEase Hao, which is a social media platform and only provides information storage services.