Agent Arena 评测 Jev Router:成本高 38% 未破前沿
Arena 在 Agent Arena 上对 Jev Router 完成超 4,700 次真实智能体会话评测,结果显示其未改进当前 Pareto 前沿:与 DeepSeek V4.1 Flash (Max) 性能相当时成本高 38%、中位模型请求延迟高 1.7x。
We evaluated Jev Router by @typesafeai on Agent Arena.
Our tests spanning more than 4,700 real-world agentic sessions show:
1. Jev Router does not improve on the current Pareto frontier. For similar performance as DeepSeek V4.1 Flash (Max), the solution costs 38% more and its median model request latency is 1.7x higher.
2. However, it does mostly route to Pareto efficient models. The most LLM it picks is DeepSeek V4.1 Flash, with GPT-6.1 Sol and GPT-6 Luna also being frequent choices.
3. Jev Router’s key strength is steerability. Its score of +10% almost matches Claude Opus 5.5 (High) (+10.48). This highlights the benefit of Jev effectively routing to stronger LLM in response to user feedback.
More insights in the thread below.
来源:arena · x.com