实战
1.模型选择
客观数据
DeepSWE: 衡量AI解决真实软件工程问题的基准测试。
DeepSWE score, 右上角区间内的模型most efficient
DesignArena: 由社区用户盲测的AI设计排行榜,包括h5、网站设计、游戏UI设计等。
代码分类 Preference vs Speed
Top-left quadrant shows models with low cost and high ratings
Intelligence Index vs. Cost per Intelligence Index Task
Artificial Analysis Intelligence Index · Weighted average cost (USD) per Artificial Analysis Intelligence Index task
Most attractive quadrant,区间内的模型更具性价比。
202609:
- 主力编码模型: GLM 5.3 Flash
- 顾问模型(WebUI设计规范): KIMI K3