### Feature hasn't been suggested before. - [x] I have verified this feature I'm about to request hasn't been suggested before. ### Describe the enhancement you want to request <p style="margin: 0.0px 0.0px 12.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">I’m a user of OpenCode and the OpenCode Zen API, and I have a suggestion that I think would be extremely valuable to the community.</span></p> <p style="margin: 0.0px 0.0px 12.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Would OpenCode consider </span><span style="font-family: 'TimesNewRomanPS-BoldMT'; font-weight: bold; font-style: normal; font-size: 20.00px; font-kerning: none">periodically benchmarking all model/effort-level combinations available through the Zen API</span><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none"> and publishing the results on the OpenCode website?</span></p> <p style="margin: 0.0px 0.0px 12.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Zen is particularly well positioned to provide this because you control the model/provider configurations and already test and benchmark the models you make available. At the moment, however, it’s difficult for users to answer a very practical question:</span></p> <p style="margin: 0.0px 0.0px 12.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'TimesNewRomanPS-BoldMT'; font-weight: bold; font-style: normal; font-size: 20.00px; font-kerning: none">Which Zen model and reasoning-effort level gives me the best coding performance for the money?</span></p> <p style="margin: 0.0px 0.0px 12.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Artificial Analysis provides excellent coding-agent benchmarks, but its coverage of OpenCode/Zen model-effort combinations is relatively limited. OpenCode could potentially fill this gap with an authoritative benchmark specifically representing the Zen experience.</span></p> <p style="margin: 0.0px 0.0px 14.0px 0.0px; font: 21.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'TimesNewRomanPS-BoldMT'; font-weight: bold; font-style: normal; font-size: 21.00px; font-kerning: none">Suggested benchmark</span></p> <p style="margin: 0.0px 0.0px 12.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">For every supported Zen model, test each applicable effort/reasoning level using the </span><span style="font-family: 'TimesNewRomanPS-BoldMT'; font-weight: bold; font-style: normal; font-size: 20.00px; font-kerning: none">same standardized set of coding-agent tasks</span><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">, for example:</span></p> <ul style="list-style-type: disc"> <li style="margin: 0.0px 0.0px 0.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Repository understanding</span></li> <li style="margin: 0.0px 0.0px 0.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Feature implementation</span></li> <li style="margin: 0.0px 0.0px 0.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Bug fixing</span></li> <li style="margin: 0.0px 0.0px 0.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Refactoring</span></li> <li style="margin: 0.0px 0.0px 0.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Test creation/fixing</span></li> <li style="margin: 0.0px 0.0px 0.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Multi-file architectural changes</span></li> </ul> <p style="margin: 0.0px 0.0px 12.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">For each model/effort combination, publish:</span></p> <ul style="list-style-type: disc"> <li style="margin: 0.0px 0.0px 0.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Overall benchmark score</span></li> <li style="margin: 0.0px 0.0px 0.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Task success rate</span></li> <li style="margin: 0.0px 0.0px 0.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Tests passed</span></li> <li style="margin: 0.0px 0.0px 0.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Average/median completion time</span></li> <li style="margin: 0.0px 0.0px 0.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Input/output/reasoning tokens</span></li> <li style="margin: 0.0px 0.0px 0.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Cost per task</span></li> <li style="margin: 0.0px 0.0px 0.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Failure rate</span></li> <li style="margin: 0.0px 0.0px 0.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Human intervention, if applicable</span></li> </ul> <p style="margin: 0.0px 0.0px 12.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">The resulting table could look something like:</span></p> Model | Effort | Score | Success | Time | Tokens | Cost/task -- | -- | -- | -- | -- | -- | -- GPT-5.6 Luna | low | … | … | … | … | … GPT-5.6 Luna | medium | … | … | … | … | … GPT-5.6 Luna | high | … | … | … | … | … GPT-5.6 Luna | max | … | … | … | … | … Gemini 3.8 Flash | high | … | … | … | … | … Muse Spark | xhigh | … | … | … | … | … <p style="margin: 0.0px 0.0px 12.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">I would also strongly recommend publishing a </span><span style="font-family: 'TimesNewRomanPS-BoldMT'; font-weight: bold; font-style: normal; font-size: 20.00px; font-kerning: none">Pareto frontier of quality vs. cost</span><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">, since that would make the data considerably more useful for selecting models.</span></p> <p style="margin: 0.0px 0.0px 14.0px 0.0px; font: 21.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'TimesNewRomanPS-BoldMT'; font-weight: bold; font-style: normal; font-size: 21.00px; font-kerning: none">Periodic re-testing</span></p> <p style="margin: 0.0px 0.0px 12.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Ideally, this would be repeated whenever the Zen catalog changes materially—perhaps </span><span style="font-family: 'TimesNewRomanPS-BoldMT'; font-weight: bold; font-style: normal; font-size: 20.00px; font-kerning: none">monthly or quarterly</span><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">—with historical results retained.</span></p> <p style="margin: 0.0px 0.0px 12.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">That would allow users to see not only which models perform best, but also how the economics and performance of the Zen ecosystem are changing over time.</span></p> <p style="margin: 0.0px 0.0px 14.0px 0.0px; font: 21.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'TimesNewRomanPS-BoldMT'; font-weight: bold; font-style: normal; font-size: 21.00px; font-kerning: none">Why I think this would be particularly useful</span></p> <p style="margin: 0.0px 0.0px 12.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">One of the most interesting things about Zen is that it gives users access to multiple models through a relatively consistent coding-agent environment. That makes it possible to compare </span><span style="font-family: 'TimesNewRomanPS-BoldMT'; font-weight: bold; font-style: normal; font-size: 20.00px; font-kerning: none">the actual model + OpenCode harness + effort configuration</span><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">, rather than relying on generic model benchmarks.</span></p> <p style="margin: 0.0px 0.0px 12.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">It would also help answer questions such as:</span></p> <ul style="list-style-type: disc"> <li style="margin: 0.0px 0.0px 0.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Is GPT-5.6 Luna high actually worth the additional cost over medium?</span></li> <li style="margin: 0.0px 0.0px 0.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Does increasing reasoning effort materially improve task completion?</span></li> <li style="margin: 0.0px 0.0px 0.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Which inexpensive models are genuinely competitive?</span></li> <li style="margin: 0.0px 0.0px 0.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Where is the quality/cost sweet spot?</span></li> <li style="margin: 0.0px 0.0px 0.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Which models are best for repository-scale engineering versus simple coding tasks?</span></li> <li style="margin: 0.0px 0.0px 0.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Does a model’s performance change significantly when operated through OpenCode compared with another agent harness?</span></li> </ul> <p style="margin: 0.0px 0.0px 12.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">A public benchmark like this could become one of the strongest differentiators of OpenCode Zen.</span></p> <p style="margin: 0.0px 0.0px 12.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Even a relatively small standardized benchmark would be useful. The computational/API cost should be quite manageable if the benchmark consists of a modest number of representative tasks, and the results could be published alongside the existing Zen model catalog.</span></p> <p style="margin: 0.0px 0.0px 12.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">I’d be very interested in seeing something like an </span><span style="font-family: 'TimesNewRomanPS-BoldMT'; font-weight: bold; font-style: normal; font-size: 20.00px; font-kerning: none">“OpenCode Zen Coding Agent Benchmark”</span><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none"> become a regularly updated public resource.</span></p> <p style="margin: 0.0px 0.0px 12.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Thanks for building OpenCode and Zen. It’s a very compelling approach, and I think transparent model/effort benchmarking would make it even more useful for serious users trying to choose the right configuration.</span></p> <p style="margin: 0.0px 0.0px 12.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Best regards,</span></p> <p style="margin: 0.0px 0.0px 12.0px 0.0px; font: 20.0px 'Times New Roman'; -webkit-text-stroke: #000000"><span style="font-family: 'Times New Roman'; font-weight: normal; font-style: normal; font-size: 20.00px; font-kerning: none">Michael</span></p> <br class="Apple-interchange-newline">