Hacker Newsnew | past | comments | ask | show | jobs | submit | pmilanez's commentslogin

Is there any good benchmarks for coding harness? In terms of tokens consumption, Pi is awesome..local model etc.

But regarding coding quality, task completion, over engineering, all related to the final result that the LLM is delivering trough the harness, is there any benchmark that shows the ups and downs of each one?

I'm struggling rotating harness because of a lack os a way to truly compared what is good or not.


One recent example but also quite specific to Kimi K3 can be found here https://xcancel.com/composio/status/2083161873357111297


it's very difficult and time-consuming to make good comparisons (especially for open ended tasks) also because there are so many configuration options

  - which env do you provide?
  - which model(s)?
  - subagents?
  - system prompt (default or custom?)?
  - agents.md
etc etc

also you kinda have to look at many runs and study their traces, if you look at too few runs the outcome variability you're drawing from is too high


Feels like a problem where you have to quantify/test for your own use-case. Otherwise just go by the generic SWE-bench scores etc.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: