Is there any good benchmarks for coding harness? In terms of tokens consumption, Pi is awesome..local model etc.
But regarding coding quality, task completion, over engineering, all related to the final result that the LLM is delivering trough the harness, is there any benchmark that shows the ups and downs of each one?
I'm struggling rotating harness because of a lack os a way to truly compared what is good or not.
But regarding coding quality, task completion, over engineering, all related to the final result that the LLM is delivering trough the harness, is there any benchmark that shows the ups and downs of each one?
I'm struggling rotating harness because of a lack os a way to truly compared what is good or not.