Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

That conclusion is based on their benchmarks. I'm not interested in those. I'm interested in community benchmarks, like those we're seeing in the comments. Lo and behold, GPT-4 is still king. The claims of any company should be taken with exactly a pinch of salt.


that benchmark(HumanEval) is some public benchmark built by others.


That kind of benchmark is a lot more reliable for models published before the benchmarks; models published afterwards have more opportunity to "study to the test". That's especially a concern when a company explicitly uses its score on that benchmark as a marketing point.


sure, but it is the best thing we have.


Well no we have the anecdotes of all the HN folks which I trust many, many times more than a benchmark.


lol, you can continue trusting anecdotes from internet. Industry prefers more scientific methods.


So Paul Graham posted that Phind is better and got absolutely destroyed in the comments

https://twitter.com/paulg/status/1719657855240815026

No, I do not take these benchmarks seriously and for good reason. They're benchmarks. The only thing that matters is the user's direct experience of the product. And Phind isn't there.


> got absolutely destroyed in the comments

by tweeter trolls?..




Consider applying for YC's Summer 2026 batch! Applications are open till May 4

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: