3 points, 1 comments on Hacker News
Source: [Hacker News](https://thenextweb.com/news/openai-astra-arc-agi-3-harness-62-7-vs-99-9-benchmark-revisions)
Designing an LLM Leaderboard That Can Survive Change An LLM comparison page is easy to sketch and hard to keep honest. The first version usually has a table, a score column, and a sort button. The second version has multiple model families, benchmark updates, pricing changes, provider outages, ...
An Anthropic safety researcher said there is a greater than 10% chance AI could "kill all humans" after a former colleague quits over safety concerns.
They’re going to do something good? Is this a joke? Sign up here to get an email whenever First Dog cartoons are published Get all your needs met at the First Dog shop if what you need is First Dog merchandise and prints Continue reading...
Developing story — details emerging. Check the source link for the latest updates.
Title length restricted. Why is HN focusing so much on the human drama of OpenAI and Anthropic and not the fact that AI just solved a millennium problem in a single week? The fact that AI did this means we should see more unsolved problems in all fields rapidly soon.
Pharmaceutical companies have embraced artificial intelligence in drug research almost universally, yet few company executives expect AI to actually make drugs succeed, according to a Citi survey released on Tuesday. “The biggest risk to the AI-powered drug discovery thesis is not that AI fails ...