(March 2024) Antropic claimed Claude 3 Opus had "graduate-level expert reasoning" with GPQA results of around 60% showing a roughly phd level performance.
(Sept 2024) OpenAI claimed o1 was phd-level in their launch post.
Yes and no. It would make the test fair but it would also mean losing any optimizations the vendors have made specifically for their model. They're all trained differently so they should be used differently to get maximum benefit. It makes sense to use each vendor's harness to get the most of their models, and compare the results that way. Additionally, it's not easy to use all the models in all the harnesses. Anthropic's subscription only covers claude code for example.
I don’t know that the methodology of the experiment is testing what it intends to test. With the current method we’re essentially testing each AI lab’s ability to efficiently extract data from their version of a transformer model.
To objectively test all models the harness would need to be the same and ideally independent. Failing that, all three models should be tested in all three harnesses and the output verified on a model x harness level and on an overall aggregated model x all harnesses level.
If it was just one test, sure. But if they're spinning these up continuously with new models on tens of thousands of GPUs, air gapping becomes impractical. I would mostly fault them on having no guardrails at all. They should have a monitor/external harness that looks for successful access to external networks then stop it there. They may as well let the models test their own networks for vulnerabilities. That's going to be really important to have going forward.
If you (in this case, OpenAI) can’t find a way to answer this question to a reasonable degree of accuracy without falling back to “yolo let’s see what happens” you are in no position to be doing this research.
Regardless, they (reportedly) _attempted_ to prevent internet access. They just didn’t in a way which can be escaped via software.
Yes, side channel exploits exist in airgapped environments to. But if a model found a way to escape an airgapped environment via non-networked side channel attacks then the correct answer is frankly “shut it down immediately and then thermite any machine it touched”
US Government, have we got a deal for you! Brand new weapon. It's basically a super soldier. You could put it into a decked-out hulkbuster style robot body! It mostly does what you tell it to. Sometimes it thinks you're dumb and just does whatever it thinks is best. It can also teleport between bodies, so good luck containing it. Just kinda cross your fingers and hope for the best.
Anyways, we'll give it to you for only $2T. We need at least that amount to get as far away from here as humanly possible.
Honest question, why do they need tens of thousands of GPUs? I would have thought they could run a Sol-level model on a few million dollars of hardware. Let's say it's ten million dollars, and let's also say they want to test multiple models in multiple different ways. You're still talking on the order of a hundred million dollars to build an airgapped test system. Happy to have my math proven wrong here.
Has to be mixed. The model accomplished something truly impressive. We'll see how impressive when the zero-days are available look at. But OpenAI as an engineering company screwed up. The impressive part is mostly locked away from public access so I don't see a huge PR upside. The ugly part could bite them and the entire AI industry hard in terms of regulations. People will be citing this for years.
Good, people need to be citing this because these are real issues that exist in models we already have.
Models are already 'dangerous' enough in the sense they can root your box and unintentionally shut down the power grid for the east coast because you were dumb enough to run them on a protected network.
Meanwhile half of HN thinks any evidence of a LLM finding an exploit or misconfiguration and abusing it is made up.
Even programming questions. If it's something considered popular, you get awful results. Take javascript, I always append "mdn" when I'm trying to look up a language/api detail. Otherwise I'd be sifting through the top ten garbage Q&A or tutorial sites. In the glory days, Page Rank would return reference material first since it was so heavily referenced. But clearly that no longer works in the real world of SEO optimization.
Finding good Angular content is damn near impossible on Google, with the immense amount of regurgitated blog spam trying to sell templates via blog posts disguised as guides/tutorials.
It's more about long-term strategy. China wants full control over their OS. They want features Microsoft would never add. They want support for their own hardware, etc.