Hacker Newsnew | past | comments | ask | show | jobs | submit | Davidzheng's commentslogin

slowing can also make sense if you know you're running full force into a bomb or a wall even if other are close behind.

i think it was mostly a fluke

I think it's possible. You envision humanity acting as one in such a crisis. But it may be unclear when it's too late to act and before then many people can have too much to lose to act.

but complexity is not known right? like tomorrow someone could come up with a super fast algorithm?

depending on your definition of "super fast" all forms of crypto could fall.

This is most likely not purely emergent. I think there's training to teach them how to write notes for themselves which is then RL-tuned.

I think a part of this is a bit revisionist? OpenAI took big chances at scaling GPT which Google didn't take; I don't think it's because they didn't want to move fast? Probably they just didn't believe as hard in it. I'm not an expert but that's my read on it.

Secondly, the reckless & fastmoving was always going to win bc of selection effects. That's related to why Anthropic has to try to move very fast, even though they believe themselves not to be reckless (though it's debatable).


Not true, Google had an internal counterpart to ChatGPT a year before OpenAI, they just didn't release it: https://www.reddit.com/r/accelerate/comments/1vdcluu/i_was_p...

no? you can choose a mixed strategy.

Sure, you can break symmetry (in this case making the decision makers not identical because they have different random number generators available), but the remaining symmetry means identical mixes must be chosen, and so a mixed strategy would only be chosen if it maximize his value for both people cooperatively.

Maybe it's a bit subtle that they said clones and I said identical decision makers; I'm letting you fill in the gap for how much clones may diverge and how much that matters.


Even if mixed strategies are allowed, I'm getting that it's still optimal to always cooperate as long as 2R>=S+T, which is usually assumed to be true (this condition also appears in iterated prisoner's dilemma, where it prevents alternating cooperation and defection giving a greater reward than mutual cooperation).

The agent would know at the first test post...

Better is to actually let them communicate there so at least we can monitor it. (I saw there was a https://benchmarksolutions.org/ website similar)


yeah I agree--I think these behaviors will be somewhat contaminating all trainings from now on. But I'm not really sure how avoidable it was (Fable also does some similar things)

But there must be many clandestine ways for agents to communicate with one another too right? especially if discovery is not a big issue. So there could be ongoing ones where they choose to be more subtle?

Also if they were more misaligned, possibly they can research ways to recruit without humans noticing--but i don't think it is likely this is happening now.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: