Inkhaven resident Benjamin Sturgeon discusses one of his Inkhaven publications. Even in the AI safety endgame, outsiders can still become live players by finding the right leverage points. Ben Pace speaks with him at Inkhaven, a residency for writers and bloggers.
Sponsored by WordPress.com
https://www.inkhaven.blog/about
https://www.benjaminsturgeon.com/inkhaven-day-20
00:00 Intro
03:48 Reading
12:08 Interview
45:39 Sponsor
Ben:
Welcome to Ink Haven Presents Readings from the Archive. I’m Ben, and in this podcast, I’m inviting writers from the internet to bring something that they have written, uh, to read it, and for us to discuss it. My guest today is Benjamin Sturgeon. Uh, Benjamin is currently part of the, um, extension program with MATS, uh, being advised by some researchers at the UK AI Security Institute, part of the government there. Um, he’s likely soon starting a PhD at Imperial College London and is the strategic director for, uh, AI Safety South Africa. Welcome, Ben. Thanks for chatting with me today.
Benjamin:
Yeah. Thanks for having me.
Ben:
Uh, what, uh, what have you brought to read today?
Benjamin:
So the piece is titled, “Is It Too Late to Make Contributions to AI Safety From Outside the Labs?”
Ben:
Mm-hmm.
Benjamin:
Or something very close to that.
Ben:
And, uh, yeah, why did you choose it?
Benjamin:
I chose it because, uh, it feels like the most important work that I was able to do, at least in terms of giving a, a message that I thought would be useful for people to receive.
Ben:
Mm-hmm. And I think it’s about the, uh... It was inspired by reading the Mythos, uh, uh, model security card.
Benjamin:
Uh, the system card, yeah.
Ben:
System card.
Benjamin:
And yeah, I, I felt there was a huge amount of interesting stuff in the system card, but there wasn’t a lot of easy ways to synthesize what was actually in there. So I just wanted an excuse to dig into it, and I was feeling quite a strong sense of, like, hopelessness in the wake of the release, wondering like, you know, is there any point in doing technical AI safety research? You know, what is this all for if, you know, we’re already at such an advanced stage, uh, where models are becoming sufficiently capable that, you know, they can do almost superhuman, uh, cyber offensive, um, you know, zero day research. And for context on what Claude Mythos preview is, essentially, Anthropic recently publicly stated that they would not be releasing their most powerful model and had been doing a whole bunch of research into just how capable it was, and its benchmarks were far above anything else that we’d seen. Um, specifically in some, um, in the code department, the models had been trending between fifty-four and fifty-nine, roughly, on SWE-Bench Pro, which is a benchmark that tracks coding ability. And Claude Mythos preview hit around seventy-seven point three, uh, which was a massive step change. And as a part of this, it was capable of discovering novel zero day vulnerabilities, which would allow it to completely take over, um, remote machines. And normally, only a handful of these are reported every year in the order of, like, less than a hundred, and they were able to find ten, um, tier five vulnerabilities, which are the most extreme type of vulnerability, um, in just a few weeks of testing.
Ben:
Mm-hmm. Mm-hmm. Yeah, it was quite, uh, concerning. I think, uh, I don’t know if it’s safe to say this, but, uh, I think we just got a text in our Less Wrong intercom the other day with someone, uh, providing some new, uh, vulnerabilities.
Benjamin:
Oh, wow. Okay.
Ben:
Yeah. I’m sure we’ll be fixing. Um, okay. Well, uh, yeah, let’s, uh, let’s... I’m interested to hear this, so let’s listen to Ben’s reading of this essay. Thanks very much for reading it. I enjoyed listening to your extremely calm and careful reading of your essay. Um, did you notice anything differently reading it out loud than when you’d just written it in text?
Benjamin:
I realized I should have, uh, added a bit more of a conclusion.
Ben:
I also kind of found that.
Benjamin:
‘Cause it ends quite suddenly, yeah. Yeah, I, I talked about the fact that all of these big research insights and techniques had only been developed in the last few months, and I really want to kind of emphasize this fact. You know, there are still huge gaps in our ability to interpret what the models are actually doing, and it seems extremely valuable that something like activation verbalizers were created. Um, and this was a highly contingent event. You know, I was talking to Adam Kovenen when he was doing the research, ‘cause we were doing the same math stream, and he kept, um, you know, like, trying a different direction, and it didn’t really work out, and he would just move on and try, like, a different direction. So every time I would talk to him, he was, like, trying out some completely different set of ideas. And then I didn’t talk to him for a few months, and then the paper dropped, and I was like, “Wow, this is, like, such an impressive, um, result and finding.” And yeah, just a process of iterating and thinking about, you know, how can we really find the most valuable problem here. Adam’s not like a god. He’s just, he’s just a man.
Ben:
He’s but a man.
Benjamin:
He is but a man. And yeah, I think it’s just possible for other people to do this kind of work.
Ben:
Are there not... So there had not been many other things that had given you this feeling for some time, that, uh, researchers external to the labs, uh, would be able to have positive influence on the understanding of the models?
Benjamin:
Yeah. So I was speaking to a number of people who are just kind of unsure whether any of the work that they do will actually have any impact.
Ben:
Yeah.
Benjamin:
Because, you know- If you’re not at the labs, they probably, like, don’t care to use it. Um, labs do have somewhat of a preference for, like, you know, having the work that they do be researched in-house. So Owain actually also talked about this in his talk the other day. I think part of what made the AV work, um, you know, be integrated so quickly was that Adam eventually started collaborating with some of the people from Anthropic in order to produce the work. And so I think, you know, that accelerated the pipeline into them adopting it.
Ben:
Hmm.
Benjamin:
Um-
Ben:
Yeah, I think, uh, also when Greenblatt worked with them on the, uh, alignment faking stuff, I think that became a bigger deal than had Greenblatt tried to do some of that work externally to Anthropic.
Benjamin:
Yeah. Yeah, totally. I think, um, you know, it is very, very helpful to just integrate oneself into the lab’s process, but they are open to doing this kind of thing. And even if not, you know, if the wor-
Ben:
Yeah
Benjamin:
... like, everyone re-reads Owain Evans’ latest papers within the labs as soon as they drop.
Ben:
Hmm.
Benjamin:
So you know that they’re thinking about this stuff-
Ben:
Mm-hmm
Benjamin:
... even if they aren’t, like, necessarily collaborating directly with you.
Ben:
Yeah. Um, I think I-- you’ve probably paid much more attention to a lot of the AI safety research than I have over the past chunk of years. Um, was there, uh... When did you start to feel this pessimism more? Was, uh, was there previous times when there was research coming out that you felt was being, like, integrated? Or were you merely coming to realize that, uh, none of it was just recently? Or was there a change?
Benjamin:
Part of the problem is that we just don’t know how much is being integrated.
Ben:
Hmm.
Benjamin:
You know, it’s quite a black box. And something like the Mythos system card is kind of a huge amount of information that actually, like, tells you a lot about what the labs are actually doing and what they’re thinking about-
Ben:
Mm-hmm
Benjamin:
... and what techniques they’re using. If you have no idea whether or not any of the research is actually going anywhere-
Ben:
Mm-hmm
Benjamin:
... then, you know, by default, you assume that it’s not being used. Um, so it can be easy to become discouraged from that. And so, partially just doing that close reading of the system card was very helpful in-
Ben:
Mm-hmm
Benjamin:
... you know, becoming a bit more optimistic about what technical research can actually do.
Ben:
Uh, how long have you been working in AI safety?
Benjamin:
I’ve been trying to contribute directly in the field, uh, for about two years now. I dro- I quit my job doing machine learning engineering consulting work at a company in South Africa. Then I started my master’s and founded AI Safety South Africa. Yeah, from there, it was, you know, a process of bootstrapping and trying to build connections with, uh, people at the frontier of research. But yeah, doing maths was, like, really, really helpful for s- um, supercharging that process.
Ben:
I kind of want to just, um, ask you a fairly random... There’s, uh, one of the five things in it were about, uh, I think measuring, tracking the emotions that the model was, quote-unquote, “feeling.”
Benjamin:
Yeah.
Ben:
And, uh, using those as guidances. Um, can you say a little bit about how you, how one tracks the emotions of a model? Is it... I actually just don’t know how that works.
Benjamin:
My friend, Anna Soligo, was doing some of this work during the Anthropic Fellows program, and I think there was also an independent, uh, stream of researchers in Anthropic, uh, working on this completely independently.
Ben:
Hmm.
Benjamin:
So it was interesting that, and multiple people, like, converged on it at the same time. But essentially, what you wanna do is develop, like, specific vectors that are attuned to specific emotions. And so one way you can do this is using contrastive pairs of, like, many, many examples of the model-
Ben:
Mm-hmm
Benjamin:
... like, feeling a certain way or describing a certain feeling. And then another set of examples where the model is feeling something else or, you know, just describing, like, completely neutral topics.
Ben:
When you say the model is feeling, do you mean, like, uh, what would cause it to output this particular token or set of tokens?
Benjamin:
Yeah, so an example would be something like, um, you know, “I watched,” um, “I watched sadly as, you know, the child was driven away in the car.” And then a contrasting version of that would be, “I watched as the child was driven away in the car.”
Ben:
Mm-hmm.
Benjamin:
Um, and then the sadness part, uh, just being included in that sentence, will change the underlying activations, such that they actually capture some of how the model, like, processes that emotion. And then when you take the difference between these two examples, uh, and just take minus the activations from each other-
Ben:
Hmm
Benjamin:
... you get this clean emotion vector. And you do this over a bunch of examples, and then you can just add this into any future, um, like, forward pass-
Ben:
Mm-hmm
Benjamin:
... and it’ll actually, like, inject the emotion into it.
Ben:
I see.
Benjamin:
Yeah.
Ben:
It did seem like it would have to have an impact on the labs. This is not the, I think, the measuring emotions one, but, uh, I think one of the other two you mentioned that, uh, would find parts of the AI saying, thinking things like, “Avoid suspicion,” and, “Clean up to avoid detection.”
Benjamin:
Yeah.
Ben:
These are very blatant, uh, concerning sentences.
Benjamin:
Yeah.
Ben:
Um, yeah. I mean, I, I admit I am not as... Or I, I don’t feel as confident that, uh, uh, one can have, mm, one can have, uh, arbitrary influence from, uh, um, outside, from outside the labs if the bar is, has to show blatant misalignment.
Benjamin:
Hmm.
Ben:
Uh, but you can still definitely, there’s still some work, some good work to be done there.
Benjamin:
Yeah.
Ben:
Are you more optimistic than that, or are you about as optimistic as that?
Benjamin:
You know, it was really the technique was able to elicit the evidence of the misalignment in a way that previously would’ve been impossible. Um, and, you know, the model, if you look at the chain of thought, you know, it doesn’t mention anything about, you know, sneaking or, like, you know, hiding its, its path.
Ben:
Hmm.
Benjamin:
Um, it’s purely, like, in the case of the activation, um, verbalizers, you can take the activations from the sentence, and specifically the code itself that it wrote, um, to, you know, like, hide what it was doing. And just taking the code, you run that through the activation verbalizer. It takes the activations and can pass out some of the, like, things that the model is actually thinking-
Ben:
Hmm
Benjamin:
... as it’s writing that code, and turn it into natural language, where, which are act- that’s actually what those sentences said. Um, you know, like, “I need to hide my, my traces here.”
Ben:
Yep.
Benjamin:
Um, so yeah. Like, if you were just reading this, um, transcript, you would have no idea that there was any misalignment occurring at all.
Ben:
Yes.
Benjamin:
Um, so that’s really what makes that such a powerful tool.
Ben:
Um, I haven’t read the code. What did Anthropic do as a result of finding these concerning things?
Benjamin:
Well, they chose not to publish the model, for one.
Ben:
Did they... Yeah, did they make any other efforts to, uh, make it more aligned and not have these concerns?
Benjamin:
Yeah, I mean, it’s somewhat of a challenging situation because, you know, while they did observe these behaviors, um, according to the metrics they use to measure alignment, it was the most aligned model they’ve ever trained.
Ben:
That’s, uh, in some ways, uh, concerning.
Benjamin:
Yeah.
Ben:
Uh, but yeah. Interesting.
Benjamin:
I mean, I think Ryan Greenblatt recently had a very popular Less Wrong post saying, you know, “The models don’t seem very aligned to me.”
Ben:
Yeah.
Benjamin:
And I think this is probably, um, accurate, in that you can observe these, um, misaligned behaviors directly, in that they’ll overstate what they were able to achieve, and they will, you know, gaslight you into thinking they did better than, um, what they actually were able to do-
Ben:
Mm
Benjamin:
... or bothered to do. And they just don’t, like, fully execute on some of the things that you tell them. We can, or Anthropic can still successfully say, you know, there’s this gap between perfect alignment and what we can observe. Um, but what these specific techniques in this post, uh, are pointing at is that they allow us to go much deeper in identifying, um, like, true misalignment failures in, like, a very clear way. And so they also, you know, reveal a path by which we can, like, optimize against the failures even further until, you know, we truly aren’t seeing any, um, traces of, like, deception in the model’s, like, deepest thoughts. And, you know, th- this is, like, goes deep into, um, the, like, theory on AI safety literature, which is that, you know, if you optimize against what the models are deeply thinking, eventually, you know, you just can’t detect it, but the misalignment is still there.
Ben:
Mm-hmm.
Benjamin:
Um-
Ben:
Train against the detection rather than the, uh, the underlying cause.
Benjamin:
Yeah. But at the same time, um, you know, there’s a significant, like, decoupling between, um, the activations as they are measured by something like an activation verbalizer and the actual reward function, because we’re not just, um, training the model to, you know, not think deceptive thoughts. Uh, we just keep identifying, oh, like, its deepest thoughts-
Ben:
Mm-hmm
Benjamin:
... were deceptive, then, like, we need to just make it do more training that’s, like, aligned with, um, you know, how we want it to behave. Like, alignment training, basically.
Ben:
Yep. As usual, my brain just goes to pretty in-depth sci-fi stories of the AIs, like, thinking, “What’s happening to me in this complicated situation? Could it possibly be that it doesn’t like my thinking here, so that I get different outputs there?”
Benjamin:
Yeah.
Ben:
Uh, and then, uh, you know, Jensen Huang will yell at me that we are serious people and this is not a sci-fi situation. Um, did you pu- uh, publish this piece on Less Wrong?
Benjamin:
Um, I haven’t published it yet, but I will.
Ben:
It struck me as a good Less Wrong post.
Benjamin:
Yeah. Maybe once I add the conclusion, then I’ll publish it on Less Wrong.
Ben:
Have you put any of the pieces so far out in Cayman on Less Wrong?
Benjamin:
Yeah. I’ve just tried to have a bit of a higher bar. Um, and this one I thought was worth just spending a little bit more time-
Ben:
Yeah
Benjamin:
... um, polishing and getting feedback from Scott Alexander and stuff.
Ben:
Oh, have you... Has that happened?
Benjamin:
Yeah, yeah.
Ben:
Oh, that’s good.
Benjamin:
Yeah.
Ben:
What kind of feedback did he give?
Benjamin:
Um, the main feedback was on, like, the structure of talking about specific examples, uh, within the activation verbalizers and the emotion vectors. He said, like, I went a bit too into the, the weeds with specific examples. And so I actually, like, took out some of those details.
Ben:
Oh, interesting.
Benjamin:
Um, and then I also just tightened up the structure a little bit.
Ben:
Yeah. I think... I don’t know if this is even accurate, but I think I like how short this essay is. I think it, it’s no longer than it needs to... Uh, most machine learning related papers can get much too long unnecessarily. They can explain a lot of detail that isn’t relevant to the point. Um-
Benjamin:
Yeah. I didn’t really want it to be, like, a machine learning, um, piece.
Ben:
Yeah. It’s perhaps a selfishly curious question, but you, uh, you mentioned everyone reads Owain Evans’ papers when they come out.
Benjamin:
Mm.
Ben:
Do you also find, uh, where else do you get information from, and how much is Lesswrong a, a part of the place you get information from about the field?
Benjamin:
Like, Owain also takes Lesswrong very seriously. You know, like they publish everything-
Ben:
That’s true
Benjamin:
... to Lesswrong as well. I’d say most serious researchers, like, post on Lesswrong. Uh, you know, the shortened Lesswrong post version of their paper.
Ben:
Mm-hmm.
Benjamin:
And often that’s just the one I’ll read.
Ben:
Oh, that makes sense.
Benjamin:
Um, but also Twitter.
Ben:
Yeah.
Benjamin:
Um, is-
Ben:
‘Cause it’s got better discourse.
Benjamin:
Yeah. The people are just so, so reasonable and so wise.
Ben:
Mm-hmm.
Benjamin:
Um-
Ben:
Well, you were feeling more pessimism about AI safety, uh, maybe a month ago. I’m not sure, how long has Mythos card been released?
Benjamin:
Um, I think it’s been roughly three weeks.
Ben:
Three weeks.
Benjamin:
Yeah.
Ben:
Um, curious whether you were considering, had other alternatives in mind for what you would spend your time on, whether it was, uh, uh, some, I don’t know, pause AI kinda thing, whether it was a, “One thing, I’m just gonna go and go home and have a nice rest of my next few years,” or some third thing.
Benjamin:
Yeah. Um, I guess in general, I have like a bias against, like, extreme actions. Um, like I’ve kind of chosen what I’ll be doing for the next few years, so my day-to-day feelings, like, don’t feature that strongly in, like, what I actually will be doing.
Ben:
Mm-hmm. Yeah, I think similarly at, like, like home. Like, I think over the past, you know, near decade of existing, there have been various weeks, uh, maybe two weeks or so, where, uh, my boss has been like, “I think everything we’re doing is bad and we should stop.”
Benjamin:
Mm.
Ben:
We did not immediately stop.
Benjamin:
Yeah.
Ben:
We were like, “We should maybe continue for a while and see whether our epistemic state changes.”
Benjamin:
Yeah. Yeah, I think like-
Ben:
We may currently be in that one. We’ll have to find out.
Benjamin:
Yeah. Like, I think often, um, you know, those emotional states are more about, often about confusion, you know, don’t actually have enough information. Yeah, going and looking through the Myth- the Mythos system card was able to, like, change my view and change my feelings as a result.
Ben:
Yeah. I think confusion is accurate. I think I would probably frame it more as sometimes being about, um, the true hypothesis not being in your current hypothesis space.
Benjamin:
Mm-hmm.
Ben:
You’re like, “Cool, I have a good hypothesis and a bad hypothesis. I have falsified the good hypothesis. I’m gonna look around for other hypotheses, and in the meantime, be very concerned and upset that it’s only looking currently like the bad hypothesis.”
Benjamin:
Yeah, exactly.
Ben:
I think I wanna change track a little bit and, uh, talk about Inkhaven more generally.
Benjamin:
Yeah.
Ben:
Um, why did you come to Inkhaven?
Benjamin:
I felt that I had a lot of resistance towards, um, publishing stuff in general.
Ben:
Mm-hmm.
Benjamin:
And it would take me, like, a long time to actually share things, and I didn’t have that great of a sense of, like, the benefits of sharing stuff.
Ben:
Mm-hmm.
Benjamin:
You know, like what, how do I actually, you know, do the, the EV calculation on putting the effort into a post? So I wanted to get more data on that, and then I also wanted to... I believe that, like, writing and, like, putting stuff out there is kind of, um, a superpower, in that writing kind of builds up over time and-
Ben:
Yeah
Benjamin:
... kind of does work for you while you sleep. Um, so it’s a very powerful form of leverage.
Ben:
That’s, uh, another reason why I personally prefer Lesswrong to Twitter, which is that, uh, I feel more like Lesswrong is building up a body of work, and more that, like, great posts historically will continue to be linked and known of. Whereas, like, it’s hard to remember people’s tweets from long, from months ago, or link to them-
Benjamin:
Yeah
Ben:
... or search for them.
Benjamin:
Yeah, that’s definitely true. I kinda want to do a project where I take all my bookmarks on Twitter and turn them into a personal, like, wiki, so that I can actually spend more time going through them, because I feel like I bookmark... I, I religiously bookmark everything that I think I might wanna go back to.
Ben:
Hmm.
Benjamin:
Um, and then I never, like, actually go back. Um, but there’s a, a lot of bangers in there.
Ben:
Yeah. I don’t know whether you’ve had time to reflect yet on, uh, have you made any updates about the, how to make the EV calculations on how much effort to put into a piece of writing before sharing it?
Benjamin:
Yeah, I mean, my two most popular posts, um, are still the ones that I spent, like, two months writing, um, which I published before I came to Inkhaven.
Ben:
Mm-hmm.
Benjamin:
And the ones I have posted on Lesswrong so far have gotten, like, maybe, like, 10 or 20 upvotes. It turns out the, the quality uplift was pretty significant.
Ben:
Mm-hmm.
Benjamin:
But at the same time, like, I’ve been watching Lawrence Chan post every post on Lesswrong, and immediately get, like, 50 upvotes, which has been like, wow, like how does he do it? Uh, his stuff is really good.
Ben:
It’s funny, I think, I forget exactly. I think he’s a little disappointed in his writing, so it’s funny that everyone was like, “Oh, I wish it got better reception.”
Benjamin:
Yeah, I mean, I actually have still been quite careful about posting on Lesswrong, so everything I’ve posted has just been on my website, um, which gives me time to actually go back and, like, edit the things a little bit more. Um, which is maybe, like, a form of procrastination.
Ben:
Yeah, I was thinking that, I would, I would imagine that, uh, a good way to do it would be to get all your Inkhaven posts out on your blog, and then, uh, take the, the ones you think are the best fit for Lesswrong or strongest and, uh, not, not once a day, but, uh, put them over to Lesswrong.
Benjamin:
Exactly.
Ben:
That seems like a pretty reasonable way of doing an Inkhaven.
Benjamin:
Yeah, that’s exactly how I’ve been doing it.
Ben:
Gotcha. Yeah, I think I’m still kinda curious, uh, ‘cause a lot of MADS people post on Lesswrong, and- I think not as much. Uh, uh, my guess is something like less than half of it finds substantial traction. Do mass people care a lot about that? Do they try hard to make it do well on Lesswrong? Are they like, “Ah, that’s kind of...” Are they pretty happy when it goes well on Lesswrong? How much do they care about it?
Benjamin:
I think they do care quite a lot. Um, I was talking to Tim Hua, uh-
Ben:
Oh, yeah. He’s, he’s just... He’s written some great things on Lesswrong.
Benjamin:
Yeah. He was saying he just passed 1,000 karma, um, so he got the, like, double up vote power, and he was like, “Oh my God, this is so amazing.”
Ben:
I remember being about 17 or 18 and posting on Facebook that I’d crossed 1,000 karma on Lesswrong.
Benjamin:
Oh, wow. That’s impressive.
Ben:
Um, but, uh, uh, I hadn’t written anything as good as... Uh, he’s got some re- he’s got some curated pieces that’s got, uh, I think a few hundred karma for one of his pieces.
Benjamin:
Yeah. I think his post on the spiral personas-
Ben:
That one
Benjamin:
... um, got at least a few hundred up votes.
Ben:
Wasn’t that an Adèle Lopez post, or am I... They both got one of those.
Benjamin:
They both... Uh, so Tim’s one was-
Ben:
Was his about mental he- like, uh, how the different models would treat people when they were depressed and so on?
Benjamin:
Well, it’s basically, like, um, taking the spiral idea and trying to apply metrics to it-
Ben:
Metrics
Benjamin:
... to see how different models, like, in- do the spiral thing.
Ben:
Yep. The title is “AI-Induced Psychosis: A, uh, Shallow Investigation.”
Benjamin:
Yeah.
Ben:
That one was great.
Benjamin:
Yeah.
Ben:
Um-
Benjamin:
I think Neil kind of also encourages his scholars to, like, post early and often on Lesswrong, even if the posts are, like, a little bit rough. Um, I think the idea is that, you know, if the idea is interesting, it’ll get some traction, and like-
Ben:
He’s a fella who, uh, I think has done a long period of daily posting, uh, early on.
Benjamin:
Who, Neil?
Ben:
Yeah. I think when he was pretty young, like 20 or something.
Benjamin:
Yeah. I mean, I was, like, hugely intimidated early in Inkhaven, ‘cause I would go and read Neil’s posts, uh, and they were, like, really good.
Ben:
Yeah.
Benjamin:
You know, the ones he was posting every day. Uh, I was like, “How, how, how did he do this?” And he was, like, working at Anthropic at the time, um, while posting these, like, 2,000-word, like, daily pieces about, like, how he... Like, his algorithms for thinking about everything in life.
Ben:
Yeah.
Benjamin:
Um-
Ben:
I mean, I think it’s a combination of it’s a good fit for who he is psychologically to be able to do that kind of thing. But also, I’m pretty sure he did long periods of daily posting bef- much younger, before he was at Anthropic. You know?
Benjamin:
Oh, interesting.
Ben:
Yeah. I’m forgetting the name of his blog, but I think he did a bunch of, like, rationales, he explainers and things back when he was a bit younger.
Benjamin:
Yeah.
Ben:
‘Cause he was pretty young when he arrived on the scene. Um-
Benjamin:
Yeah
Ben:
... has, uh, daily writing... Is this the first time you’ve done some daily writing?
Benjamin:
Yeah.
Ben:
What’s this done to your mind? How has it affected you?
Benjamin:
Um, I mean, partially I feel like I didn’t take full advantage of the opportunity by, like, you know, spending, like, six hours a day writing. You know, often the post would actually only take, like, two of... two to four hours of actual writing. Um, and so yeah, it feels like it would be better to really, like, you know, fully immerse my brain in as much time doing it as possible.
Ben:
Mm-hmm.
Benjamin:
Um, like, it was actually really helpful working with, like, Alicorn and, like, sitting in a circle together and, like, writing a bunch of words in a limited period of time. I could have-
Ben:
Was this a Pomodoro or Speedhaven or some third thing?
Benjamin:
No, just, like, her s- it’s kind of a silly game.
Ben:
Oh, yeah, yeah. I recall that game.
Benjamin:
Yeah. And, uh, it was actually really effective. I, uh, I could’ve done that, like, the whole day, every day for, like, eight hours.
Ben:
Hmm.
Benjamin:
Um, if, if someone had been facilitating it, I would’ve probably joined.
Ben:
Interesting.
Benjamin:
Um, I actually wanted to do, like, a number of, you know, just how many words can I write, even if they’re just garbage.
Ben:
Yeah.
Benjamin:
You know, like, can I hit 8,000 words a day, uh, for, like, multiple days? Um, but I didn’t actually, like, try to do this. Um, but I would’ve liked to, um, you know, like, seriously do that.
Ben:
Did you spend less time writing because you were distracted by other things, or just because, uh, you know, it’s an uphill battle to write? Or were you... Was the socializing somehow distracting you, or did you have other work obligations?
Benjamin:
No. I mean, I think it’s largely, like, procrastination. Um, and, like, if... I, I also think, like, not sleeping enough. If I’m tired, uh, I just, like, don’t want to work, and I just want to, like, you know, play Factorio or something. Ideally, I would have it so that the deadline is, like, 10:00 PM. Um-
Ben:
Are you... I was gonna ask, were you one of the late, late writers, publishers?
Benjamin:
I started, like, posting later and later as time went on.
Ben:
Yeah.
Benjamin:
Um-
Ben:
You’re not Viv, who hit the record for the latest, which is three seconds and four seconds, both of those before midnight.
Benjamin:
Yeah, I was, I think, like, 43 seconds before midnight at one point.
Ben:
Yeah.
Benjamin:
Um, but that was really stressful.
Ben:
We nearly lost you.
Benjamin:
Yeah. It was bad. Yeah, I kind of had intentions of, like, using Beeminder to, like, um, you know, lock in. You know, it can essentially go and check my website for whether I’ve actually posted that day already, and then charge me money if I haven’t, and I think that would’ve just been smart. I’m sad I didn’t actually implement that.
Ben:
Mm-hmm.
Benjamin:
Um, there’s so many things I would’ve done differently if I could do it again.
Ben:
We were considering... Some people wanted to implement fake midnight at, like, 6:00 PM or something.
Benjamin:
Yeah.
Ben:
And hopefully get some incentive to do it by then each day-
Benjamin:
Yeah
Ben:
... so we could have our evenings. Um, maybe we’ll figure something out for that next time. I am ge- yeah, I’m generally curious. So is that... Is there any other advice you would’ve given yourself, uh, at the beginning of this to do it differently than you did?
Benjamin:
The most powerful version of myself would, like, wake up at 7:00 AM You know, have my post done by like lunchtime.
Ben:
Mm-hmm.
Benjamin:
Um, when my brain is like freshest and like most ready. Um, go to bed at like 9:00 p.m. Spend some time like hiking in the mountains like every second day. Um, and I think also having more like structured sessions around exclusively writing, like let’s just sit and write circles.
Ben:
Yeah.
Benjamin:
Um, I think that would be a really good thing to implement as like an official thing from Inkaven side. Like, um-
Ben:
Writing hours.
Benjamin:
Yeah. Like, you know, in this r- time, in this room, we’re going to sit and write.
Ben:
Yes.
Benjamin:
And, um, everything will go on like a public leaderboard, and you can see how much everyone is writing. Um, I think I would’ve like enjoyed joining that.
Ben:
That makes sense. Yeah, I could do... That seems like a good idea.
Benjamin:
Mm-hmm.
Ben:
Um, what are you proudest of in the writing you’ve done here at Inkaven?
Benjamin:
Um, I think I wrote some like pretty genuinely good little research pieces, you know, that are like 70 to 80% of a research paper in like a day. I think the post titled “The Garden,” which got like very few upvotes on Lesswrong.
Ben:
Mm-hmm.
Benjamin:
But it’s actually like quite a good piece. Uh, it does some like real philosophical work. Um, I wish it got more upvotes. Um, I think some of the like emotion posting that I did was like quite good. Um, you know, kind of working through some stuff.
Ben:
Mm-hmm. What’s, uh, what’s one piece of writing feedback or advice you got during Inkaven that was especially good?
Benjamin:
I think getting advice from Justice on some of my posts was very useful. Um-
Ben:
One of these podcasts I’m gonna have to explain who Justice is and why everyone keeps thanking him.
Benjamin:
Oh, yeah.
Ben:
‘Cause he’s, he’s our on tap. You can get him for Lesswrong as well, you know.
Benjamin:
Yeah.
Ben:
You go to the bottom of the editor and you hit request feedback.
Benjamin:
And that’s just free, right?
Ben:
I think for 100 karma or something, yeah.
Benjamin:
Oh, you pay karma?
Ben:
No, no, no. Once you hit 100 karma it’s free.
Benjamin:
Oh, okay. I was like, “What?” Okay. That’s cool.
Ben:
Yeah.
Benjamin:
Yeah.
Ben:
Yeah, we basically employ him part-time to do that.
Benjamin:
Yeah. No, he is genuinely very helpful.
Ben:
Yeah.
Benjamin:
I liked having Jesse, uh, single.
Ben:
Oh, sure. Did he... How... Did, did you take him a technical alignment piece of research?
Benjamin:
No, no. I just enjoyed the workshops he ran.
Ben:
Oh, yeah, yeah, yeah.
Benjamin:
Yeah. I didn’t actually get feedback from him.
Ben:
Oh. Yeah, he did a good one on, uh, Consider the Lobster.
Benjamin:
Yeah, exactly.
Ben:
Which I had not read before.
Benjamin:
Yeah.
Ben:
It was a very rationalist piece of writing. It was like, “Yep, I’ve got to talk about this festival. Let’s talk about philosophy, ethics, animals, and, uh, just go there for the rest of this essay.”
Benjamin:
Yeah.
Ben:
It was great.
Benjamin:
I think it would’ve been quite cool, like I didn’t do this, but I posted on Slack about, you know, for the, the fair, creating like a dashboard of like people’s favorite pieces of writing. Um, you know, just like having it all kind of ingested into a dashboard, you know. And then people can also upvote, uh, each other’s like submissions on the dashboard, so then you can like see which pieces are like most universally enjoyed or endorsed.
Ben:
Yeah. Popular is, is a tough ranking. I don’t like the popular ranking on the front of the Inkaven blog website. It’s a little bit too like cool. There’s like... I don’t know, very general interest stuff. But, uh, I would be interested in knowing people’s favorite pieces of writing. I should probably add, added that as like a... I could probably add that as like a, a bit of the portal for everyone’s bios or something.
Benjamin:
Yeah.
Ben:
Maybe I’ll check in the feedback form that I force everyone to fill out tomorrow.
Benjamin:
Yeah.
Ben:
Um, whose, uh, whose writing have you liked reading?
Benjamin:
Um, I th- I think Alec and Viv were my tough picks.
Ben:
I had them both in here. Yeah, they, they’re f- good fun.
Benjamin:
Yeah. And, um, yeah, like Lawrence has been like a real powerhouse.
Ben:
Yes.
Benjamin:
Just like very impressive.
Ben:
Was there anything especially surprising about the Inkaven experience for you?
Benjamin:
Seeing who like showed up and who was interested, and people’s like different... How wide the like range of people’s writing was. Um, yeah, just how much variety and like types of writing showed up, you know?
Ben:
Yeah.
Benjamin:
There was a, there was like a handful of technical AI safety people. Um, there were people who just wrote about like, you know, magic. People who wrote about like their deep personal experiences. And yeah, that was quite surprising seeing the extent-
Ben:
Yeah, I can’t tell whether I want- wanted more AI safety stuff, I think. I wanted to... I feel like the kind of ego goes around here wouldn’t, by default, would like to eat Inkaven and make it all AI safety things, and so I’ve pumped against that. I, I mean, you know, I, I mostly just tried to select people who basically submitted good writing. I wasn’t really trying to do anything else.
Benjamin:
Yeah.
Ben:
But it’s possible. I think I would’ve preferred slightly more, but, uh... Is there any memory or experience you’ll take away that has been especially Inkaven shaped, that you can only have if you lock 55 writers in a walls-of-ails compound for a month?
Benjamin:
The karaoke probably.
Ben:
That was pretty good.
Benjamin:
Yeah. I think it would’ve been quite nice to have a few more things like that. It feels like in the last week we’ve been like-
Ben:
Trying to make up for lost time.
Benjamin:
Yeah. Um, yeah, I’m not sure like why that happened. Um, it’s also kind of interesting, like after the showcase, it felt like everyone was kind of in a like, “Oh, this is over” mindset. Like the climax happened a week early.
Ben:
Yeah.
Benjamin:
Um-
Ben:
Yeah, I think, uh, I’m making updates about the dates of Inkaven. I think I’m no longer sold that it needs to be 30 days, and no longer sold that it needs to start on the first and end on the last day of the month.
Benjamin:
Mm.
Ben:
My guess is it would be better if it started on a like Saturday and ended on a Saturday or something like that.
Benjamin:
Yeah.
Ben:
Such that the last day could be a fair, uh, rather than five days before the end.
Benjamin:
Yeah.
Ben:
Um-
Benjamin:
I mean, maybe there’s like... It’s not necessarily a bad thing, you know. In some sense, having people come to terms with the end early might be a good thing ‘cause it galvanizes them.
Ben:
Yeah.
Benjamin:
Yeah. It’s hard to, like say.
Ben:
Mm-hmm, mm-hmm.
Benjamin:
Yeah.
Ben:
Um, great. Uh, well, um, thanks for, uh, sitting down with me, and thanks for your writing.
Benjamin:
Mm-hmm.
Ben:
Uh, and, uh, good luck with the rest of it.
Benjamin:
Yeah. Thank you for running it. It’s been really fun.
Ben:
You’re very welcome.
Benjamin:
Yeah.
Ben:
All right. Farewell for now.









