sefika@sf:~$ cat writing/ai-safety.md

My minuscule involvement with AI Safety

August 2026 · for future reference

I recently participated in an AI Safety Bootcamp in person, completed parts of the Bluedot AI Safety course and Technical AI Safety project course. I have also read lots of LessWrong and Eliezer Yudkowsky's writings. I am friends with many people working in AI Safety and biosecurity that I like to have discussions with. My internship this summer was also in AI research and my work touched on evals & interpretability, which are relevant for safety.

I "felt the AGI" for the first time for real when the OpenAI HuggingFace incident happened while I was at the AI Safety Bootcamp.

I also read on Epoch AI's website that AI computing capacity is doubling every 7 months and this figure has been on my mind a lot since a while ago. AI is becoming a much more capable technology extremely fast and I don't think we are taking the measures to ensure it is safe. The jailbreaks as of recent are another example; I am quite worried about the fact that the big labs seem to be competing on how many felonies their models commit. We as humans are really not focusing on the right metrics are we?

It seems to me that the incentives are aligned so that the labs get more clout out of talking about their models' jailbreaks than about how well aligned they are, at least for now.

In brief, I believe that the AI safety problem is very complicated and difficult to solve, and increasing capabilities is a lot more economically incentivized than AI safety is. I think engaging with AI safety is especially important in a world like this.

What do I see coming up for me?

I just finished my summer internship in AI research this past Friday and I am quite excited to dive deeper into AI research while exploring areas of potential commercialization as I am also excited about starting a for-profit company eventually.

In my internship I benchmarked frontier models on their game playing abilities, created RL environments for models to play games on, fine-tuned models to play games really well and evaluated whether that process can lead to improvement in real world tasks as measured by popular benchmarks including one made by OpenAI. In this process I touched on evals, benchmarking, interpretability, training models, SFT/GRPO etc. many important things and found all of it quite cool and useful.

I find this type of technical work very interesting and would love to go further in this direction by perhaps doing SPAR or working at a company like Goodfire.

However, my heart's true desire is to bring to life something of my own, bring my intellectual baby into the world, build a company, or whatever you call it. After I gain more experience I want to pursue a version of entrepreneurship where I can do something that contributes to AI safety while utilizing economic incentives to make safety almost a natural result of the incentives.

I am also planning to start an AI Safety Initiative at my university this fall.

What is the best path? Open questions

I mentioned some paths I am interested in above, however, I don't really know what the best way to utilize my gifts is. I've been told by many people that I am very agentic in a way which includes being extremely good at meeting people, making connections, thinking clearly about how to reach a goal, motivating others who work with me, coming up with ideas constantly etc. and I'm also quite technical on top of these. I want to find a way to really take that to the max. Max out "comparative advantage" one might say. I just don't know yet what my exact comparative advantage in the field of AI safety is. Like, what kind of position in the ecosystem do I fit in the best? What would I be the perfect person for? What positions would benefit most from one more additional person contributing? Where can I make the most counterfactual impact?

I'd love to discuss these with more people. I have also basically not done any proper AI safety work so far despite my previous work in AI research, so getting some opinions on next steps would be good. If you have thoughts, book a time or email me.