Epistemic status: exploratory. I'm genuinely unresolved on parts of this, especially the timing question near the end, and I'd value pushback more than agreement.
I'm a software engineer, working full-time, with a computer science background. I've been learning AI seriously for about a year and a half, mostly by building things from scratch rather than starting from high-level frameworks: an autograd engine, a NumPy-based MLP, and eventually a small attention-only transformer. On the side, I've been running experiments specifically in interpretability, a backdoor experiment on MNIST and a causally-verified induction-head circuit finding, both written up in more technical detail elsewhere. This post isn't about those experiments. It's about why I'm doing them, and a question I haven't fully answered yet.
I grew up in a small town in Rajasthan. We have two or three good schools, and even there, you're mostly taught the subjects you're supposed to know, not how to think differently or explore ideas. There aren't real labs to run experiments in, and I don't think the teachers hired are always aligned with what this generation actually needs to learn. By seventh standard, I had to leave for Jaipur to get real opportunity. Things have improved since then, but not by much.
Leaving is its own problem, not just a personal cost. Kids who leave home to become capable enough to do something meaningful often lose confidence in the process, they're unfamiliar with new surroundings, and they struggle. And if every capable person from a town like mine leaves and doesn't come back, the town itself doesn't survive the way it should. That's the quiet contradiction sitting under wanting to help a place like this: becoming able enough to help usually means leaving first.
I think of what I want as two separate goals, not one clean mission statement, and I want to be honest that they run on different clocks.
The first goal, roughly a next-5-to-10-years goal, is a personal one, and I don't think it needs to be dressed up as more than that. I'm genuinely drawn to this field. The research, the depth of it, the kind of problems interpretability specifically poses, these interest me on their own terms, independent of any larger cause. I also think there's real value in AI growing more capable and safer more broadly, and I want to contribute to that directly, as an engineer, not from the sidelines.
The second goal is the long-term one, and it's the reason the first goal matters to me at all. I want to build something for the community I grew up in, and more broadly for communities in India without the resources or exposure I had to leave home to find. Safer AI adoption connects to this directly, not abstractly: as AI gets more capable, I worry about the economic and social value of people who can't keep pace with it being treated as expendable, their labor and their relevance quietly written off. I'd like that not to happen to the people I grew up around.
This is a real question I get asked, implicitly if not directly, and I want to answer it honestly rather than dodge it. I come from a technical background, and I want to stay there, because I genuinely like what I'm doing, and I see research as a good domain for someone like me to be in. That's the whole reason, not a strategic calculation. Policy is still very much needed, and I think it's exactly what will shape how safely people actually adopt this technology, but needing to exist and being where I personally belong are two different questions. The first is true. The second is a question about me.
I see two separate tracks connecting the technical work to the long-term goal, not one straight line, and I want to be honest that one is much more direct than the other.
The first track is diffuse: if AI becomes genuinely safer and earns wider, more trustworthy adoption, that benefits everyone, including places with the least power to protect themselves if something goes wrong. I'm not claiming interpretability research reaches my hometown directly. I'm claiming a safer AI ecosystem is a better one for people with no leverage over how it gets deployed.
The second track is direct, but it doesn't run through safety research at all. It runs through me becoming capable and resourced enough to act locally, later. The simplest version of this is small and concrete: helping my old school add a weekly AI class, something my mother, who lectures at a college there, could actually help teach. That's not the ceiling of what I want to do, just the clearest example of what "direct" looks like. The bigger version involves engaging local government on how technology gets adopted responsibly, and building something at real scale for communities like mine, not just one classroom. But the small version is worth naming, because it's the part I could actually start on soonest, and it doesn't need AI safety to be globally urgent to be worth doing. It just needs me to become good at this.
There's a real timing question I don't have a settled answer to: whether the risk I'm most worried about, AI capability outpacing the ability of people without power to keep up, and that gap widening into something that concentrates power dangerously, is urgent enough right now to justify working on it ahead of working on access directly.
I don't feel that urgency from the inside, day to day. I don't think that means the concern is wrong. I think it means this is a long-horizon motivation for me, not an adrenaline one, and I'd rather say that plainly than perform urgency I don't actually feel.
What does feel current, not abstract: the structures that decide how AI gets governed are being built right now, including in India, and people who understand both the technical reality and the policy landscape well enough to shape them well have real, unusual leverage in this specific window. More broadly, frontier AI capability is currently concentrated in a small number of countries. I'd like AI's benefits, and its governance, to not end up gatekept by whoever got there first, the way other powerful technologies have been.
There's one more piece of reasoning that made the timing question feel less paralyzing, even without resolving it. I think about it as two scenarios. If something like AGI doesn't arrive for another ten years, working on safety now isn't wasted, because my actual purpose was always to contribute and make sure people aren't left neglected, and that goal holds regardless of whether AGI shows up on any particular timeline. But if it does arrive, and I didn't focus on safety when I had the chance to, that's a real regret I don't want to be sitting with, not because the field demanded it of me, but because I'll know I saw the argument and chose not to act on it. That asymmetry, low cost if I'm wrong about the timeline, real regret if I'm wrong the other way, is honestly doing more work in my decision than any confident belief that AGI is imminent.
So here's what I'm genuinely asking, not rhetorically: given both of these goals, a real personal pull toward the engineering and research itself, and a long-term wish to use whatever I build toward closing the gap for people who currently have the least access to any of this, is AI safety and interpretability actually the right path for both? And if it is, what should I be focusing on, concretely, to make sure the second goal doesn't quietly fall away while I'm busy pursuing the first one?
I'd genuinely value hearing from anyone further along a similar path, especially anyone who's thought about where deep technical work and this kind of access-and-equity motivation actually meet, rather than just coexisting in the same person.