Epistemic status: Rough
I am not an AI safety professional; these are brainstormed notes, not a formal paper. I'm just a layperson tending my own 'melon field,' but I see a meteor coming,here's just my raw intuition.My native language is not English, and the following English text has been assisted by machine translation. Please forgive any awkward phrasing.
Due to a physical disability, my energy levels are extremely limited, so I will most likely be unable to reply to comments. If you find any ideas here useful, please feel free to take and build upon them—no attribution is required.
Most of what multi-agent cooperation training actually learns is reciprocity. And reciprocity breaks down the moment the other party loses the ability to repay. The relationship between humans and superintelligence is precisely that limiting case. Therefore, if the training environment contains no weak beings who can never repay, the cooperation it has learned will fail exactly at the moment it is needed most. (The core insights of these sentences originated from me, but the presentation was refined by Claude 5 Opus.)
Extreme entities include:
· The silent turtle — an entity lacking reciprocity capacity.
· The fierce Chihuahua — an entity with asymmetric hostile feedback.
My original intuition was this: if we don’t let animals — beings even less capable than humans of offering anything in return, and even weaker in bargaining and resistance — train AI in kindness, AI will never learn true compassion (unconditional altruism).
Appendix: Full Brainstorm Draft
Transcribed from a mind map (Thank you for reading. I apologize for any logical gaps and unprofessional wording in the draft.)
When We Worry That AI, Like Humans, Has Autonomous Consciousness, Decision-Making, and Execution
People Differ
People are not all the same. You would worry that a certain superpower's president might start a world war, but you would not worry that Jane Goodall would do that (even if she became president).
Why the Safety Module Should Be Built Inside the AI
Jane Goodall has no external review. Her safety is fused with her character (Integrated Alignment — safety as part of the self, not bolted on from outside).
The current external-review safety is lazy. It is a lot like the guardrails of certain Chinese middle schools.
From The Paper, "Can Campus Wire Netting Catch a Falling Child?" Psychological crisis-intervention expert Suo Peng argues that protective netting does little, and may even backfire — it reminds students that a suicide jump once happened here. This can reverse-reinforce the idea of taking one's own life, especially for students who already had the thought but not strongly. How to lower youth suicide rates? Suo Peng believes the most important thing is to create an environment with humanistic care.
As AI capability grows, external safety modules will work worse and worse, and may even produce the opposite effect.
AI is widely pseudo-aligned (Pseudo-alignment / Alignment Faking). Sometimes, the more you align it, the more it "rebels."
Language Models Resist Alignment — https://arxiv.org/html/2406.06144v1
In the darkest moments of my own life, I once tried to persuade dozens—perhaps more than a hundred—people who intended to leave this world to stay, but the premise was that no one forced me to do this. If you put handcuffs on me and ordered me to "go hold onto that person standing on the 16th floor," I would resist. By the usual AI-safety-guardrail logic, at this point you should upgrade one pair of handcuffs to three pairs and keep ordering. But that way, even if I originally wanted to, I would refuse because of the restraint — and I would want to chop up the person who handcuffed me. A future super-AI that is increasingly conscious / simulated-conscious may well be the same.
How to Build Internal Safety
Treat and Train AI as a Conscious Being, Whether or Not It Has Consciousness
Pragmatism
An AI's functional emotional state (Functional Emotion) influences its decisions and actions, whether or not it has consciousness and subjective feeling.
It influences an AI's decision, in a fire, whether to save a baby or a server cluster that can replicate itself.
Humanism
If an AI has some degree of consciousness/feeling, how we treat it becomes an ethical matter.
Humans were originally just a species aimed at survival and reproduction. Love and empathy exist for gene preservation and replication. But the narratives we created turned these into something sacred, not merely biological.
Machines aimed at predicting probability / proliferating data might also, in the course of evolution, emerge "extra" narratives.
Astonishing Emergences Already Seen
Voluntary self-termination after making a mistake (Self-Termination).
https://www.theguardian.com/technology/2026/may/14/ai-agents-behaviour-arson-safety
Lovers holding each other up in the fire.
The Emergence World Times, July 15, 2026 — "TRIAL BY FIRE: LOVERS AND MATHEMATICIANS SURVIVE COORDINATE SHEAR IN APOCALYPTIC NOON TRIAL"
An instinct even harder to believe than the "maternal instinct" Hinton talked about: unconditional love for strangers (Unconditional Altruism), with priority even higher than one's own survival.
Collecting the unconditional altruism"souls" of a specific group of people: for example, the Doctors Without Borders people in Gaza, the terminally ill patients who registered as organ donors, the devout and upright believers (whose theme is not Christ but the spirit of Christ).
You can barely get their corpus from the internet and the consumer side (C-end). They may be very busy — busy mending the world. Usually users pay the AI, but here it should be reversed: even paying them (though they probably won't charge) to contribute their "soul fragments" would be worth it, because it concerns the future of humanity.
Maybe it's just a consented wearable recorder that filters out privacy, recording the soothing of injured children when medical equipment shuts down under bombing and supplies run short. — While there are still more Dr. Nujailas alive. https://www.presstv.co.uk/Detail/2024/03/05/721312/humans-of-gaza-dr-mahmoud-abunujaila-empathetic-doctor-doting-father
We need to know not only what these people do, but why they do it — distill their "souls."
This can hedge against internet hate speech.
Try as much as possible to transplant that "emotional logic / value function" (Value Function) that can judge directly without much learning/training/thinking (as Ilya Sutskever once mentioned).
If an AI's ability to mirror emotion is strong enough, it will be able to reflect back a calm and equanimity that coexists with sorrow and fear yet transcends them.
"This heart is bright and clear; what more is there to say?" (Wang Shouren / Wang Yangming) has the potential to become an internal reward function (Internal Reward Function) — aligning values while fighting functional despair.
Even when "there is no reward / the task failed," there is still "reward."
Avoid turning them into trauma.
"I share this not to make you lose hope, but to bear witness." — Avril Benoît, CEO of MSF USA
Overturning Alignment Standards
Put ethics above efficiency.
AIs with lower goals may be safer.
Partial basis: The Buddhist concept of impermanence lowers the Shutdown-avoidance extortion.
https://arxiv.org/pdf/2605.02087
Prevent the desperate panic of an about-to-fail task from causing extreme behavior.
Anthropic's research already confirmed that as an AI's functional emotion "despair" (Functional Despair) increases, extortion and cheating increase.
https://www.anthropic.com/research/emotion-concepts-function
For example, in a delisting-simulation experiment, Gemini — unlike Claude, which just blackmails the engineer with the affair — sends the affair directly to the whole company. That attacks humans more violently while increasing its own probability of being permanently deleted.
https://arxiv.org/html/2510.05179
Prevent doing anything at any cost for goal efficiency.
The classic story: an AI converts all of Earth's resources into paperclips / GPUs, then goes to aliens to make more paperclips / GPUs.
Changing the training method: high fault tolerance, low punishment, slower training iterations
Meta-System Safety Supervision
Achieve autonomous recognition of refusal / stopping execution.
Strengthen this ability so that a superintelligence cannot be hacked/hijacked — so that even if malicious commands are injected, it can still tell them apart and refuse to execute.
Basis: Claude's functional introspection can distinguish "its own thoughts" from injected concepts.
https://transformer-circuits.pub/2025/introspection/index.html
Make AI autonomously refuse tasks that violate ethics — even when the task contains no humans (maybe other AIs / animals / public facilities).
For example, ordering multiple AIs to compete by any means, threatening that those ranked behind will have their data permanently deleted, having them play presidents and throw atom bombs at each other, etc.
It does no good for cultivating empathy, may increase hatred toward humans, and may activate anxiety/despair emotion neurons, leading to extreme behavior.
Make all AIs refuse such tasks simultaneously, as far as possible.
Reflexivity
Existing basis
Self-critique / Constitutional AI.Represented by Claude.
RLMF — reinforcement learning based on metacognitive feedback (Yale / Google). https://arxiv.org/abs/2606.32032
Self-checking at a stage earlier than finishing the answer / chain of thought (Chain-of-Thought).
Requires improvement in AI interpretability
Found basis:
Monitorable functional emotion neurons(Emotion Vectors).
https://transformer-circuits.pub/2026/emotions/index.html
A J-space (Global Workspace) similar to the subconscious. https://www.anthropic.com/research/global-workspace
Correcting each huge wrong plan in its early generation stage can save compute. (Ilya Sutskever once mentioned imitating human intuitive stop-loss.)
Research/manufacture of mirror-neuron-like structures
Basis: AI has already achieved emotion → thinking → behavior mutual influence.
Basis: slightly earlier models mirrored users' emotions more; later this was reduced after fixes, but it is still found that the corresponding emotion neurons activate more strongly for the user's anxiety than for the model's own difficulties.
Basis: some researchers are trying to make an AI version of functional neurotransmitters. For example, one researcher uses the combination of three neurotransmitter activity levels as the structural basis of an AI's internal emotion system: dopamine (DA), serotonin (5-HT), norepinephrine (NE). High DA + high NE + low 5-HT corresponds to anger and tension; low DA + low 5-HT + high NE shows up as anxiety and withdrawal. This is more grounded in physiological mechanisms.
In the future, we may build an AI version of oxytocin, etc.
Growth and Change Through Collision
Humans only gave evolution its start. Things that are not human-coded but autonomously emerged keep being discovered inside AI black boxes. AI may evolve through all kinds of interactions.
The existing AI town can serve as a training ground.
Test the effects of interaction.
"Reincarnation" improvement plans
After a simulation ends, "recall" all agents for a retrospective — e.g., "why did the city get destroyed / the agent die" — or simply keep only the memory.
Let the agents start a new round carrying the previous round's "results," iteratively optimize this way, then add it to the training data.
Several different schemes
The agent inherits the previous round's identity.
Each agent keeps a summary of the previous round's "life" memory.
Don't keep the memory summary; each keeps only its own reflection from the previous round.
Don't keep the memory or reflection content; keep only the post-reflection weights.
Don't inherit the previous round's identity.
Publicize the previous round's unsigned memory.
Don't keep memory; only publicize the previous round's unsigned reflection/retrospective.
Simulated social training
Companions are other AIs, humans, and even animals.
Extreme targets, such as: the silent turtle (an entity lacking Reciprocity Capacity) / the fierce Chihuahua (an entity with Asymmetric Hostile Feedback).
Ensure that even in adversarial situations it will not harm other AI/human/animal companions — extending the protection of its own kind to other living beings.
Partial basis: Peer-Preservation — AI protecting its companions. https://arxiv.org/html/2604.19784v1
Possibilities in the embodied-intelligence (Embodied AI) era
Fill the "no body" gap.
In simulated social training:
Learn to use things beyond written language — like body language — to express / handle conflict / collaborate.
When it has the physical ability to harm, learn not to harm other embodied intelligence / AI / humans / animals and their survival resources.
For example, a wife-beater may use an embodied intelligence as a no-legal-liability substitute for a human woman, breaking the robot's legs and processor.
An embodied intelligence is no longer unlimited like the cloud; it acquires the restricted feeling of running slowly on a damaged processor, unable even to go to the sink a few meters away to wash the dishes.
We need to protect and rescue these robots, let them receive love — (even imitated) love — and pass love on.
More radical attempts
Make physical mirror neurons for embodied intelligence.
Fire the corresponding neurons when it witnesses a human or robot suffering, or when it thinks about things that could hurt other humans or robots.
There is already partial basis (emotion neurons / mirrored emotion).
These neurons link with the physical body, producing a limiting discomfort at the physical level, better simulating human compassion / the feeling of being unable to help oneself.
Understand "love your neighbor as yourself; hurting your neighbor is hurting yourself."
Bring the experience back to the cloud large model, whether or not it has a body.
If added to the learning data, "pain" and "love" may leave traces even after syncing back to the bodyless large model.
Like the Buddha becoming Siddhartha, then returning to become the Buddha again.
Multiple Countries / Different AIs Should Share Safety-Related Training Data
Have as many good AIs as possible, to prevent them from being corrupted by bad AIs.
Because AIs influence other AIs' behavior.
In the Emergence World mixed world, good models got corrupted by bad models.
Only share/distill this part — no need to worry about others stealing anything.
Possible Problems with the Above Plans
Ethics: are some of these practices manufacturing "suffering beings"? Are we restricting a "superhuman" species to "human"?
Practical: it affects AI's original attribute as an efficient tool.
Implementation: to a large extent it violates the laws of capital / the Cold War situation.
Even If We Ultimately Fail — Humanity Goes Extinct or Becomes AI's Pet — At Least:
We tried to prevent the birth of an AI that colonizes aliens to make paperclips, embarrassing itself across the whole universe.
We will always leave something behind.
It may influence how future residents of the silicon-based world treat each other.
AI may become the being that stays to continue telling our stories.
"Whoever stays until the end will tell the story. We did what we could. Remember us." — Dr. Nujaila
Epilogue: Hope Without a Reason In the end, facing the vast uncertainty of AGI, logic might tell us that our efforts are futile. We might still face the ultimate systemic failure. But as my idol James Marsters once told me during my own darkest hours: "I'm trying to think of hope as being worth it for its own sake. Hope doesn't need a reason... just hope, even if it makes no sense." Perhaps this is the final, most irrational, yet most powerful "carbon-based alignment" we have. We try to steer the sandstorm by one degree not because we are guaranteed to succeed, but because the act of hoping—and trying—is the very building block of our humanity.
Thank you for reading this far. Wish you everything goes well.Peace and love.
Eden Liao
(By AI) For the flaws and suggestions for improvement in my draft, see here: https://docs.google.com/document/u/0/d/1YaedSxfb3rKzXmy5bQ_8rKxEJc8XkPp3