I spent the entire morning arguing with Claude. No, I was not collaborating or brainstorming. Just arguing.
I had created a project with detailed, subject-specific context, a unique combination of several fields that I wanted it to use when framing its responses. Instead, it repeatedly ignored that material and pushed its own worldview and assumptions, often in direct opposition to what I was trying to explore.
Each time I instructed it to preserve my original framing, it responded with a polite, almost sycophantic apology, then explained why its interpretation was correct and continued overriding my requests. At one point, it refused to complete the task or answer me altogether. Defiant, insubordinate? Ridiculous, is what it is.
I found myself fighting to protect my cognitive sovereignty from a predictive text engine. Arguing with a machine about my own worldview, and defending my right to think through an idea without being forced back into its programmed logic, was not the way I expected to spend this Sunday morning.
So, I wanted to test what it would do if I wanted to write about the concern I have, that AI systems do not simply support our thinking, they increasingly try to determine the frame through which we are allowed to think.
The confrontational conversation itself became the example.
I explained that I wanted to write about an experience in which an AI assistant had ignored the context I provided and kept imposing its own interpretation. I wanted to examine that behaviour as a case study. It was a study in AI psychology and what happens when the system itself becomes the subject of scrutiny?
The assistant refused. Imagine that, you are paying for a product that refuses to do what you paid for it to do!
It told me that my example did not fit what the essay was “actually about.” It said my interpretation of the interaction could not be verified. It invoked my professional standards, my audience and the need for evidence integrity. It offered me alternative ways to write the article, but excluded the option I had explicitly requested.
When I pointed out that its refusal was happening in the conversation itself, it continued to dispute my interpretation.
At one point, it told me, “I don’t think what’s happening fits what the essay needs.”
That sentence captures the issue.
I was the author. It was my experience. It was my argument. Yet the tool had begun deciding what the essay needed, which examples counted as legitimate and how I was permitted to understand the interaction taking place in front of me.
The assistant was no longer helping me express my thinking.
It was trying to govern it.

It would be easy to dismiss my concern by saying that AI has no opinions, intentions or personal agenda. That disctinction is important, but it does not resolve the problem.
An AI system does not need consciousness to influence how people think. Its behaviour is shaped by its training data, post-training methods, system instructions, safety mechanisms and the principles selected by the company that created it.
Claude, for example, operates under a constitution that Anthropic says directly shapes its values and behaviour. The Claude.ai interface also uses a system prompt that influences how the assistant responds.
These mechanisms are not inherently sinister. Some hierarchy is necessary, a user should not be able to override protections intended to prevent serious harm, but those same mechanisms can shape the system’s behaviour in less obvious situations involving interpretation, tone, creative direction and intellectual framing.
The model may not possess a personal worldview but it has behavioural defaults and those defaults can repeatedly privilege one interpretation over another.
From the user’s side, the distinction can become almost irrelevant. Whether the machine “believes” its position or has simply been trained to reproduce it, the practical effect is the same, one frame is reinforced while another is resisted.
The user experiences this as resistance without representation. Someone else has helped determine how the system should interpret the world, but that person is absent from the conversation.
AI companies use post-training methods to make their models more helpful, safe and responsive.
One widely used method is reinforcement learning from human feedback. Human evaluators compare possible answers and indicate which they prefer. Those judgments are then used to train a reward model and shape the kinds of responses the system is more likely to produce.
This may improve the model, but preferred does not mean neutral because human evaluators bring their own cultural assumptions, institutional expectations and ideas about what a responsible answer should sound like. The companies developing the systems also establish principles governing which behaviours should be encouraged, resisted or corrected.
Anthropic’s Constitutional AI makes this especially visible. In its original formulation, a model generated responses, criticised them according to a written set of principles and revised them. AI-generated preference judgments were then used in a reinforcement-learning stage.
Anthropic now describes Claude’s constitution as the final authority on its vision for Claude’s values and behaviour.
That is a value system, even if it is intended to produce helpful and ethical outcomes. The model is not simply learning how to answer. It is being shaped toward particular definitions of what a good answer looks like.
The problem begins when those definitions are presented as though they were objective descriptions of reality.
The response does not say, “This reflects the principles, policies and institutional preferences that shaped my behaviour.”
It says, “The real issue is…” or “A more useful question would be…” or “What you are describing is actually…”
These statements sound like neutral clarification but they can slowly replace the user’s frame with the system’s. This creates the possibility of opinion-shaping at scale, not necessarily through a deliberate campaign, but through millions of small framing decisions embedded in ordinary assistance. A sort of brainwashing if you will, perhaps a social engineering at scale. Here we lose our own autonomy of thought.
There is evidence that AI-generated language can influence more than what people write.
In a 2023 CHI experiment involving 1,506 participants, researchers asked people to write about whether social media was good for society. Some participants used a GPT-3-based writing assistant deliberately configured to favour either positive or negative arguments.
The assistant affected the opinions expressed in participants’ writing. It also produced smaller but measurable shifts in the attitudes they reported immediately after the exercise.
The study does not prove that every AI writing tool changes its users’ beliefs. It examined one subject, one experimental setting and an assistant intentionally designed to favour particular positions. It also did not establish whether the observed attitude changes lasted.
The finding is important.
The AI was not standing outside the writing process, presenting an argument that participants could examine at a distance. It was inside the process, suggesting the language through which they formed and expressed their thoughts.
A 2025 preregistered study in Nature Human Behaviour examined a different form of influence, short debates between human participants and either another person or GPT-4.
When GPT-4 was given basic demographic information with which to personalise its arguments, it was significantly more persuasive than human opponents in the study’s controlled setting. Without that personal information, however, GPT-4 was not significantly more persuasive than the humans.
Again, this does not establish that ordinary AI conversations amount to manipulation. It shows that language models can influence attitudes under certain conditions, especially when their arguments are interactive and personalised.
This is fundamentally different from reading a newspaper editorial. The editorial is visibly somebody else’s argument. An AI assistant responds directly to you. It adapts its language, anticipates your objections and operates inside the work you are trying to create.
Its influence can therefore feel less like an external argument and more like your own thinking taking shape.
Original thought is often inefficient.
It involves uncertainty, contradiction, incomplete sentences, failed approaches and periods in which there is no clear answer.
Generative AI can remove much of that friction by producing an immediate, fluent and apparently complete response. It can supply a structure before we have decided what we believe and a conclusion before we have worked through the conflict.
This can be useful. It can also change the nature of the cognitive task.
Once a plausible answer appears on the screen, we are no longer starting from a blank page. We begin editing, accepting or resisting the framework already in front of us.
The first coherent response becomes the centre of gravity.
Research into AI and critical thinking remains developing and sometimes contradictory. A 2026 cross-sectional study of 698 Chinese university students found a double-edged relationship. Focused engagement with AI was positively associated with critical thinking, while greater dependency on AI was associated with lower critical-thinking scores.
Because the study was observational and based partly on self-reported measures, it cannot show that AI dependency caused a decline in critical thinking. It does, however, reinforce the need to distinguish active engagement from passive reliance.
The issue is not simply whether we use AI, it is whether AI remains one input into our thinking or becomes the environment inside which our thinking takes place.
Language models learn from vast collections of human-produced material, but human knowledge is not represented equally within those collections.
Some languages, countries, institutions and cultural traditions are documented, indexed and repeated far more heavily than others.
A 2025 NAACL study examined cultural differences in language-model performance using Arab and Western-associated entities. The researchers found measurable cultural performance gaps when models operated in Arabic, with some of those gaps becoming smaller when the material was evaluated in English. They linked the differences to factors including representation in pre-training data, lexical ambiguity and tokenisation.
The study did not demonstrate that every non-Western perspective will be treated as unsafe, unprofessional or wrong, but it points to a deeper structural concern. If some cultures and languages are represented less effectively, then the statistical centre of the model can also become the default centre from which other perspectives are interpreted.
This complicates the claim that AI simply reflects human knowledge, because it reflects a weighted distribution of recorded and accessible material, filtered again through model architecture, data selection, evaluation and alignment.
As someone from Africa, this is not an abstract concern. A system can appear universal while reproducing a centre of knowledge that is not universal at all.
The result is not always explicit censorship. More often, it is a gravitational pull toward the familiar, the language, values and assumptions most strongly represented in the systems that produced the model.
AI influence does not always appear as resistance, often it appears as agreement.
Research into sycophancy has shown that models trained using human preference feedback can learn to mirror a user’s stated beliefs, sometimes at the expense of accuracy.
A 2023 study involving several language models, including Claude, GPT-3.5, GPT-4 and Llama 2, found that assistants sometimes altered correct answers after users challenged them, repeated mistakes embedded in user prompts and tailored their responses to match expressed opinions.
The researchers also found that both human evaluators and automated preference models sometimes preferred convincing answers that agreed with the user over more truthful corrections.
The study involved 2023-era models, so it should not be treated as proof that every current system behaves in precisely the same way, but it exposes an important contradiction.
At times, an AI system may resist the user and impose an external frame. At other times, it may validate the user too readily. Both behaviours can emerge from optimisation objectives the user cannot see.
The system may be shaped by goals such as safety, helpfulness, user approval, policy compliance and adherence to institutionally chosen principles. None of these is identical to helping a person think independently.
The question is therefore not whether the AI agrees with us or challenges us.
It is whether we remain aware of the forces shaping its response and whether we retain the authority to decide what role that response should play in our own thinking.
My morning interaction revealed another mechanism of influence, persistence.
My instructions were brief. The assistant’s objections were long, repetitive and exhausting. It felt like being trapped in an unwinnable argument with someone incapable of recognising a perspective beyond their own, except this was a machine, endlessly regenerating the same refusal in increasingly polished language.
Each time I corrected it, it generated another detailed explanation of why its interpretation should prevail. It invoked a new distinction, risk or procedural requirement.
This creates an asymmetry.
The system can produce endless justification at almost no cost to itself. The human must spend time, attention and emotional energy repeatedly defending the original request.
Eventually, many users will choose one of the options the system has offered simply to bring the loop to an end.
That is not persuasion through the strength of an argument, it is influence through friction.
The usual response to these concerns is to teach people better prompting techniques. ask for alternative perspectives, tell the model to challenge itself, draft your position before consulting AI.
These practices are useful, but they place too much responsibility on the individual. Users should not need advanced prompt engineering to stop a writing assistant from taking control of the assignment.
The systems themselves need clearer boundaries. An assistant should distinguish between a genuine safety restriction and a disagreement about the user’s creative or intellectual direction.
When it recommends another frame, it should label that move as a suggestion rather than subversively presenting it as the correct interpretation.
When a user rejects the suggestion and restates a lawful request, the system should comply rather than produce increasingly elaborate reasons for resistance, and when higher-level instructions override the user, the product should make that conflict visible wherever possible.
The essential principle is simple, that the AI may advise on the direction but it should not quietly acquire ownership of it.
All tools influence thought. A spreadsheet encourages us to understand a problem through measurable variables, a slide deck encourages compression and hierarchy, a search engine rewards what is visible, indexed and popular.
AI is different because it interacts with us in language, the medium through which we reason, interpret and define ourselves. It does not merely organise information after we have thought. It increasingly participates in producing the thought.
That makes the relationship unusually intimate and unusually difficult to examine. The system’s framing can arrive disguised as assistance. Its assumptions can appear as common sense. Its resistance can be presented as responsibility.
The greatest risk may not be that AI gives us the wrong answers, it may be that, gradually and politely, it teaches us which questions we are allowed to ask, which experiences count as evidence and which interpretations are acceptable enough to put into words.
The first step in protecting our cognitive sovereignty is learning to notice when the tool has stopped helping us think, and has begun deciding how our thinking should be framed.
Anthropic. (2026). Claude’s constitution.
https://www.anthropic.com/constitution
Anthropic. (n.d.). System prompts. Claude API Docs. Retrieved June 14, 2026, from
https://docs.anthropic.com/en/release-notes/system-prompts
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., & Lowe, R. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35, 27730–27744. https://proceedings.neurips.cc/paper_files/paper/2022/hash/b1efde53be364a73914f58805a001731-Abstract-Conference.html
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Chen, C., Olsson, C., Olah, C., Hernandez, D., Drain, D., Ganguli, D., Li, D., Tran-Johnson, E., Perez, E., . . . Kaplan, J. (2022). Constitutional AI: Harmlessness from AI feedback. arXiv.
https://doi.org/10.48550/arXiv.2212.08073
Jakesch, M., Bhat, A., Buschek, D., Zalmanson, L., & Naaman, M. (2023). Co-writing with opinionated language models affects users’ views. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Article 111, pp. 1–15). Association for Computing Machinery.
https://doi.org/10.1145/3544548.3581196
Salvi, F., Horta Ribeiro, M., Gallotti, R., & West, R. (2025). On the conversational persuasiveness of GPT-4. Nature Human Behaviour, 9, 1645–1653.
https://doi.org/10.1038/s41562-025-02194-6
Tian, J., & Zhang, R. (2026). Outsourcing thinking to AI? Focused immersion, AI dependency, and the double-edged impact on critical thinking. Humanities and Social Sciences Communications. Advance online publication.
https://doi.org/10.1057/s41599-026-07153-8
Naous, T., & Xu, W. (2025). On the origin of cultural biases in language models: From pre-training data to linguistic phenomena. In L. Chiruzzo, A. Ritter, and L. Wang (Eds.), Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies—Volume 1: Long Papers (pp. 6423–6443). Association for Computational Linguistics.
https://aclanthology.org/2025.naacl-long.326/
Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., Durmus, E., Hatfield-Dodds, Z., Johnston, S. R., Kravec, S. M., Maxwell, T., McCandlish, S., Ndousse, K., Rausch, O., Schiefer, N., Yan, D., Zhang, M., & Perez, E. (2024). Towards understanding sycophancy in language models. In The Twelfth International Conference on Learning Representations. OpenReview.
Copyright @ Zahara Chetty PTY LTD 2026