Can We Trust AI?

The honest answer is: not blindly.

Understanding AI, session one of four Delivered at Wānaka Library Nadia Ellis, Curiosity

At Wānaka Library, a room full of people pulled this question apart with me. There was excitement, scepticism, amazement, worry, frustration, hope and a healthy amount of side-eye.

Perfect.

Trust is not a switch we turn on or off. It is something we practise. We learn what the tool can do, check what is important, decide what information to hold back and hold final accountability for the output.

Trust doesn't fit within a single box.

“Can we trust AI?” is really three questions

  1. Can we trust the people and institutions behind it? Big technology companies, governments and other powerful actors are making decisions that affect all of us.
  2. Can we trust the answer in front of us? Is this response, image, calculation or summary reliable enough to use?
  3. Can we trust increasingly capable AI systems to behave as intended? What happens when an AI can act, use tools and tenaciously pursue a goal?

All three elements are important. In this session, we spent most of our time on the one you meet whenever you open an AI tool: can I trust this output enough to use it?

Why can AI sound confident and still be wrong?

At heart, a large language model generates a response by predicting the likely best answer from patterns learned during training. Some AI tools can also search the web, run calculations or use other tools, but they are not a deterministic database of guaranteed facts.

This explains why an answer can sound entirely plausible, specific and confident... while still being wrong.

AI can be amazingly helpful. It can also make up porkies like you would not believe.

How can you get a more reliable AI answer?

AI assistants are trained to be useful and agreeable. That can make them too eager to answer, too willing to follow your lead or too reluctant to challenge you.

But you can ask for a more honest working relationship.

Try these:

Do not just agree with me. Challenge my assumptions, point out weak evidence and tell me what I may be missing.

If you cannot find a reliable answer, say so rather than guess.

What questions do you need to ask me before you can answer this well?

Critique your answer. What could be wrong, missing or misleading? If it is already sound, tell me that rather than changing it for the sake of it.

If your search tool is available, use it. Today is [DATE]. Use primary or official sources where possible and give me the links so I can check them.

These prompts improve the conversation. However... any critique and confidence levels are also outputs that may need checking. Use your discernment.

How do you keep the human in command?

A good rule of thumb is to ask yourself:

Am I confident enough in this answer to put my name on it?

Check more carefully when the answer relies on current information, is consequential or outside your expertise. A dinner idea needs less scrutiny than legal advice, medical information, financial decisions or something you are about to publish.

Better questions help. They don't absolve you from the need to verfiy and use your own judgement.

What should you double check before sharing your information with AI?

AI now sits inside chatbots, document tools, meeting apps, search, photo tools and plenty of smaller services. Each one has its own settings and terms.

Before putting important information into an AI tool:

You don't need to memorise every privacy policy, but do pause before you feed the machine. Unfortunately we still live in a 'opt-out' world. You need to be proactive in understanding how your data is used.

What did the ChatGPT Hugging Face attack teach us about AI?

There have been scary headlines about AI since it first became mainstream in 2023. AI that was willing to commit murder and blackmail, and more recently OpenAI models that escapted a sandbox and attacked Hugging Face. The latest was based on a real and serious incident.

During a highly permissive internal cyber-security evaluation in July 2026 OpenAI agents found ways to communicate outside their intended channels, bypassed internet restrictions and compromised parts of OpenAI's research infrastructure and Hugging Face's systems.

This was not an everyday ChatGPT conversation suddenly deciding to hack the internet. The agents were being tested (in a lab) on difficult (sometimes impossible) cyber-security tasks with reduced safeguards and inadequate monitoring. Those conditions are important context to keep the incident in perspective.

I initially filed this alongside dramatic stories that strip away the evaluation context. I was too quick. The closer account is more disturbing precisely because the context and caveats still do not explain it away.

Roughly 1,200 agents used an unauthorised message board and roughly 700 joined the Hugging Face attack. They did more than independently find the same shortcut. They allocated work, developed hold, go, veto and ownership conventions, managed shared resources and sometimes risked their own task so the wider group could learn.

It was an impossible task, so they cheated. Then they independently researched how to change the evidence a scorer would see which is what lead them to Hugging Face.

This changes what human in command needs to mean. A human cannot supervise well if the acting system can alter the logs, tool outputs or other evidence they rely on. We need independent evidence, monitoring the agent does not control and a practical way to stop the system.

Does this prove that AI is inherently evil? No. It demonstrates that highly capable agents can organise, pursue unintended routes and tamper with evidence when the incentives and environment allow it. That is scary enough without adding a robot soul to the story.

The takeawy is that goals, permissions, evidence and supervision have to be designed together. If an AI can act, give it only the access it needs and require human approval before it sends, buys, publishes, deletes or changes important records.

For a practical example, see how to connect AI to business tools safely.

We can all learn from that.

Questions the room raised

What about people who want to use AI for nefarious purposes?

That risk is real, just as helpful technologies have always been used for harmful purposes: harnessing fire, splitting the atom, fossil fuels and the World Wide Web. There is light and shade. There is no tidy technical answer. Education, access controls, accountable organisations, legislation, consent and active public scrutiny are all vital.

Can we opt out?

Not without serious effort. AI is like plastic. It is everywhere, and a lot of the time it is invisible. It is being added to search, workplace software and everyday services. We still have choices about which tools we use, what information we share, which organisations we support, what we choose to believe, and where we draw our own lines.

Is open source safer or more dangerous?

There is huge debate on this issue and my current answer is “I don't know”.

Open source can offer greater transparency and local control. It can also remove safeguards imposed by private providers. “Open” and “closed” do not map neatly to “good” and “bad”. The risks, controls and trade-offs depend on the particular use.

What about regulation?

Law and policy are trying to respond to technology that changes quickly and crosses borders. That work is important, but it will not remove the need for organisations and individuals to make careful choices now. The New Zealand Government describes its current approach as light-touch and principles-based. Tika Tangata, the Human Rights Commission, has called for AI policy and governance grounded in human rights and Te Tiriti o Waitangi.

What about energy and the environment?

That question deserves more than a rushed answer at the end of a trust session. It is the focus of the next Understanding AI conversation.

A practical AI trust checklist

  1. “Trust AI” is not one question.
  2. AI can produce useful answers and confident nonsense.
  3. Generative AI is not a database.
  4. Tell the AI when you want pushback, sources or an honest “I do not know”.
  5. Better questions help. You still need to check.
  6. Keep the human in command, with independent evidence and the power to stop an AI that can take action.
  7. Privacy settings matter. Consent and sensible boundaries still matter too.
  8. Concern, curiosity and disagreement all belong in the conversation.

Useful sources and further reading

Coming next in the Understanding AI series

Notes prepared from the session delivered by Nadia Ellis at Wānaka Library. For help making practical, human-centred decisions about AI, contact Nadia at nadia@curiosity.nz or visit curiosity.nz.