AI Safety & What to Watch For
Hallucination, privacy, bias and accountability — the four risks that actually affect ordinary users, and the habits that contain them.
What you'll be able to do
- Identify the specific conditions that make hallucination likely
- Decide what is safe to type into a chatbot and what is not
- Explain why the person who publishes AI output is the one accountable for it
Assumes: Lesson 2 — How Chatbots Actually Work
Four risks that matter
Plenty is written about long-term AI risk. This lesson is about the four things that can affect you personally, this week.
1. It will be confidently wrong
Not occasionally — routinely, and in a specific pattern. Hallucination risk is predictable, which makes it manageable.
High risk:
- Precise numbers — statistics, prices, dates, dosages, measurements.
- Citations — paper titles, authors, case numbers, URLs. Fabricated references are the classic failure, and they look completely real.
- Niche topics with little written about them.
- Leading questions. Ask “why does X cause Y?” and you will get reasons, whether or not X causes Y.
Low risk:
- Summarising or rewriting text you supplied.
- Generating options you will judge yourself.
- Explaining well-documented, mainstream subjects.
The rule that follows: the more specific and checkable a claim is, the more it needs checking. And note the trap — asking the model “are you sure?” is worthless. It will apologise and revise whether or not it was wrong the first time.
2. Whatever you type leaves your machine
Depending on the product and its settings, what you send may be stored, retained for a period, reviewed by staff, or used to improve future models. Terms vary and change.
You do not need to memorise anyone’s privacy policy. One rule covers it:
Do not type anything you would not be comfortable appearing outside your control.
Do not send: passwords or keys, card or bank details, government ID numbers, medical records, other people’s personal data, confidential client or employer material under NDA.
Generally fine: your own general questions, public information, creative work, code with secrets removed, documents you have redacted.
Two practical points. Redaction is the control you actually hold — swap real names for placeholders and the task usually works just as well. And telling the model to forget something does nothing; that instruction is just more text, it has no authority over storage.
If you handle regulated data at work, your employer’s rules on approved tools override all of this.
3. Bias is in the material, not the motive
The model learned from text people wrote. That text contains every imbalance of the societies that produced it — whose expertise gets described as authoritative, whose experiences are treated as default, which languages and regions have thorough coverage and which have almost none.
Reproducing patterns faithfully means reproducing skews faithfully. Nobody has to intend it.
In practice this shows up as: uneven answer quality across cultures and languages, default assumptions about who does which job, and thin or subtly wrong treatment of places and communities that are under-represented online.
Being aware of it is most of the defence. When output touches people, groups or cultures, read it the way you would read a stranger’s opinion rather than a reference work.
4. You are accountable, not the tool
If you publish it, send it, submit it or act on it, it is yours. Every provider’s terms say so, and so does every professional and academic norm.
This has a practical consequence people underestimate: AI output raises your review burden, it does not lower it. A confident, fluent, well-formatted paragraph is harder to read sceptically than a rough human draft — it looks finished. That fluency is exactly what makes an unverified claim inside it dangerous.
The habits, in short
- Verify anything specific and consequential — independently, not by asking again.
- Redact before you paste.
- Read output about people with the same scepticism you would apply to a stranger’s opinion.
- Never publish what you have not read properly yourself.
- Use it to prepare, not to decide, where the stakes are medical, legal or financial.
Remember: none of this makes AI unsafe to use. It makes it a tool with known failure modes — which is true of every tool. What separates people who get real value from it from people who get burned is simply knowing which failure mode is in play.
Go deeper
Quick Quiz
Test what you just learned. Pick the best answer for each question.
Q1 Which answer should you be most suspicious of?
Q2 You want help rewording a sensitive work document. What is safest?
Q3 Why can AI output carry bias even with no ill intent anywhere?
Q4 You publish a report containing an AI-invented statistic. Who is responsible?