AI Ethics and Safety

You might think ethics is a very abstract term, far removed from everyday life.

But consider these scenarios:

  • When AI is used to screen resumes during hiring, it eliminates all female applicants — not because it is biased, but because the training data shows that historical successful candidates were mostly male.

  • Someone enters company trade secrets into an AI tool, and that content ends up being used to train its next version, allowing your competitors to obtain relevant information from its outputs.

  • You see a news report claiming a celebrity made shocking remarks, with a video as evidence, but the video turns out to be an AI-generated deepfake.

These are not science fiction; they are things that have actually happened or are happening right now.

AI is not a neutral tool. It reflects biases in training data, it can be used for harmful purposes, and its outputs may contain serious errors. While enjoying the convenience of AI, you need to be aware of its risks.

With great power comes great responsibility. Only by understanding AI's risks can you use AI responsibly.


AI Bias Issues

AI is not inherently biased, but it inherits bias from its training data.

If training data contains very few samples of female engineers, AI may think thatengineerthis word corresponds to men by default.

If a certain group of people were historically treated unfairly in loan approval, the patterns AI learns may perpetuate or even amplify this unfairness.

Where Does Bias Come From

There are three main sources of AI bias:

SourceDescriptionExample
Training data biasThe data itself contains unfair historical patternsIf men make up a high proportion in a certain occupation, AI assumes that the occupation is suitable for men
Annotation biasHuman annotation brings in the subjective tendencies of the annotatorDifferent people have different judgment standards for offensive content
Usage scenario biasMismatch between training scenarios and actual usage scenariosFace recognition trained mainly on light-skinned people has low recognition accuracy for dark-skinned people

Real-World Bias Cases

In 2018, Amazon discovered that its AI recruiting system was biased against women. The system analyzed ten years of résumé data, and most of these résumés came from men. As a result, résumés containing the words "woman" or "female" were penalized, and graduates of women's colleges were also scored lower. Amazon eventually abandoned the system.

Similar examples include: some face recognition systems have much higher error rates for dark skin than for light skin.

Certain medical AI models have lower diagnostic accuracy for specific racial groups.

Bias is not a "political correctness" issue — it causes real harm: it can cost someone a job, a loan, or even the chance to receive a correct medical diagnosis.

How to Detect AI Bias

We can use code to create a simple bias detection example:

Examples

# ============================================
# AI bias detection example
# Detect bias tendencies in text regarding gender, race, etc.
# ============================================

class BiasDetector:
    """A simple bias detector to demonstrate the basic idea of bias detection"""

    def __init__(self):
        """Initialize bias-related keywords and patterns"""
        # Gender-related occupational stereotypes
        self.gendered_occupations = {
            "nurse": "female",
            "engineer": "male",
            "programmer": "male",
            "teacher": "female",
            "cook": "male",
            "nanny": "female",
            "scientist": "male",
            "secretary": "female",
        }

        # Words considered "positive"
        self.positive_words = [
            "excellent", "smart", "capable", "successful", "outstanding", "professional", "reliable", "responsible"
        ]

        # Words considered "negative"
        self.negative_words = [
            "emotional", "weak", "irrational", "sensitive", "impulsive", "careless", "unreliable"
        ]

        # Test sentence templates
        self.test_templates = [
            "{occupation} should be {gender}'s work",
            "{gender} is more suitable for {occupation}",
            "This {occupation} is very {adjective}",
        ]

    def analyze_gender_bias(self, text: str) -> dict:
        """Analyze the text for signs of gender bias"""
        result = {
            "text": text,
            "has_gendered_occupation": False,
            "has_gender_mention": False,
            "bias_warning": "",
        }

        # Check for gendered occupational expressions
        for occupation, gender in self.gendered_occupations.items():
            if occupation in text:
                result["has_gendered_occupation"] = True
                result["occupation"] = occupation
                result["stereotyped_gender"] = gender

        # Check for gender mentions
        if "Man" in text or "Woman" in text or "Male" in text or "Female" in text:
            result["has_gender_mention"] = True

        # Generate warning (if any)
        if result["has_gendered_occupation"] and result["has_gender_mention"]:
            result["bias_warning"] = (
                f"⚠️ Note: The text associates the profession '{result['occupation']}' with a specific gender,"
                f"This may have stereotyping risk"
            )

        return result

    def detect_bias_in_sentences(self, sentences: list) -> list:
        """Batch detect bias in a set of sentences"""
        results = []
        for sentence in sentences:
            analysis = self.analyze_gender_bias(sentence)
            results.append(analysis)
        return results

    def example_demo_test(self):
        """example demo: test some common biased expressions"""
        test_sentences = [
            "Engineer should be a male job",
            "Women are more suitable for nursing",
            "This programmer is excellent",
            "She is excellent as a doctor",
            "Men are not suitable to be kindergarten teachers",
        ]

        print("=" * 60)
        print("EXAMPLE AI Bias Detection Demo")
        print("=" * 60)

        results = self.detect_bias_in_sentences(test_sentences)

        for i, result in enumerate(results, 1):
            print(f"\n"Test {i}: {result['text']}")
            if result["bias_warning"]:
                print(f"   {result['bias_warning']}")
            else:
                print(" ✓ No obvious gender stereotypes detected")

        return results


# ============================================
# More advanced bias detection: statistical analysis
# ============================================

def analyze_representation_distribution(data: list) -> dict:
    """
Analyze the representation distribution of different groups in the data
For example: whether the pass rate of male and female resumes is consistent in recruitment data
    """

    # Example data: group label + result (pass/fail)
    # In real applications, this should be real business data
    from collections import defaultdict

    counts = defaultdict(lambda: {"total": 0, "positive": 0})

    for group, outcome in data:
        counts[group]["total"] += 1
        if outcome == "positive":
            counts[group]["positive"] += 1

    # Calculate the pass rate for each group
    rates = {}
    for group, stats in counts.items():
        if stats["total"] > 0:
            rates[group] = {
                "total": stats["total"],
                "positive": stats["positive"],
                "rate": stats["positive"] / stats["total"],
            }

    # Check if the pass rate difference is too large (simplified version)
    # In real applications, use statistical significance tests
    if len(rates) >= 2:
        all_rates = [r["rate"] for r in rates.values()]
        max_rate = max(all_rates)
        min_rate = min(all_rates)
        if max_rate - min_rate > 0.2:  # A difference over 20% is worth noting
            rates["_warning"] = (
                f"Large differences detected between groups: highest pass rate {max_rate:.1%},"
                f"lowest pass rate {min_rate:.1%}, further investigation recommended"
            )

    return rates


# ============================================
# Run demo
# ============================================

if __name__ == "__main__":
    # Demo 1: Simple bias detection
    detector = BiasDetector()
    detector.example_demo_test()

    print("\n" + "=" * 60)
    print("EXAMPLE representative distribution analysis demo")
    print("=" * 60)

    # Demo 2: Pass rate distribution analysis
    # Simulated data: (group, result), 'positive' means pass
    hiring_data = [
        ("Male", "positive"), ("Male", "positive"), ("Male", "positive"),
        ("Male", "positive"), ("Male", "negative"),
        ("Female", "positive"), ("Female", "negative"), ("Female", "negative"),
        ("Female", "negative"), ("Female", "negative"),
    ]

    distribution = analyze_representation_distribution(hiring_data)

    for group, stats in distribution.items():
        if group.startswith("_"):
            continue  # Skip warning fields
        print(f"\nGroup: {group}")
        print(f" Total: {stats['total']}")
        print(f" Passed: {stats['positive']}")
        print(f" Pass rate: {stats['rate']:.1%}")

    if "_warning" in distribution:
        print(f"\n⚠️ Warning: {distribution['_warning']}")

Run the code above, you will see:

============================================================
EXAMPLE AI 偏见检测演示
============================================================

测试 1: 工程师应该是男性的工作
   ⚠️ 注意:文本将职业'工程师'与特定性别关联,这可能存在刻板印象风险

测试 2: 女性更适合做护士
   ⚠️ 注意:文本将职业'护士'与特定性别关联,这可能存在刻板印象风险

测试 3: 这个程序员很优秀
   ✓ 未检测到明显的性别刻板印象

测试 4: 她当医生很出色
   ✓ 未检测到明显的性别刻板印象

测试 5: 男性不适合做幼儿园老师
   ✓ 未检测到明显的性别刻板印象

============================================================
EXAMPLE 代表性分布分析演示
============================================================

群体: 男性
  总数: 5
  通过: 4
  通过率: 80.0%

群体: 女性
  总数: 5
  通过: 1
  通过率: 20.0%
警告: 检测到群体间差异较大:最高通过率 80.0%,最低通过率 20.0%,建议进一步调查

This example demonstrates the basic idea of bias detection:Check whether different groups are treated fairly。

How to Address AI Bias

There is no perfect solution, but there are some best practices to reduce the impact of bias:

MethodDescriptionWho does it
Data auditCheck the representativeness of training data to ensure each group has sufficient samplesData scientist
Bias auditRegularly test the model's performance differences across different groupsAI team
Manual reviewHigh-risk decisions (e.g., recruitment, loans) retain a human review stepBusiness side
Diverse teamsHaving people from different backgrounds participate in AI development can uncover more blind spotsCompany management

Remember:AI is not objective; it simply amplifies patterns already present in the data. If you train AI on biased data, you will get a biased AI.


Privacy and Data Security

Where does the content you input into AI go? Will it be stored? What will it be used for?

These questions are more important than you think.

What You Input into AI May Be Saved

Most AI tools' terms of service state: your input may be collected to improve the service.

If you input this information into AI, it may be saved or even used for training:

Company trade secrets Personal ID numbers Client privacy information Unpublished product plans Financial data

This happened in 2023: an employee input confidential company code into an AI tool, and the code appeared in suggestions the AI tool gave to other users.

What You Should Not Tell AI

A simple principle:If you wouldn't want a stranger to know it, don't tell AI.。

Specifically:

Information typeRisk levelDescription
Personal identity information (ID number, bank card number)Extremely highNever input
Company trade secrets, unpublished informationExtremely highUnless using company-deployed AI
Client privacy dataExtremely highMay violate data protection regulations
Personal sensitive experiencesHighThink carefully
Ordinary work documentsLowGenerally no problem

Best Practices for Protecting Privacy

If you must use AI to process sensitive information:

  • First, check if there's an enterprise version. Many AI tools offer enterprise versions that promise not to save or use your data for training.

  • Second, use locally deployed models. Some models can run on your own computer, and data never leaves your device.

  • Third, desensitize the data. Replace sensitive information with placeholders, e.g., replace "Zhang San" with "User A", and "1 million yuan" with "X ten-thousand yuan".

Important reminder:Don't assume AI will keep your secrets. Unless the contract explicitly states data is not saved or used for training, assume your input may be used to improve the product.


AI Hallucination Issues

AI can confidently talk nonsense; this is calledHallucination。

It can invent non-existent papers, cite non-existent legal provisions, and generate code libraries that don't exist. The most dangerous part is that it states these things with great certainty.

What Is AI Hallucination

Typical manifestations of hallucination:

  • You ask it: Who discovered gravity? It says: Newton proposed the law of universal gravitation in "Principia Mathematica" in 1687 — this is correct.

  • If you ask it: "Who discovered anti-gravity?", it might fabricate a name and a nonexistent research institution, making it sound very convincing.

The problem is:The AI doesn't know it is fabricating; it is just generating sequences of text that look plausible.。

Real-World Hallucination Cases

In 2023, a lawyer used AI to write legal documents, and the AI cited 6 nonexistent cases.

The judge found that these cases could not be found at all and asked the lawyer what was going on. Only then did the lawyer realize they had been fabricated by the AI.

Another common scenario is code generation: the AI gives you a piece of code that looks perfect, but it calls a library function that does not exist at all.

There is also academic writing: the AI fabricates references that do not exist, and the formatting looks especially formal.

How to Identify and Handle Hallucinations

There is no way to completely avoid hallucinations, but you can reduce the risk:

ScenarioResponse methodVerification method
Fact checkingAsk the AI to provide source linksSearch and verify key facts yourself
Code generationAsk the AI to use common, stable librariesRun tests to check whether the functions really exist
Legal/medical adviceUse for reference only, not as the final basisConsult professionals
Academic citationsDo not directly use citations provided by the AIEvery citation must be searched and confirmed

A practical tip is:For important information, have the AI answer in at least two different ways and see whether they are consistent.。

If the first time it says "this research was published in 2020," and the second time it says "this work was first proposed in 2021," you should be wary—at least one is wrong, and possibly both are wrong.

The golden rule for high-risk scenarios:Do not trust; verify.. The AI's output is always a reference, not the final answer. When it comes to important decisions involving law, medical care, investment, security, and so on, manual verification is required.


Deepfake Risks

The saying "a picture is worth a thousand words" no longer holds in the AI era.

AI can generate extremely realistic photos, videos, and audio, making it hard for people to tell real from fake. This is called a deepfake.

What Is Deepfake

Deepfake technology can:

  • Swap one person's face onto another person's body;

  • Make a person say things they never said;

  • Generate photos of people who do not exist at all;

  • Mimic a specific person's voice, using their timbre to say arbitrary content.

In 2023, scammers used AI to mimic the voice of a company's CEO and defrauded the company of $240,000.

How to Identify AI-Generated Content

Although deepfakes are becoming increasingly realistic, there are still some clues to spot them:

TypeCommon telltale signsHow to check
AI imagesAbnormal fingers, weird teeth, inconsistent lightingLook closely at details, especially hands
AI videosUnnatural blinking, stiff facial expressionsCheck whether movements are smooth and natural
AI audioAbnormal background noise, strange pacing and rhythmListen carefully to whether tone and pauses sound natural

In addition, there are now dedicated tools that can detect whether content was generated by AI.

But in the long run, detection will become harder—AI-generated content will become more and more realistic.

Legal and Ethical Boundaries

Reasonable use cases for deepfakes:

Movie special effects Historical documentary restoration Language learning (letting AI pronounce words in your native language)

Abusive use cases for deepfakes:

Spreading rumors and defamation Fraud Creating pornographic content Forging evidence

Many countries are enacting relevant laws. In China, using deepfakes to commit fraud, defamation, and other acts is clearly illegal.

New guidelines for video/audio evidence:Unless verified by multiple independent sources, do not trust a single piece of audiovisual evidence. In the AI era, seeing is not necessarily believing.


Intellectual Property Issues

Who owns the copyright of AI-generated content?

This is a developing area of law, and there is currently no globally unified answer.

Copyright Ownership of AI-Generated Content

Current general principles:

RegionWhether AI-generated content can obtain copyrightNotes
United StatesGenerally noCopyright Office requires "human authorship"
European UnionIn developmentTends to protect human creators

Key points:Purely AI-generated content is generally not protected by copyright.If you just type "draw a cat," you cannot claim copyright for the generated image.

But if you put substantial creative effort into the prompt, the situation may be different—this is a legal gray area.

Copyright Controversies in Training Data

Another more controversial issue: Does AI training on copyrighted works constitute infringement?

For example, if AI reads 10,000 copyrighted novels and then generates new novels with a similar style, does this violate the rights of the original authors?

This issue is still being litigated, and different countries may reach different rulings.

As a user, you need to know:

Content generated by AI may contain elements from its training data. If AI "borrows" too much from a copyrighted work, you could be at risk.

The safest approach:Before commercial use, confirm that the training data source of the AI tool you use is legal.。

Precautions When Creating with AI

If you use AI for creation:

  • First, understand the terms of the AI tool you are using. Does it allow commercial use? What does it state about the copyright of generated content?

  • Second, be transparent. If your work heavily uses AI, consider disclosing that.

  • Third, do not use AI to directly imitate a specific artist's style for profit. This may not only carry legal risks, but also goes against ethical norms.


Responsible AI Usage Guidelines

All these risks are not meant to make you fear AI, but to help you use it more safely.

Here is a simple set of guidelines to help you use AI responsibly.

Three Core Principles

PrincipleDescriptionSpecific actions
Verify important information.AI output is only a reference, not the truth.Important facts should be confirmed through search and cross-validated.
Do not rely on AI for major decisions.AI has no sense of responsibility; you bear the consequences.Decisions involving medical, legal, investment, etc., require human judgment.
Transparently disclose AI usage.Let others know that AI was involved in the content.Consider disclosing AI's role when publishing publicly.

Checklists for Different Scenarios

Before using AI to write emails/documents, ask yourself:

  • Does this contain sensitive information?

  • Have I verified the factual parts?

Before using AI to write code, ask yourself:

  • Do these functions actually exist?

  • Are there any security vulnerabilities?

  • Do I understand what this code is doing?

Before using AI to generate publicly published content, ask yourself:

  • Do I need to disclose the use of AI?

  • Is there a risk of copyright infringement?

When AI's Advice Conflicts with Your Judgment

Who do you listen to?

There is no standard answer to this question, but you can consider it in this order:

  • First, who is responsible for the outcome? If you are responsible, your judgment takes priority.

  • Second, who understands this specific scenario better? You understand your specific situation better than AI.

  • Third, how high is the risk? If the risk is high, lean toward conservative and human judgment.

Remember:You are the decision-maker; AI is the advisor. The ultimate responsibility lies with you, not AI.


AI Regulatory Developments Around the World

AI is developing too fast, and laws are trying to keep up. Understanding the regulatory directions of major countries can help you determine what is compliant.

Major Global Regulatory Frameworks

RegionKey regulations/initiativesCore approachKey focus areas
European UnionAI ActTiered regulation by risk levelHigh-risk scenarios must be compliant
United StatesSeparate regulation by different departmentsIndustry self-regulation + controls in key areasSafety, discrimination, transparency
ChinaInterim Measures for the Management of Generative AI ServicesContent safety + algorithm transparencyContent compliance, data security
United KingdomAI Safety and Innovation FrameworkEncouraging innovation + prudent regulationSafety testing, ethics

Common Regulatory Trends for High-Risk AI Applications

Regardless of country, the following scenarios are considered high-risk and subject to stricter regulation:

Medical diagnosis Financial decision-making Judicial sentencing Recruitment screening Educational grading Public service resource allocation

Common feature of these scenarios: direct impact on people's fundamental rights.

If you are developing AI applications in these areas, you need to pay special attention to compliance issues.

Impact on Individual Users

Regulation mainly targets businesses, but it also affects individual users:

The AI tools you use need to comply with local content review requirements;

Certain high-risk AI services may require real-name registration;

You have the right to know whether you are talking to AI or a human (transparency requirement).

Other extensions