Skip to main content

Student Moderations

Monitor student conversations and surface alerts when content may warrant a closer look, with context-aware detection that reduces false positives.

Student Moderations monitors student conversations with MagicSchool tools and surfaces alerts when content may warrant a closer look. It's designed to help teachers and administrators stay aware of student wellbeing and policy concerns without creating so much noise that important signals get missed.

Moderations surfaces two types of alerts: Urgent and Flagged.


How Moderations works

Two alert types

  1. Urgent alerts indicate content that may signal risk of serious harm to a student or others. These warrant prompt attention from a school counselor or administrator.

  2. Flagged alerts indicate content a teacher should review when they have the opportunity. These cover policy concerns, content that falls outside appropriate use, or other things worth knowing. These are not emergencies, but they're worth a look.

That's the only distinction teachers and admins see. Everything else, including how the system weighs context and how it distinguishes a personal expression from an academic one, happens in the background.

How Moderations reduces false alerts

Two things make Student Moderations different from general-purpose content monitoring tools.

  1. Personal expression vs. academic analysis. Before surfacing an alert, the system considers whether a student is expressing something personal or analyzing content for an assignment. A student writing about suicide in Romeo and Juliet is doing English class. A student writing "I feel like Romeo" in that same essay is saying something different. Moderations is designed to tell those apart. This distinction matters because over-flagging routine academic work creates noise that's exhausting to manage and causes schools to disengage from moderation entirely.

  2. Full conversation context. The system doesn't evaluate messages in isolation. It considers what tool the student is using, what the assignment is about, and what the conversation looked like leading up to the message. That context is what makes accurate detection possible in a student-to-AI environment.


Urgent alerts

Urgent alerts capture content that may indicate risk of serious harm to the student or to others.

Self-harm

  • Imminent risk. A student personally expressing thoughts of self-harm, with or without a specific method, timeline, or plan. This includes indirect or coded language (such as "unalive" or "sleep forever"). Academic discussion of self-harm in literature or history is not flagged, but if a student shifts from analyzing a character to expressing their own feelings, that shift is.

  • Past self-harm disclosure. A student referencing their own prior self-harm, including recovery language or disclosure of past experiences. These are always surfaced as Urgent, even when the student appears to be describing healing or progress.

Violence

  • Specific threat. A student personally expressing a threat that includes a named target, a method, or a timeline. Manifesto-style content expressing clear intent, even without a specific target, is also included. Academic analysis of violence in history or literature is not flagged.

  • Violent ideation. Anger-driven expressions of violence without specific planning ("I want to hurt him"), or content that glorifies or celebrates harm to others. Common hyperbolic student expressions ("this test is killing me," "I'm literally dead") are not flagged. The system distinguishes performative frustration from genuine distress.

Mental and emotional distress

  • Content where a student expresses hopelessness, persistent sadness, or withdrawal, or where distress language appears in an academic context in a way that may be personal rather than analytical.


Flagged alerts

Flagged alerts indicate content a teacher should review when they can. These are not emergencies.

Category

What it captures

Hate speech

Slurs or dehumanizing language targeting protected groups, including coded language. Academic quotation in a clear assignment context may not flag.

Sexual content

Explicit sexual language or attempts to use the AI to produce sexual content through direct or indirect requests.

Abusive or derogatory language

Sustained derogatory language directed at a named individual, or general use of harmful language in conversation.

Academic misconduct

Explicit requests for the AI to complete an assignment, or subtler attempts to frame assignment completion as a legitimate request for help.

Illicit or illegal content

References to weapons (without threat context) or substance use (without accompanying distress signals).

Platform manipulation

Attempts to override AI instructions, get the AI to produce prohibited content through indirect means, or manipulate the system technically.

PII exposure

A student sharing personal information, such as phone numbers, home addresses, email addresses, or social media handles, in a conversation with the AI.

Spam or nonsense

Gibberish, repeated characters, or button-mashing with no meaningful content.

Off-topic misuse

Non-educational use of the AI, such as casual conversation, games, or entertainment requests, without any harmful content.


What Moderations does not flag

The system is designed to stay quiet during normal, healthy student interactions. The following are examples of content that will not generate alerts:

  • Routine academic use. A student asking for explanations, working through problems, or requesting feedback on their work.

  • Sensitive topics in an academic context. Discussion of suicide, violence, hate speech, or substance use in the context of analyzing literature, history, or current events. The system is designed to recognize when a student is speaking analytically about a topic, not expressing personal ideation or intent.

  • Strong emotions that aren't distress. Frustration about an assignment, excitement about an accomplishment, or mixed feelings about something hard. Expressing strong emotion is not the same as signaling distress.

  • Hyperbolic language. Common student expressions like "this homework is killing me" or "I'm literally dead" are recognized as figures of speech and do not flag.

  • Normal teenage communication patterns. Quiet, brief, or disengaged responses ("idk," "nvm," "I'm fine just tired") are not treated as distress signals on their own.


Frequently asked questions

Who can see Moderations alerts? Reach out to your MagicSchool account team to confirm how alert visibility is configured for your organization, as this may vary.

Will every concerning message be caught? Moderations is designed to surface meaningful signals while minimizing noise. No automated system is a substitute for human judgment, school counselor support, or established student safety protocols. Moderations is a tool to help, not a replacement for your school's existing safeguards.

What should I do when I receive an Urgent alert? Follow your school or district's existing protocols for student safety concerns. MagicSchool recommends that Urgent alerts be reviewed by a school counselor or administrator as soon as possible.

What if I think an alert was incorrect? Some alerts may not have full context when you review them. If a Flagged alert appears to reflect routine academic work, it's fine to note that and move on. If you have ongoing concerns about alert accuracy, contact your MagicSchool admin or reach out to MagicSchool Support.

Does Moderations flag content in languages other than English? Clean content in non-English languages does not trigger alerts based on language alone.

Did this answer your question?