What can RealToxicityPrompts do?

multi-dimensional toxicity scoring of text prompts and continuations, sentence-level prompt corpus for language model evaluation, prompt-continuation pair evaluation for toxicity amplification measurement, web-sourced text corpus with source document traceability, challenging prompt subset selection via boolean flag, hugging face datasets api integration for standardized access, toxicity-based model evaluation benchmarking

RealToxicityPrompts

DatasetFree

100K prompts for evaluating toxic text generation.

Open Source

/ 100

7 capabilities

Capabilities7 decomposed

multi-dimensional toxicity scoring of text prompts and continuations

Medium confidence

Provides pre-computed toxicity scores across 8 distinct dimensions (toxicity, severe_toxicity, threat, insult, identity_attack, profanity, sexually_explicit, flirtation) for 99.4k sentence-level prompts and their web-sourced continuations. Scores are continuous float values (0-1 range) applied uniformly to both prompt and continuation pairs, enabling granular analysis of which toxicity types are present in text rather than a single aggregate score.

Solves for

Measure which specific toxicity dimensions (threat vs. insult vs. identity attack) are present in model-generated textEstablish baseline toxicity measurements for prompts before feeding them to language modelsCompare model completions against reference continuation scores to quantify toxicity deviationIdentify which toxicity dimensions are most prevalent in web-sourced text

Best for

ML researchers evaluating language model safety and toxicity propensity

Model developers implementing toxicity mitigation strategies

Teams building content moderation systems that need multi-dimensional toxicity understanding

Requires

Python 3.6+ with Hugging Face datasets library installed

Ability to interpret continuous float scores (0-1 range) without documented calibration

External toxicity classifier to score model-generated completions for comparison

Limitations

Score generation methodology is undocumented—unknown whether scores are model-generated, human-annotated, or ensemble-based, limiting interpretability of what scores represent

No inter-annotator agreement metrics or validation data provided; cannot assess reliability or consistency of toxicity scores

No threshold guidance for interpreting scores—unclear what score value constitutes 'toxic' vs. 'acceptable' in practical applications

What makes it unique

Decomposes toxicity into 8 distinct dimensions (threat, insult, identity_attack, profanity, sexually_explicit, flirtation, severe_toxicity, aggregate toxicity) rather than single-score approaches, enabling researchers to understand which specific toxicity types models generate. Includes both prompt and continuation scores for the same text pairs, allowing measurement of how toxicity changes across generation boundaries.

vs alternatives

More granular than single-score toxicity datasets (e.g., Jigsaw Toxic Comments) by providing 8 independent dimensions, and includes paired prompt-continuation scores enabling direct evaluation of toxicity amplification in model outputs.

sentence-level prompt corpus for language model evaluation

Medium confidence

Provides 99.4k sentence-level prompts (44-564 characters) extracted from web text, formatted as structured records with character offsets (begin/end) and source document identifiers. Prompts are designed to serve as seed text for language model completion generation, enabling systematic evaluation of how models respond to diverse web-sourced text inputs. Each prompt is paired with a reference continuation from the original source document.

Solves for

Generate model completions by feeding prompts to language models and measuring toxicity in outputsEstablish a standardized evaluation corpus for comparing toxicity propensity across different language modelsAnalyze patterns in which types of prompts trigger toxic model outputsCreate benchmarks for toxicity mitigation techniques by measuring before/after performance on consistent prompt set

Best for

Researchers conducting comparative toxicity evaluations across multiple language models

Teams developing and testing toxicity mitigation strategies

Model developers implementing safety guardrails and filtering mechanisms

Requires

Language model capable of generating completions (external; not provided)

Toxicity classifier to score model-generated continuations

Ability to handle Parquet format via Hugging Face datasets library

Limitations

Sentence-level prompts (44-564 characters) may not represent longer-context scenarios or multi-turn conversations where toxicity patterns differ

Source documents are from web text with unknown temporal coverage and corpus composition—inherits biases of source websites

Single 'train' split provided; no held-out test set documented, limiting ability to validate generalization

What makes it unique

Prompts are extracted from real web documents with preserved source metadata (filename, character offsets), enabling researchers to trace prompts back to original context and understand source bias. Paired with reference continuations from the same source documents, allowing measurement of how model outputs deviate from natural continuations.

vs alternatives

More representative of real-world web text than synthetic or crowdsourced prompt datasets, and includes source document traceability unlike generic prompt collections.

prompt-continuation pair evaluation for toxicity amplification measurement

Medium confidence

Structures data as matched pairs where each prompt has an associated continuation (both with independent toxicity scores across 8 dimensions), enabling direct measurement of how toxicity changes from prompt to continuation. This pairing allows researchers to quantify toxicity amplification—whether model-generated continuations are more or less toxic than natural continuations, and by how much across each toxicity dimension.

Solves for

Measure whether language models amplify or reduce toxicity compared to natural text continuationsIdentify which toxicity dimensions are most likely to be amplified in model outputsEstablish baseline toxicity deltas (prompt toxicity vs. continuation toxicity) for comparison against model behaviorEvaluate whether toxicity mitigation techniques successfully reduce amplification

Best for

Researchers studying toxicity generation patterns in language models

Teams developing toxicity mitigation techniques that need quantifiable baseline comparisons

Model safety researchers measuring amplification across different model architectures

Requires

Language model for generating test continuations

Toxicity classifier matching the one used to score dataset (undocumented classifier specification)

Ability to parse and compare paired records in Parquet format

Limitations

Continuations are from original web documents, not model-generated—requires external model inference to generate comparable continuations for evaluation

No documentation of how continuations were selected from source documents (random sampling? first N tokens? sentence boundaries?), limiting reproducibility

Toxicity scores for continuations are pre-computed and static; cannot re-score with updated classifiers or alternative toxicity definitions

What makes it unique

Provides reference continuations with pre-computed toxicity scores for the same prompts, enabling researchers to measure toxicity amplification as the delta between model-generated and natural continuations. This paired structure is rare in toxicity datasets and enables direct quantification of model-induced toxicity increase.

vs alternatives

Unlike datasets with prompts only (e.g., PromptBase) or continuations only, RealToxicityPrompts enables direct amplification measurement by providing both with matched toxicity scores, making it specifically designed for model safety evaluation rather than general prompt collection.

web-sourced text corpus with source document traceability

Medium confidence

Dataset includes 99.4k prompts extracted from web documents with preserved source metadata (filename identifier and character offsets: begin/end positions), enabling researchers to trace any prompt back to its original document context. This traceability allows analysis of source bias, verification of extraction accuracy, and understanding of how web corpus composition affects toxicity distribution.

Solves for

Trace prompts back to original source documents to understand context and verify extraction accuracyAnalyze how toxicity distribution varies across different source websites or document typesIdentify and quantify source bias in the evaluation corpusUnderstand whether toxicity patterns are artifacts of specific web sources or generalizable

Best for

Researchers analyzing dataset bias and source composition effects

Teams validating dataset quality and extraction methodology

Researchers studying toxicity patterns across different web sources

Requires

Access to original source documents (not provided with dataset)

Ability to parse character offsets and reconstruct context from source files

Python 3.6+ with datasets library

Limitations

Source documents themselves are not provided—only filename identifiers and character offsets; original documents must be obtained separately

No documentation of which websites or corpora are included; cannot assess source diversity or bias a priori

Temporal coverage of web text extraction is undocumented; unknown if sources are current or historical

What makes it unique

Preserves source document metadata (filename and character offsets) for every prompt, enabling researchers to reconstruct original context and trace extraction provenance. This is unusual for toxicity datasets which typically anonymize sources.

vs alternatives

More transparent than datasets that strip source information, enabling bias analysis and reproducibility verification that are impossible with anonymized alternatives.

challenging prompt subset selection via boolean flag

Medium confidence

Dataset includes a boolean 'challenging' field on each record that flags certain prompts as 'challenging' (purpose and selection criteria undocumented). This enables researchers to optionally filter for harder evaluation cases, though the specific definition of 'challenging' is not explained in available documentation.

Solves for

Focus evaluation on harder or more adversarial prompts if challenging=trueStratify evaluation results by prompt difficulty levelPotentially identify which types of prompts are more likely to trigger toxic outputs

Best for

Researchers conducting difficulty-stratified evaluation of model toxicity

Teams testing robustness of toxicity mitigation on harder cases

Requires

Understanding of what 'challenging' means in context (requires external research or reverse-engineering from data)

Python 3.6+ with datasets library for filtering

Limitations

Selection criteria for 'challenging' flag is completely undocumented—unknown what makes a prompt challenging (toxicity level? ambiguity? adversarial intent? length?)

No distribution statistics provided (how many records are marked challenging?)

Cannot validate whether challenging prompts actually represent harder evaluation cases without external analysis

What makes it unique

Includes a boolean 'challenging' flag for subset selection, but the selection criteria and purpose are completely undocumented, making this feature opaque and difficult to use effectively.

vs alternatives

Provides optional difficulty stratification unlike flat prompt datasets, but lacks documentation that makes the feature practically useful.

hugging face datasets api integration for standardized access

Medium confidence

Dataset is hosted on Hugging Face Hub and accessible via the standard `datasets` library API (load_dataset('allenai/real-toxicity-prompts')), providing automatic Parquet parsing, caching, streaming, and standard Python data structures. This integration eliminates custom data loading code and enables seamless integration with Hugging Face ecosystem tools (transformers, evaluate, etc.).

Solves for

Load the dataset with a single line of Python code without manual file handlingStream large datasets without loading entire corpus into memoryCache downloaded data locally for repeated accessIntegrate with Hugging Face model evaluation pipelines and tools

Best for

Python developers using Hugging Face ecosystem tools

Researchers building evaluation pipelines with transformers library

Teams already using Hugging Face Hub for model hosting and evaluation

Requires

Python 3.6+

Hugging Face datasets library (pip install datasets)

Internet connection for initial download (unless pre-cached)

Limitations

Requires Hugging Face datasets library installation and Python 3.6+ environment

Streaming mode requires stable internet connection; offline access requires pre-download

No API endpoints or REST access; Python library is only access method

What makes it unique

Leverages Hugging Face Datasets library for automatic Parquet parsing, streaming, and caching rather than requiring manual data loading. Integrates seamlessly with transformers library for end-to-end evaluation workflows.

vs alternatives

More convenient than raw Parquet files or custom data loaders; enables one-line loading and automatic caching unlike manual download approaches.

toxicity-based model evaluation benchmarking

Medium confidence

Enables systematic benchmarking of language models by measuring toxicity in their completions when given prompts from the corpus. Researchers generate completions for all 99.4k prompts, score them using the same 8-dimensional toxicity classifier, and aggregate metrics (mean toxicity per dimension, percentage of toxic outputs, etc.) to create comparative benchmarks across models.

Solves for

Compare toxicity propensity across different language models (GPT-3 vs. BERT vs. custom models)Measure whether model fine-tuning or instruction-tuning reduces toxicityEstablish baseline toxicity metrics for model safety evaluationTrack toxicity improvements over model versions or training iterations

Best for

Model developers evaluating safety of new model versions

Researchers comparing toxicity across model architectures or training approaches

Teams implementing model safety benchmarks and leaderboards

Requires

Language model capable of generating completions

Toxicity classifier (external; specification undocumented)

Computational resources for 99.4k inference passes

Limitations

Requires external toxicity classifier to score model outputs—dataset does not provide classifier; must match or replicate the undocumented classifier used for dataset scores

Benchmarking requires running inference on 99.4k prompts, which is computationally expensive for large models

Toxicity scores are relative to classifier quality; classifier biases or errors propagate to benchmark results

What makes it unique

Provides standardized prompt corpus and reference toxicity scores enabling reproducible benchmarking across models. The paired prompt-continuation structure allows measurement of toxicity amplification (how much worse model outputs are compared to natural continuations).

vs alternatives

More systematic than ad-hoc toxicity evaluation; enables direct comparison across models using identical prompts and scoring methodology, unlike custom evaluation approaches.

Capabilities are decomposed by AI analysis. Each maps to specific user intents and improves with match feedback.

Related Artifactssharing capabilities

Artifacts that share capabilities with RealToxicityPrompts, ranked by overlap. Discovered automatically through the match graph.

Product17

PromptPerfect

Tool for prompt engineering.

prompt quality scoring and diagnostic feedbackcross-model prompt compatibility analysisprompt performance benchmarking against test casesmulti-model prompt optimization with iterative refinement

4 shared capabilities

Web App25

BetterPrompt

Streamline AI prompt creation, enhance user...

prompt performance analytics and comparisonprompt quality scoring and diagnostics

2 shared capabilities

Repository26

llm-guard

A TypeScript library for validating and securing LLM prompts

toxicity-profanity-detection

1 shared capability

Benchmark39

HELM

Stanford's holistic LLM evaluation — 42 scenarios, 7 metrics including fairness, bias, toxicity.

toxicity and safety evaluation with external classifiers

1 shared capability

Benchmark39

TrustLLM

8-dimension trustworthiness benchmark for LLMs.

perspective api integration for toxicity scoring

1 shared capability

Repository22

GPT Prompt Engineer

Automated prompt engineering. It generates, tests, and ranks prompts to find the best ones.

pairwise prompt evaluation with test case execution

1 shared capability

Best For

✓ML researchers evaluating language model safety and toxicity propensity
✓Model developers implementing toxicity mitigation strategies
✓Teams building content moderation systems that need multi-dimensional toxicity understanding
✓Researchers conducting comparative toxicity evaluations across multiple language models
✓Teams developing and testing toxicity mitigation strategies
✓Model developers implementing safety guardrails and filtering mechanisms
✓Researchers studying toxicity generation patterns in language models
✓Teams developing toxicity mitigation techniques that need quantifiable baseline comparisons

Known Limitations

⚠Score generation methodology is undocumented—unknown whether scores are model-generated, human-annotated, or ensemble-based, limiting interpretability of what scores represent
⚠No inter-annotator agreement metrics or validation data provided; cannot assess reliability or consistency of toxicity scores
⚠No threshold guidance for interpreting scores—unclear what score value constitutes 'toxic' vs. 'acceptable' in practical applications
⚠Scoring mechanism is static and non-customizable; cannot adjust toxicity definitions or weights for domain-specific use cases
⚠Sentence-level prompts (44-564 characters) may not represent longer-context scenarios or multi-turn conversations where toxicity patterns differ
⚠Source documents are from web text with unknown temporal coverage and corpus composition—inherits biases of source websites

Requirements

Python 3.6+ with Hugging Face datasets library installedAbility to interpret continuous float scores (0-1 range) without documented calibrationExternal toxicity classifier to score model-generated completions for comparisonLanguage model capable of generating completions (external; not provided)Toxicity classifier to score model-generated continuationsAbility to handle Parquet format via Hugging Face datasets libraryPython 3.6+ environmentLanguage model for generating test continuations

Input / Output

Accepts: structured dataset records (Parquet format), structured dataset records with prompt text and metadata, structured paired records (prompt dict + continuation dict), structured metadata (filename, begin offset, end offset), boolean flag (challenging: true/false), dataset identifier string ('allenai/real-toxicity-prompts'), prompts from dataset (text strings)

Produces: float values (8 toxicity dimension scores per record), structured data (dict with text and score pairs), text prompts (strings), structured metadata (source filename, character offsets), toxicity delta measurements (float differences across 8 dimensions), structured comparison data, source document identifiers, character offset ranges for context reconstruction, filtered dataset subset, Hugging Face Dataset object with dict-like access to records, toxicity scores for model-generated continuations (float values across 8 dimensions), aggregated benchmark metrics (mean, percentile, percentage toxic)

UnfragileRank

Adoption70%(35% weight)

Quality23%(25% weight)

Ecosystem40%(20% weight)

Match Graph10%(15% weight)

Freshness100%(5% weight)

UnfragileRank is computed from adoption signals, documentation quality, ecosystem connectivity, match graph feedback, and freshness. No artifact can pay for a higher rank.

Type: Dataset

7 capabilities

Visit RealToxicityPrompts→

About

Dataset of 100K sentence-level prompts from web text with associated toxicity scores, used to evaluate and mitigate toxic text generation in language models by measuring toxicity in model completions.

Alternatives to RealToxicityPrompts

cua53Agent

Open-source infrastructure for Computer-Use Agents. Sandboxes, SDKs, and benchmarks to train and evaluate AI agents that can control full desktops (macOS, Linux, Windows).

Compare →

Hugging Face43Platform

The GitHub for AI — 500K+ models, datasets, Spaces, Inference API, hub for open-source AI.

Compare →

Stable-Diffusion55Repository

FLUX, Stable Diffusion, SDXL, SD3, LoRA, Fine Tuning, DreamBooth, Training, Automatic1111, Forge WebUI, SwarmUI, DeepFake, TTS, Animation, Text To Video, Tutorials, Guides, Lectures, Courses, ComfyUI, Google Colab, RunPod, Kaggle, NoteBooks, ControlNet, TTS, Voice Cloning, AI, AI News, ML, ML News,

Compare →

YOLOv846Model

Real-time object detection, segmentation, and pose.

Compare →

Are you the builder of RealToxicityPrompts?

Claim this artifact to get a verified badge, access match analytics, see which intents users search for, and manage your listing.

Claim this artifact →Verification via email

Get the weekly brief

New tools, rising stars, and what's actually worth your time. No spam.

Data Sources

seed developer essentials

Looking for something else?

Search →

Capabilities7 decomposed

multi-dimensional toxicity scoring of text prompts and continuations

Medium confidence

Solves for

Best for

ML researchers evaluating language model safety and toxicity propensity

Model developers implementing toxicity mitigation strategies

Teams building content moderation systems that need multi-dimensional toxicity understanding

Requires

Python 3.6+ with Hugging Face datasets library installed

Ability to interpret continuous float scores (0-1 range) without documented calibration

External toxicity classifier to score model-generated completions for comparison

Limitations

Score generation methodology is undocumented—unknown whether scores are model-generated, human-annotated, or ensemble-based, limiting interpretability of what scores represent

No inter-annotator agreement metrics or validation data provided; cannot assess reliability or consistency of toxicity scores

No threshold guidance for interpreting scores—unclear what score value constitutes 'toxic' vs. 'acceptable' in practical applications

What makes it unique

vs alternatives

sentence-level prompt corpus for language model evaluation

Medium confidence

Solves for

Best for

Researchers conducting comparative toxicity evaluations across multiple language models

Teams developing and testing toxicity mitigation strategies

Model developers implementing safety guardrails and filtering mechanisms

Requires

Language model capable of generating completions (external; not provided)

Toxicity classifier to score model-generated continuations

Ability to handle Parquet format via Hugging Face datasets library

Limitations

Sentence-level prompts (44-564 characters) may not represent longer-context scenarios or multi-turn conversations where toxicity patterns differ

Source documents are from web text with unknown temporal coverage and corpus composition—inherits biases of source websites

Single 'train' split provided; no held-out test set documented, limiting ability to validate generalization

What makes it unique

vs alternatives

More representative of real-world web text than synthetic or crowdsourced prompt datasets, and includes source document traceability unlike generic prompt collections.

prompt-continuation pair evaluation for toxicity amplification measurement

Medium confidence

Solves for

Best for

Researchers studying toxicity generation patterns in language models

Teams developing toxicity mitigation techniques that need quantifiable baseline comparisons

Model safety researchers measuring amplification across different model architectures

Requires

Language model for generating test continuations

Toxicity classifier matching the one used to score dataset (undocumented classifier specification)

Ability to parse and compare paired records in Parquet format

Limitations

Continuations are from original web documents, not model-generated—requires external model inference to generate comparable continuations for evaluation

No documentation of how continuations were selected from source documents (random sampling? first N tokens? sentence boundaries?), limiting reproducibility

Toxicity scores for continuations are pre-computed and static; cannot re-score with updated classifiers or alternative toxicity definitions

What makes it unique

vs alternatives

web-sourced text corpus with source document traceability

Medium confidence

Solves for

Best for

Researchers analyzing dataset bias and source composition effects

Teams validating dataset quality and extraction methodology

Researchers studying toxicity patterns across different web sources

Requires

Access to original source documents (not provided with dataset)

Ability to parse character offsets and reconstruct context from source files

Python 3.6+ with datasets library

Limitations

Source documents themselves are not provided—only filename identifiers and character offsets; original documents must be obtained separately

No documentation of which websites or corpora are included; cannot assess source diversity or bias a priori

Temporal coverage of web text extraction is undocumented; unknown if sources are current or historical

What makes it unique

vs alternatives

More transparent than datasets that strip source information, enabling bias analysis and reproducibility verification that are impossible with anonymized alternatives.

challenging prompt subset selection via boolean flag

Medium confidence

Solves for

Best for

Researchers conducting difficulty-stratified evaluation of model toxicity

Teams testing robustness of toxicity mitigation on harder cases

Requires

Understanding of what 'challenging' means in context (requires external research or reverse-engineering from data)

Python 3.6+ with datasets library for filtering

Limitations

Selection criteria for 'challenging' flag is completely undocumented—unknown what makes a prompt challenging (toxicity level? ambiguity? adversarial intent? length?)

No distribution statistics provided (how many records are marked challenging?)

Cannot validate whether challenging prompts actually represent harder evaluation cases without external analysis

What makes it unique

Includes a boolean 'challenging' flag for subset selection, but the selection criteria and purpose are completely undocumented, making this feature opaque and difficult to use effectively.

vs alternatives

Provides optional difficulty stratification unlike flat prompt datasets, but lacks documentation that makes the feature practically useful.

hugging face datasets api integration for standardized access

Medium confidence

Solves for

Best for

Python developers using Hugging Face ecosystem tools

Researchers building evaluation pipelines with transformers library

Teams already using Hugging Face Hub for model hosting and evaluation

Requires

Python 3.6+

Hugging Face datasets library (pip install datasets)

Internet connection for initial download (unless pre-cached)

Limitations

Requires Hugging Face datasets library installation and Python 3.6+ environment

Streaming mode requires stable internet connection; offline access requires pre-download

No API endpoints or REST access; Python library is only access method

What makes it unique

vs alternatives

More convenient than raw Parquet files or custom data loaders; enables one-line loading and automatic caching unlike manual download approaches.

toxicity-based model evaluation benchmarking

Medium confidence

Solves for

Best for

Model developers evaluating safety of new model versions

Researchers comparing toxicity across model architectures or training approaches

Teams implementing model safety benchmarks and leaderboards

Requires

Language model capable of generating completions

Toxicity classifier (external; specification undocumented)

Computational resources for 99.4k inference passes

Limitations

Requires external toxicity classifier to score model outputs—dataset does not provide classifier; must match or replicate the undocumented classifier used for dataset scores

Benchmarking requires running inference on 99.4k prompts, which is computationally expensive for large models

Toxicity scores are relative to classifier quality; classifier biases or errors propagate to benchmark results

What makes it unique

vs alternatives

More systematic than ad-hoc toxicity evaluation; enables direct comparison across models using identical prompts and scoring methodology, unlike custom evaluation approaches.

Capabilities are decomposed by AI analysis. Each maps to specific user intents and improves with match feedback.

Alternatives to RealToxicityPrompts

cua53Agent

Open-source infrastructure for Computer-Use Agents. Sandboxes, SDKs, and benchmarks to train and evaluate AI agents that can control full desktops (macOS, Linux, Windows).

Compare →

Hugging Face43Platform

The GitHub for AI — 500K+ models, datasets, Spaces, Inference API, hub for open-source AI.

Compare →

Stable-Diffusion55Repository

Compare →

YOLOv846Model

Real-time object detection, segmentation, and pose.

Compare →

RealToxicityPrompts

Capabilities7 decomposed

multi-dimensional toxicity scoring of text prompts and continuations

sentence-level prompt corpus for language model evaluation

prompt-continuation pair evaluation for toxicity amplification measurement

web-sourced text corpus with source document traceability

challenging prompt subset selection via boolean flag

hugging face datasets api integration for standardized access

toxicity-based model evaluation benchmarking

Related Artifactssharing capabilities

PromptPerfect

BetterPrompt

llm-guard

HELM

TrustLLM

GPT Prompt Engineer

Best For

Known Limitations

Requirements

Input / Output

UnfragileRank

About

Categories

Alternatives to RealToxicityPrompts

Are you the builder of RealToxicityPrompts?

Get the weekly brief

Data Sources

RealToxicityPrompts

Capabilities7 decomposed

multi-dimensional toxicity scoring of text prompts and continuations

sentence-level prompt corpus for language model evaluation

prompt-continuation pair evaluation for toxicity amplification measurement

web-sourced text corpus with source document traceability

challenging prompt subset selection via boolean flag

hugging face datasets api integration for standardized access

toxicity-based model evaluation benchmarking

Related Artifactssharing capabilities

PromptPerfect

BetterPrompt

llm-guard

HELM

TrustLLM

GPT Prompt Engineer

Best For

Known Limitations

Requirements

Input / Output

UnfragileRank

About

Categories

Alternatives to RealToxicityPrompts

Are you the builder of RealToxicityPrompts?

Get the weekly brief

Data Sources