Catch inappropriate names. Allow harmless ones.
Inappropriate Usernames checks usernames before they reach a public leaderboard. It uses a small, explainable production classifier instead of a runtime LLM: fast enough for modest hardware, reproducible and easy to operate.
The goal is not to block as much as possible. The model should catch genuinely inappropriate names while allowing harmless words such as Essex, Assistant or Fjord.
What is inside
- Multilingual datasets for German, English, Spanish, French, Japanese, Portuguese, Russian, Ukrainian and Chinese.
- Character-based features that cover spelling variants, leetspeak and separators.
- Carefully curated CLEAN and HARD-NEGATIVE sets to keep false positives low.
- A small FastAPI endpoint for single and batch predictions.
- Deterministic builds and documented model provenance.
Under the hood
The pipeline combines raw, compact and canonical username representations. Character n-grams feed a TF-IDF branch and binary canonical features. Logistic Regression then produces a probability; the release threshold is documented and evaluated separately.
The stored reference configuration reaches 99.74% recall, 99.30% precision and a 0.22% false-positive rate on the product test. A separate holdout with external or held-out product names serves as a plausibility and regression test.
Try it and contribute
You can train and evaluate the classifier locally, or run it against your own username list. New hard negatives, reproducible failure cases, language-list reviews and documentation improvements are especially welcome.
git clone https://codeberg.org/scovillo/inappropriate-usernames.git
cd inappropriate-usernames
python build_dataset.py
python train/train_model.py
python evaluate_product_test.py
The original software is licensed under MIT. See NOTICE for curated datasets, external word lists and model provenance.