Skip to content

Repository files navigation

PhishGuard AI — Email Phishing Detection System

Flask web application that classifies email text as phishing or safe using a TF-IDF + Logistic Regression pipeline, with a rule-based layer that corrects for false positives on newsletters and platform notifications. Includes a Gmail integration that scans a connected inbox concurrently and scores messages in real time.

Model

  • Pipeline: TF-IDF (40,000 features, unigrams + bigrams, sublinear TF) → Logistic Regression (class_weight="balanced")
  • Trained on 18,096 labeled emails (14,476 train / 3,620 test, 80/20 stratified split)
  • Test accuracy: 98.26% | ROC-AUC: 0.9969
  • Metrics are computed at training time and stored in model/metadata.pkl; the dashboard reads them directly rather than hardcoding values.

Hybrid rules layer

The raw ML probability is adjusted by adjust_threat_with_rules() in app.py: emails from a whitelisted set of sender domains (GitHub, Google, Microsoft, etc.) or matching known digest/newsletter phrasing are capped at 12% probability, unless the body also contains an urgency trigger ("verify your account," "suspend," "compromised," etc.), in which case the ML score is used unmodified. Classification threshold is 70% probability.

Gmail scanner

Connects via Google OAuth2 (read-only Gmail scope), fetches the 50 most recent inbox messages, and scores them concurrently with a 20-worker ThreadPoolExecutor — each worker instantiates its own isolated credentials/service object to stay thread-safe. Results link back to the live message in Gmail.

Setup

1. Install dependencies

pip install -r requirements.txt

2. Google OAuth (required only for the Gmail scanner)

  1. Create a project in the Google Cloud Console and enable the Gmail API.
  2. Configure the OAuth consent screen (External, scope .../auth/gmail.readonly) and add your test Gmail account under Test Users.
  3. Create an OAuth Client ID (Web application) with redirect URI http://127.0.0.1:5001/gmail/callback.
  4. Download the client JSON, save it as credentials.json in the repo root (git-ignored — not included in this repo), or set its contents as the GCP_CREDENTIALS_JSON environment variable for deployment.

3. Environment variables

  • FLASK_SECRET_KEY — session signing key. Falls back to a random key per process if unset, which invalidates sessions on restart; set explicitly for anything beyond local testing.
  • GCP_CREDENTIALS_JSON — optional, alternative to a local credentials.json (used in the Vercel deployment).

4. Train the model (optional)

The trained pipeline is already checked into model/. To retrain from scratch, place Phishing_Email.csv (email text + label columns) in the repo root and run:

python train_model.py

This overwrites model/phishing_pipeline.pkl and model/metadata.pkl. The CSV itself isn't included in this repo (~50 MB).

5. Run

python app.py

Serves at http://127.0.0.1:5001.

Folder structure

phish-guard-ai/
├── app.py                # Flask app: routes, Gmail OAuth/scanning, chart generation, rules layer
├── train_model.py        # Training pipeline: load CSV → clean → TF-IDF → Logistic Regression → serialize
├── text_utils.py         # Shared text-cleaning functions used by both app.py and train_model.py
├── model/
│   ├── phishing_pipeline.pkl   # Fitted TF-IDF + Logistic Regression pipeline
│   └── metadata.pkl            # Accuracy, ROC-AUC, sample counts, vocab size
├── templates/             # Jinja2 templates (onboarding, dashboard, analyzer, results, Gmail scanner, settings)
├── static/css/            # Stylesheet
├── screenshots/           # UI screenshots referenced below
├── requirements.txt
└── vercel.json             # Vercel deployment config

Screenshots

Onboarding Dashboard
Onboarding Dashboard
Text analyzer Gmail scanner
Analyzer Gmail scanner

Known limitations

  • The rule-based whitelist (trusted sender keywords, digest phrases) is a fixed list, not learned — it needs manual extension for new platforms.
  • OAUTHLIB_INSECURE_TRANSPORT is set unconditionally to allow local HTTP OAuth callbacks; this is fine for 127.0.0.1 testing but should not be relied on for anything internet-facing.
  • No persistent storage: scan results live only in the Flask session and are lost on restart.

About

Resources

Stars

16 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages