Classification Confusion Matrix, Precision, Recall & F1-Score Evaluator

Evaluate binary and multiclass machine learning classification models. Input True Positives (TP), False Positives (FP), True Negatives (TN), and False Negatives (FN) to compute Precision, Recall, Specificity, F1-Score, Balanced Accuracy, and Matthews Correlation Coefficient (MCC).Documentation & FAQs ↓

Classification Confusion Matrix, Precision, Recall & F1-Score Evaluator — Interactive Console
Runs locally in your browser • Instant output
Real-World Model Presets
2x2 Confusion Matrix Heatmap
Actual \ Pred
Predicted Positive
Predicted Negative
Actual Positive
True Positive (TP)28.3%
False Negative (FN)6.7% (Miss)
Actual Negative
False Positive (FP)5.0% (False Alarm)
True Negative (TN)60.0%
Actual Positives: 105 | Negatives: 195
Total Population: 300
Essential Machine Learning MetricsReal-time Calculation
Overall Accuracy88.3%
(TP + TN) / Total
Precision (PPV)85.0%
TP / (TP + FP)
Recall / Sensitivity81.0%
TP / (TP + FN)
Specificity (TNR)92.3%
TN / (TN + FP)
F1-Score0.829
Harmonic Mean of Prec & Recall
Balanced Accuracy86.6%
(Sensitivity + Specificity) / 2
Matthews Corr (MCC)0.741
Balanced metric (-1.0 to +1.0)
False Alarm Rate (FPR)7.7%
FP / (FP + TN)
Advertisement

Step-by-Step Instructions

4 Steps to Completion
1Phase 1

Enter your model's classification counts: True Positives (TP), False Positives (FP), True Negatives (TN), and False Negatives (FN).

2Phase 2

Optionally adjust the Beta weighting factor if your domain prioritizes recall (Beta > 1) or precision (Beta < 1).

3Phase 3

Review the real-time metrics dashboard displaying Precision, Recall, F1, Specificity, and MCC.

4Phase 4

Export the diagnostic summary report or copy metrics formatted for LaTeX tables and technical documentation.

Practical Use Cases & Real-World Scenarios

Medical Diagnostics & Healthcare AI
Balance high Sensitivity/Recall to ensure life-critical conditions are never missed while tracking Specificity to minimize unnecessary invasive testing.
Financial Fraud Detection Engineering
Fine-tune F-beta scores where False Negatives (undetected fraud) cost drastically more than False Positives (flagged card alerts).
Spam & Content Moderation Classifier Tuning
Maximize Precision to ensure legitimate user emails and creator posts are never incorrectly quarantined by automated moderation algorithms.
Related Editorial Guide on AnshuTechy11 Best Machine Learning Tools for Model Training in 2026
Read Tutorial →
Share With Your Community

Found this tool helpful? Share it with colleagues & friends:

100% free, private in-browser utility with zero server uploads.