Model Card for FakeNewsSLM
This model is a specialized Small Language Model (SLM) based on the RoBERTa architecture. It's designed to detect stylistic patterns, unprofessional grammar, and sensationalist emotions often found in fabricated news.
Model Details
Model Description
FakeNewsSLM is highly optimized for text-classification tasks to flag suspicious articles. Rather than acting as a universal fact database, it excels at identifying deceptive emotional language, specifically the tone of angry internet blogs and sensationalized reporting.
- Developed by: Kunalv
- Model type: Small Language Model
- Language: English
- License: Apache-2.0
- Base Model: roberta-base
Uses
Direct Use
Researchers and developers can utilize this model to automatically classify news text based on writing style and emotional tone. It's perfectly designed for integration into content moderation pipelines to catch highly emotional misinformation.
Downstream Use
It can be integrated into journalism verification systems to flag suspicious articles for human review before publication.
Out-of-Scope Use
This model shouldn't be used as an absolute truth verification system. It doesn't possess a real-time fact database and cannot verify highly nuanced satire or sophisticated professional disinformation. If a fabricated article uses perfect professional journalism grammar, the model will likely classify it as real news.
Bias, Risks, and Limitations
Because the model was fine-tuned specifically on the ISOT dataset, it relies entirely on detecting semantic structures and stylistic cues rather than verifying objective facts. It's highly effective at catching angry internet blogs but can be easily tricked by fabricated text that mimics professional journalism perfectly.
Recommendations
Users should be made aware of the limitations. Always use this model as an assistive tool for human moderators rather than an automated decision-maker. Future upgrades may require a Retrieval-Augmented Generation (RAG) pipeline to compare claims against hard evidence.
How to Get Started with the Model
Use the code below to get started with the model.
from transformers import pipeline
classifier = pipeline(model="Kunalv/FakeNewsSLM")
result = classifier("Your news text goes here")
print(result)
Training Details
Training Data
The model was fine-tuned on the ISOT Fake News dataset. The true news data was rigorously pre-processed utilizing regular expressions to strip publisher prefixes like "Reuters", forcing the model to learn genuine linguistic patterns rather than memorizing publisher names.
Training Procedure
Training Hyperparameters
The model was trained on dual NVIDIA T4 GPUs for exactly 2 epochs utilizing a maximum tokenization length of 512. This extended length allows the model to catch fake news indicators hidden deep within the article body.
Evaluation
Testing Data, Factors and Metrics
Testing Data
The evaluation was conducted on a strictly isolated vault of over 4,000 unseen validation articles from the ISOT dataset.
Metrics
Accuracy was utilized to measure the model's ability to correctly classify the semantic and grammatical structures of the text.
Results
The model achieved 100% validation accuracy on the isolated test set. This high score is attributed to the vast semantic and grammatical differences between professional journalism and sensationalized fake news within the specific dataset.
- Downloads last month
- 31
Model tree for Kunalv/FakeNewsSLM
Base model
FacebookAI/roberta-base