Fulltext available Open Access
DC FieldValueLanguage
dc.contributor.advisorPutzar, Larissa-
dc.contributor.authorVan, Anna-
dc.date.accessioned2026-08-04T08:39:50Z-
dc.date.available2026-08-04T08:39:50Z-
dc.date.issued2025-09-09-
dc.identifier.urihttps://hdl.handle.net/20.500.12738/19704-
dc.description.abstractMachine learning and transformers have become increasingly involved in modern applications, particularly for text generation and classification. This work presents the current state of art text classifiers with transformers, with a focus on email spam detection. It will evaluate and compare the performance across a selection of BERT-Transformer models: The base model BERT and specialized models RoBERTa and DistilBERT. Each of them is trained on datasets Ling-Spam, SpamAssassin and Enron-Spam as well as on a combined dataset, to examine how dataset and model architecture influence the classification performance. The study highlights the differences in efficiency, generalization and stability among the models, showing how lighter models like DistilBERT excel on smaller datasets while BERT benefits from larger, more complex data. It further underscores how model choice should consider dataset size, complexity and deployment priorities. These results offer guidance for selecting and implementing spam filters and emphasize the trade-offs between cost and performance priorities.en
dc.description.abstractMaschinelles Lernen und Transformer haben in modernen Anwendungen zunehmen an Bedeutung gewonnen, insbesondere im Bereich der Textgenerierung und -klassifikation. Diese Arbeite befasst sich mit dem aktuellen Stand der Technik von Textklassifikatoren, mit einem Fokus auf Transformern und der Erkennung von E-Mail-Spam. Dabei werden eine Auswahl von BERT-basierten Transformermodelle untersucht und verglichen: das Basismodell BERT und die spezialisierten Varianten RoBERTa und DistilBERT. Jedes dieser Modelle wird auf den Datensätzen Ling-Spam, SpamAssassin und Enron-Spam sowie einem kombinierten Datensatz trainiert, um zu analysieren, wie Datensatz und Modellarchitektur die Leistung bei der Klassifizierung beeinflussen. Die Studie hebt die Unterschiede in Effizienz, Generalisierungsfähigkeit und Stabilität der Modelle hervor und zeigt, das leichtere Modelle wie DistilBERT auf kleineren Datensätzen bessere Leistungen erbringen, während BERT bei größeren und komplexeren Daten überzeugt. Dadurch wird unterstrichen, dass die Modellwahl die Größe und Komplexität des Datensatzes sowie die Anforderungen an den Einsatz berücksichtigt werden muss. Die Ergebnisse liefern Orientierung bei der Auswahl und Implementierung von Spamfiltern und verdeutlichen die Abwägung zwischen Kosten- und Leistungsprioritäten.de
dc.language.isoenen_US
dc.subject.ddc004: Informatiken_US
dc.titleBERT-based spam email classification - exploring model effectiveness across different datasetsen
dc.typeThesisen_US
openaire.rightsinfo:eu-repo/semantics/openAccessen_US
thesis.grantor.departmentFakultät Design, Medien und Information (ehemalig, aufgelöst 10.2025)en_US
thesis.grantor.departmentDepartment Medientechnik (ehemalig, aufgelöst 10.2025)en_US
thesis.grantor.universityOrInstitutionHochschule für Angewandte Wissenschaften Hamburgen_US
tuhh.contributor.refereeSchumann, Sabine-
tuhh.identifier.urnurn:nbn:de:gbv:18302-reposit-242766-
tuhh.oai.showtrueen_US
tuhh.publication.instituteFakultät Design, Medien und Information (ehemalig, aufgelöst 10.2025)en_US
tuhh.publication.instituteDepartment Medientechnik (ehemalig, aufgelöst 10.2025)en_US
tuhh.type.opusBachelor Thesis-
dc.type.casraiSupervised Student Publication-
dc.type.dinibachelorThesis-
dc.type.driverbachelorThesis-
dc.type.statusinfo:eu-repo/semantics/publishedVersionen_US
dc.type.thesisbachelorThesisen_US
dcterms.DCMITypeText-
tuhh.dnb.statusdomainen_US
item.advisorGNDPutzar, Larissa-
item.openairetypeThesis-
item.languageiso639-1en-
item.fulltextWith Fulltext-
item.grantfulltextopen-
item.cerifentitytypePublications-
item.creatorOrcidVan, Anna-
item.creatorGNDVan, Anna-
item.openairecristypehttp://purl.org/coar/resource_type/c_46ec-
Appears in Collections:Theses
Files in This Item:
File Description SizeFormat
BA_BERT-based_spam_email_classification.pdf1.23 MBAdobe PDFView/Open
Show simple item record

Google ScholarTM

Check

HAW Katalog

Check

Note about this record


Items in REPOSIT are protected by copyright, with all rights reserved, unless otherwise indicated.