Classified Adverts Collection

The dataset contains 355 classified advertisements organized into 15 semantic categories and represented as structured JSON objects for supervised multi-class text classification. Each advertisement includes a unique identifier, category identifier and title, advertisement title, full advertisement...

Ausführliche Beschreibung

Gespeichert in:
Bibliographische Detailangaben
Datum:2026
1. Verfasser: Zharkov, Dmytro
Sprache:Ukrainisch
Veröffentlicht: DataverseUA 2026
Schlagworte:
Online Zugang:https://doi.org/10.48788/DVUA/BM3ACV
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
Назва журналу:Open Data Repository of the National Academy of Sciences of Ukraine

Institution

Open Data Repository of the National Academy of Sciences of Ukraine
_version_ 1874636220838969345
author Zharkov, Dmytro
author2 Zharkov, Dmytro
author_facet Zharkov, Dmytro
Zharkov, Dmytro
author_sort Zharkov, Dmytro
collection DSpace
description The dataset contains 355 classified advertisements organized into 15 semantic categories and represented as structured JSON objects for supervised multi-class text classification. Each advertisement includes a unique identifier, category identifier and title, advertisement title, full advertisement text, and an LLM-assisted summary. The accompanying category data provide category identifiers and titles together with category-level bag-of-words (BOW) and TF-IDF representations derived from the advertisement corpus. The corpus consists predominantly of Ukrainian-language advertisements and includes naturally occurring mixed Ukrainian–Russian content. The texts preserve characteristics of real-world advertisements, including spelling variations, colloquial language, repetitions, commercial information, and stylistic variability. The dataset covers multiple thematic domains, including furniture, commercial premises and rentals, cosmetics, perfumery, healthcare and beauty products and services, medical products, and equipment. The dataset is intended for research and educational purposes and can be used for supervised text classification, evaluation and benchmarking of machine learning and large language model (LLM)-based classifiers, natural language processing research, feature engineering, and comparative evaluation of text classification methods.
first_indexed 2026-08-27T01:00:16Z
id doi-10-48788-DVUA-BM3ACV
institution Open Data Repository of the National Academy of Sciences of Ukraine
language Ukrainian
last_indexed 2026-08-27T01:00:16Z
publishDate 2026
publisher DataverseUA
record_format dspace
spelling doi-10-48788-DVUA-BM3ACV2026-08-26T02:00:01ZClassified Adverts Collectionhttps://doi.org/10.48788/DVUA/BM3ACVZharkov, DmytroDataverseUAThe dataset contains 355 classified advertisements organized into 15 semantic categories and represented as structured JSON objects for supervised multi-class text classification. Each advertisement includes a unique identifier, category identifier and title, advertisement title, full advertisement text, and an LLM-assisted summary. The accompanying category data provide category identifiers and titles together with category-level bag-of-words (BOW) and TF-IDF representations derived from the advertisement corpus. The corpus consists predominantly of Ukrainian-language advertisements and includes naturally occurring mixed Ukrainian–Russian content. The texts preserve characteristics of real-world advertisements, including spelling variations, colloquial language, repetitions, commercial information, and stylistic variability. The dataset covers multiple thematic domains, including furniture, commercial premises and rentals, cosmetics, perfumery, healthcare and beauty products and services, medical products, and equipment. The dataset is intended for research and educational purposes and can be used for supervised text classification, evaluation and benchmarking of machine learning and large language model (LLM)-based classifiers, natural language processing research, feature engineering, and comparative evaluation of text classification methods.Computer and Information Sciencetext classificationmachine learningUkrainian languageNLPdocument classificationlarge language modelsuk2026-08-25Zharkov, DmytroClassified advertisement texts organized into semantic categories and represented as structured JSON data for text classification and natural language processing research.
spellingShingle Computer and Information Science
text classification
machine learning
Ukrainian language
NLP
document classification
large language models
Zharkov, Dmytro
Classified Adverts Collection
title Classified Adverts Collection
title_full Classified Adverts Collection
title_fullStr Classified Adverts Collection
title_full_unstemmed Classified Adverts Collection
title_short Classified Adverts Collection
title_sort classified adverts collection
topic Computer and Information Science
text classification
machine learning
Ukrainian language
NLP
document classification
large language models
url https://doi.org/10.48788/DVUA/BM3ACV
work_keys_str_mv AT zharkovdmytro classifiedadvertscollection