Classified Adverts Collection
The dataset contains 355 classified advertisements organized into 15 semantic categories and represented as structured JSON objects for supervised multi-class text classification. Each advertisement includes a unique identifier, category identifier and title, advertisement title, full advertisement...
Gespeichert in:
| Datum: | 2026 |
|---|---|
| 1. Verfasser: | |
| Sprache: | Ukrainisch |
| Veröffentlicht: |
DataverseUA
2026
|
| Schlagworte: | |
| Online Zugang: | https://doi.org/10.48788/DVUA/BM3ACV |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| Назва журналу: | Open Data Repository of the National Academy of Sciences of Ukraine |
Institution
Open Data Repository of the National Academy of Sciences of Ukraine| _version_ | 1874636220838969345 |
|---|---|
| author | Zharkov, Dmytro |
| author2 | Zharkov, Dmytro |
| author_facet | Zharkov, Dmytro Zharkov, Dmytro |
| author_sort | Zharkov, Dmytro |
| collection | DSpace |
| description | The dataset contains 355 classified advertisements organized into 15 semantic categories and represented as structured JSON objects for supervised multi-class text classification. Each advertisement includes a unique identifier, category identifier and title, advertisement title, full advertisement text, and an LLM-assisted summary. The accompanying category data provide category identifiers and titles together with category-level bag-of-words (BOW) and TF-IDF representations derived from the advertisement corpus.
The corpus consists predominantly of Ukrainian-language advertisements and includes naturally occurring mixed Ukrainian–Russian content. The texts preserve characteristics of real-world advertisements, including spelling variations, colloquial language, repetitions, commercial information, and stylistic variability. The dataset covers multiple thematic domains, including furniture, commercial premises and rentals, cosmetics, perfumery, healthcare and beauty products and services, medical products, and equipment.
The dataset is intended for research and educational purposes and can be used for supervised text classification, evaluation and benchmarking of machine learning and large language model (LLM)-based classifiers, natural language processing research, feature engineering, and comparative evaluation of text classification methods. |
| first_indexed | 2026-08-27T01:00:16Z |
| id | doi-10-48788-DVUA-BM3ACV |
| institution | Open Data Repository of the National Academy of Sciences of Ukraine |
| language | Ukrainian |
| last_indexed | 2026-08-27T01:00:16Z |
| publishDate | 2026 |
| publisher | DataverseUA |
| record_format | dspace |
| spelling | doi-10-48788-DVUA-BM3ACV2026-08-26T02:00:01ZClassified Adverts Collectionhttps://doi.org/10.48788/DVUA/BM3ACVZharkov, DmytroDataverseUAThe dataset contains 355 classified advertisements organized into 15 semantic categories and represented as structured JSON objects for supervised multi-class text classification. Each advertisement includes a unique identifier, category identifier and title, advertisement title, full advertisement text, and an LLM-assisted summary. The accompanying category data provide category identifiers and titles together with category-level bag-of-words (BOW) and TF-IDF representations derived from the advertisement corpus. The corpus consists predominantly of Ukrainian-language advertisements and includes naturally occurring mixed Ukrainian–Russian content. The texts preserve characteristics of real-world advertisements, including spelling variations, colloquial language, repetitions, commercial information, and stylistic variability. The dataset covers multiple thematic domains, including furniture, commercial premises and rentals, cosmetics, perfumery, healthcare and beauty products and services, medical products, and equipment. The dataset is intended for research and educational purposes and can be used for supervised text classification, evaluation and benchmarking of machine learning and large language model (LLM)-based classifiers, natural language processing research, feature engineering, and comparative evaluation of text classification methods.Computer and Information Sciencetext classificationmachine learningUkrainian languageNLPdocument classificationlarge language modelsuk2026-08-25Zharkov, DmytroClassified advertisement texts organized into semantic categories and represented as structured JSON data for text classification and natural language processing research. |
| spellingShingle | Computer and Information Science text classification machine learning Ukrainian language NLP document classification large language models Zharkov, Dmytro Classified Adverts Collection |
| title | Classified Adverts Collection |
| title_full | Classified Adverts Collection |
| title_fullStr | Classified Adverts Collection |
| title_full_unstemmed | Classified Adverts Collection |
| title_short | Classified Adverts Collection |
| title_sort | classified adverts collection |
| topic | Computer and Information Science text classification machine learning Ukrainian language NLP document classification large language models |
| url | https://doi.org/10.48788/DVUA/BM3ACV |
| work_keys_str_mv | AT zharkovdmytro classifiedadvertscollection |