Classified Adverts Collection (doi:10.48788/DVUA/BM3ACV)
(Ukrainian Classified Advertisements Dataset)

View:

Part 1: Document Description
Part 2: Study Description
Part 5: Other Study-Related Materials
Entire Codebook

Document Description

Citation

Title:

Classified Adverts Collection

Identification Number:

doi:10.48788/DVUA/BM3ACV

Distributor:

DataverseUA

Date of Distribution:

2026-08-25

Version:

1

Bibliographic Citation:

Zharkov, Dmytro, 2026, "Classified Adverts Collection", https://doi.org/10.48788/DVUA/BM3ACV, DataverseUA, V1

Study Description

Citation

Title:

Classified Adverts Collection

Subtitle:

Dataset for Large-Scale Text Classification of Ukrainian Classified Advertisements

Alternative Title:

Ukrainian Classified Advertisements Dataset

Identification Number:

doi:10.48788/DVUA/BM3ACV

Authoring Entity:

Zharkov, Dmytro (V.M. Glushkov Institute of Cybernetics of the NAS of Ukraine)

Producer:

V.M. Glushkov Institute of Cybernetics of the NAS of Ukraine

Distributor:

DataverseUA

Access Authority:

Zharkov, Dmytro

Depositor:

Zharkov, Dmytro

Date of Deposit:

2026-04-24

Holdings Information:

https://doi.org/10.48788/DVUA/BM3ACV

Study Scope

Keywords:

Computer and Information Science, text classification, machine learning, Ukrainian language, NLP, document classification, large language models

Abstract:

The dataset contains 355 classified advertisements organized into 15 semantic categories and represented as structured JSON objects for supervised multi-class text classification. Each advertisement includes a unique identifier, category identifier and title, advertisement title, full advertisement text, and an LLM-assisted summary. The accompanying category data provide category identifiers and titles together with category-level bag-of-words (BOW) and TF-IDF representations derived from the advertisement corpus. The corpus consists predominantly of Ukrainian-language advertisements and includes naturally occurring mixed Ukrainian–Russian content. The texts preserve characteristics of real-world advertisements, including spelling variations, colloquial language, repetitions, commercial information, and stylistic variability. The dataset covers multiple thematic domains, including furniture, commercial premises and rentals, cosmetics, perfumery, healthcare and beauty products and services, medical products, and equipment. The dataset is intended for research and educational purposes and can be used for supervised text classification, evaluation and benchmarking of machine learning and large language model (LLM)-based classifiers, natural language processing research, feature engineering, and comparative evaluation of text classification methods.

Methodology and Processing

Sources Statement

Data Sources:

Classified advertisement texts organized into semantic categories and represented as structured JSON data for text classification and natural language processing research.

Characteristics of Source Notes:

The corpus comprises 355 classified advertisement records distributed across 15 semantic categories. The textual content is predominantly Ukrainian and includes mixed Ukrainian–Russian language features characteristic of real-world user-generated advertisements. Records preserve spelling variations, colloquial expressions, commercial information, repetitions, and stylistic variability. Each advertisement is associated with a single category label.

Documentation and Access to Sources:

The dataset is accompanied by a README file describing the dataset purpose, JSON file structure, advertisement and category objects, category labeling, dataset characteristics, and recommended uses. The dataset includes adverts.json with 355 structured advertisement records and categories.json with 15 category definitions and corresponding bag-of-words (BOW) and TF-IDF representations.

Data Access

Other Study Description Materials

Other Study-Related Materials

Label:

adverts.json

Text:

A structured collection of advertisements

Notes:

application/json

Other Study-Related Materials

Label:

categories.json

Text:

A set of target categories enriched with collections of characteristic keywords specific to each class

Notes:

application/json

Other Study-Related Materials

Label:

README.txt

Text:

Documentation describing the dataset.

Notes:

text/plain