Abstract

The aim of this paper is to propose a framework addressing Big Data quality aspects in the context of Web data. The identified research gap, based on a review of existing literature, is that Big Data quality is typically evaluated using broad and general dimensions, such as data consistency, without providing detailed methodologies for measurement or specifying metrics suitable for continuous quality monitoring. This paper focuses on challenges identified in case studies involving the use of Web data, particularly in the context of large-scale web scraping, to derive information about enterprise characteristics.

Recommended Citation

Maślankowski, J.(2026). Advancing the Measurement of Big Data Quality in Web Sources. In M. Valenta, B. Mannová, R. Pergl, A. Przybylek, M. Lang, H. Linger, C. Schneider, N. Iivari, & E. Insfran (Eds.), Making ISD Sustainable: Reloaded with AI and Automation (ISD2026 Proceedings). Prague, Czech Republic: Czech Technical University in Prague. ISBN: 978-80-01-07585-2. https://doi.org/10.62036/ISD.2026.187

Paper Type

Poster

DOI

10.62036/ISD.2026.187

Share

COinS
 

Advancing the Measurement of Big Data Quality in Web Sources

The aim of this paper is to propose a framework addressing Big Data quality aspects in the context of Web data. The identified research gap, based on a review of existing literature, is that Big Data quality is typically evaluated using broad and general dimensions, such as data consistency, without providing detailed methodologies for measurement or specifying metrics suitable for continuous quality monitoring. This paper focuses on challenges identified in case studies involving the use of Web data, particularly in the context of large-scale web scraping, to derive information about enterprise characteristics.