Developing Intelligent Systems That Continuously Monitor and Validate Data Quality Across Large Distributed Systems
Main Article Content
Abstract
Ensuring high-quality data in large-scale distributed systems is essential for the reliability of real-time analytics, automated decision-making, and regulatory compliance in data-driven enterprises. Traditional data quality techniques, largely based on static rule-based approaches, are insufficient to address the scale, velocity, and complexity of modern distributed environments. This study presents the design and evaluation of an intelligent data quality monitoring system that integrates rule-based validation, machine learning models, metadata analysis, and adaptive feedback loops. The proposed architecture supports both real-time and batch processing, and was implemented using distributed computing frameworks such as Apache Kafka and Spark. Empirical evaluations conducted using synthetic IoT sensor data and real-world NYC taxi trip records demonstrated that the system outperformed traditional methods in terms of precision, recall, F1 score, and scalability. Furthermore, the system exhibited adaptive capabilities through feedback-driven learning and self-healing mechanisms, enabling it to respond effectively to evolving data patterns. These results confirm the system’s practicality and effectiveness in maintaining trustworthy data within high-volume, dynamic distributed environments. The study concludes with recommendations for future enhancements, including the integration of explainable AI and decentralized validation techniques.

Citation Metrics:
Downloads
Citation Metrics & Similar Scopus Articles
-
Zhang J. (2027)Spectrally Partitioned All-Optical Memristors for Highly Linear Neuromorphic Vision SystemsNano Micro Letters, 19(1)
-
Najafi E. (2027)Analyzing the process of dynamic and tense discourse systems in the poem "katibeh" of Akhavan salesLanguage Related Research, 17(4), 263-291
-
Sheng T. (2027)Photothermal Superhydrophobic Textiles: An Emerging Paradigm from Passive Water Repellency to Active Thermal ManagementNano Micro Letters, 19(1)
Article Details

Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
References
2. Barr, R., Razo Zapata, I., Batarseh, F. A., & Harfoushi, O. (2020). Machine learning for data quality and anomaly detection in data intensive systems. Journal of Big Data, 7, 65.
3. Batini, C., & Scannapieco, M. (2016). Data and Information Quality: Dimensions, Principles and Techniques. Springer.
4. Berti Équille, L. (2015). Quality aware data integration. In T. Sellis & E. Bertino (Eds.), Data Management in Pervasive Systems (pp. 57–83). Springer.
5. Cappiello, C., & Pernici, B. (2006). Quality aware web service composition. In Proceedings of the 2006 IEEE International Conference on Web Services (pp. 211–218). IEEE.
6. Chu, X., Ilyas, I. F., Krishnan, S., & Wang, J. (2016). Data cleaning: Overview and emerging challenges. In Proceedings of the 2016 International Conference on Management of Data (pp. 2201–2206). ACM.
7. Ilyas, I. F., Chu, X., & Ouzzani, M. (2015). Data cleaning: A machine learning perspective. Data Engineering Bulletin, 38(2), 40–47.
8. Jagadish, H. V., Lakshmanan, L. V., Srivastava, D., & Thompson, K. (2014). Managing and mining massive data: From biology to physics. Communications of the ACM, 57(7), 64–73.
9. Müller, H., & Heiler, S. (2018). Enabling data quality monitoring through integrated metadata management. Journal of Information and Data Management, 9(2), 33–45.
10. Müller, H., & Heiler, S. (2018). Enabling data quality monitoring through integrated metadata management. Journal of Information and Data Management, 9(2), 33–45.
11. Pipino, L. L., Lee, Y. W., & Wang, R. Y. (2002). Data quality assessment. Communications of the ACM, 45(4), 211–218.
12. Rekatsinas, T., Chu, X., Ilyas, I. F., & Ré, C. (2017). HoloClean: Holistic data repairs with probabilistic inference. In Proceedings of the 2017 ACM SIGMOD International Conference on Management of Data (pp. 119–134).
13. Wang, R. Y., & Strong, D. M. (1996). Beyond accuracy: What data quality means to data consumers. Journal of Management Information Systems, 12(4), 5–33.
14. Zhu, X., Song, Y., Lin, C., & Yu, Y. (2019). Real time data quality monitoring in distributed data warehouses. IEEE Access, 7, 105672–105684.


















