×

img Accessibility Controls

Research Projects Banner

Research Projects

Scalable and Compliant Data Platforms for Federated and Decentralized Data Processing

Implementing Organization

Indian Institute Of Technology Delhi
Principal Investigator
Dr. Kaustubh Beedkar
Indian Institute Of Technology Delhi
kbeedkar@cse.iitd.ac.in

Project Overview

With data privacy laws like GDPR, HIPAA, CCPA, and India's DDPA, centralized processing of sensitive data faces increasing restrictions. Traditional systems reliant on centralized access no longer suffice for regulated industries. The need to extract insights from data spread across independent sources—while staying compliant—has driven the rise of federated data systems, enabling analysis without moving or directly accessing raw data. This research tackles building an efficient, compliant data management framework for federated environments, balancing utility with privacy and regulations. It focuses on three areas: compliant dataset search in federated repositories, federated query processing, and platforms for decentralized data processing. The project's objectives include: 1) To develop integrated metadata-based and approximate data summarization-based methods dataset search; 2) To design and implement a policy-driven query processing framework; 3) To build a decentralized data computing platform that will focus on interoperability between federated data platforms; and 4) To validate the proposed techniques in real-world scenarios, demonstrating their effectiveness in terms of compliance, efficiency, and accuracy. The core hypothesis of this research is that integrated metadata-based and approximate summary-based techniques, combined with policy-aware query processing, can enable effective and compliant exploration and processing of federated datasets. The project's evaluation will include: 1) Dataset Search: Design and test techniques using metadata and approximate data summaries (e.g., histograms, sketches) for efficient dataset discovery in federated settings. Experiments will: a) evaluate data summaries' effectiveness for search tasks; b) measure method accuracy; and c) assess scalability, memory use, and search time efficiency. 2) Policy-Driven Query Processing: Implement a policy-driven query engine that can enforce regulatory compliance during query processing. Experiments will focus on evaluating data masking functions and the effectiveness of error metric (i.e., minimizing the change in G1, G2, or G3 error) and their impact on downstream tasks. 3) Decentralized Data Platform: Build a decentralized computing prototype with Apache Wayang and run analytics on federated datasets (e.g., healthcare, finance) to assess its scalability and interoperability. Overall, the research aims to advance both theory and practice in federated data processing. Specifically, (a) it will explore integrating approximate data summaries with metadata for compliant dataset search without raw data access; and (b) provide a robust framework for extracting insights from sensitive, distributed datasets. This framework will aid regulated sectors like healthcare, finance, and government, enabling actionable insights while ensuring compliance. It also sets the stage for future decentralized analytics, shifting from centralized to privacy-aware systems.
Funding Organization
Quick Information
Area of Research
Engineering Sciences
Focus Area
Computer Engineering
Start Date
03 Jun 2025
End Date
02 Jun 2028
Status
ongoing
Output
No. of Research Paper
00
Technologies (If Any)
00
No. of PhD Produced
00
Publications
00
No. of Patents
Filed : 00
Grant : 00
arrowtop
Latest Updates
Loading…