Track: Privacy & Open Source AI Tools
β Problem Description:
Every day, people share sensitive documents like pension statements, social security notices, insurance forms, and tax documents with AI assistants, online tools, or third-party services. The moment you upload that invoice to ChatGPT or another cloud tool, your personal data (name, address, ID numbers, diagnoses) is out of your hands forever.
Thereβs no safe way to use AI on private documents today.
β Primary Objective:
Build an open-source tool that lets anyone redact personal information from PDFs locally. Using AI to detect and mask sensitive data, the redacted document can subsequently be safely shared with any AI service without privacy risk.
The twist: Redaction should be reversible. A unique key (stored only in the userβs wallet or device) maps redacted tokens back to real values. The owner can always un-redact their own document.
Super challenge: Also redact images from scanned documents.
β Datasets:
Participants will work with sample PDF documents containing synthetic PII, including mock invoices, salary statements, insurance letters, and tax forms (text-based PDFs). Participants are encouraged to generate additional synthetic test data. No real personal data should be used.
β Available APIs & Services:
Participants may use any open-source NLP/NER library suitable for local execution, for example:
β Supporting Materials:
Sample synthetic test PDFs (hospital invoices, social security notices, etc.)
β Mentorship: Check online session on 4 March 26, 10:00, Gmeet