Joining and splitting text corrupted addresses.
One source put state and ZIP into the city field; another put a suite number there. The parser was corrected, regression tests added and affected records re-imported. Fixing code alone was not enough.
A program that collects public lists from portals and files and combines them into one dataset. Source records are retained; questionable addresses and duplicates need review.
A public website does not necessarily make its data easy to collect. This project started with finding and reaching sources: when one collection method failed, another had to be tried. Only then could the work on column names, addresses and duplicate records begin.
The Python workflow discovered and fetched public rosters, retained source snapshots and read CSV, XLSX and JSON formats. Different column names were mapped to a common structure, records validated and duplicates checked before projection into a usable table. A review queue and CSV/Markdown output generation were also implemented. Generating an output does not mean automatically sending emails.
A portal, API or downloadable roster.
Try a collection route. Keep the source snapshot.
Map CSV, XLSX or JSON into a shared structure.
Separate address parts and validate values.
Check repeated records and sensitive fields.
Checked records feed a directory or export.
Data came from public portals, APIs and files. A failed collection method did not end the whole workflow: alternative methods were available. Information that could not be retrieved was not something to replace with guesses. Retrieved source snapshots were retained so the result could later be checked against the original.
Structured extractors separate fields. Different source column names map into a shared structure.
Address components are separated, values checked and duplicates identified. A matching column name does not guarantee matching content.
Suspicious records enter review. AI can assist, but corrections and quality checks must precede promotion of the data.
The checked structure can feed a directory or an export of selected fields. A public source does not automatically make every subsequent use of personal data appropriate.
One source put state and ZIP into the city field; another put a suite number there. The parser was corrected, regression tests added and affected records re-imported. Fixing code alone was not enough.
A replayed import introduced duplicates. The dataset was rebuilt. A procedural limitation remained: the duplicate guard depended on the projected table, which therefore had to be rebuilt before re-import.
For providers marked as residential, the program hid street addresses using the recorded provider type. A new or unexpected type description could evade that check; this is not a guarantee that every sensitive field is always detected.
I set the data’s purpose, target definitions and budget constraints. AI agents I directed helped research sources, write extractors, author tests and implement corrections. Review findings had to become fixes before data was promoted.
The database and files were local. Public sources were fetched externally, and AI-assisted review used an external model service. Public data and local storage do not automatically mean unrestricted use or entirely local processing.
Sources: project source review and my operating experience. Examples are illustrative, not customer data.
Public sources arrive in varied formats. A column name can match while the content does not: a matching column does not mean matching data.
I defined the scope and the quality requirements, and I did the final review.
AI assisted with the fetcher, the normalisation and the tests.
An address parsing bug taught that real variations need regression cases. And one precise nuance: the duplicate import recovery and process rule is not permanent idempotency. It is a recovery rule for one situation, not a guarantee that the same import will always be safe to repeat.
I can help you build with AI while you keep control of the result.