Data collection

Public-data scraper

A program that collects public lists from portals and files and combines them into one dataset. Source records are retained; questionable addresses and duplicates need review.

← Back to projects

Project scope

Before cleaning the data, you have to reach it.

A public website does not necessarily make its data easy to collect. This project started with finding and reaching sources: when one collection method failed, another had to be tried. Only then could the work on column names, addresses and duplicate records begin.

The Python workflow discovered and fetched public rosters, retained source snapshots and read CSV, XLSX and JSON formats. Different column names were mapped to a common structure, records validated and duplicates checked before projection into a usable table. A review queue and CSV/Markdown output generation were also implemented. Generating an output does not mean automatically sending emails.

How the work happens

One failed attempt need not end the search.

6 steps
WorkflowWhen something needs attention
  1. 01Find a public source

    A portal, API or downloadable roster.

  2. 02Retrieve and retain

    Try a collection route. Keep the source snapshot.

  3. 03Parse and map fields

    Map CSV, XLSX or JSON into a shared structure.

  4. 04Check the addresses

    Separate address parts and validate values.

  5. 05Check duplicates and restrictions

    Check repeated records and sensitive fields.

  6. 06Produce a usable output

    Checked records feed a directory or export.

Simplified view of alternative collection methods. Missing information is not invented.

What did the workflow do?

  1. Find the source and try a collection route

    Data came from public portals, APIs and files. A failed collection method did not end the whole workflow: alternative methods were available. Information that could not be retrieved was not something to replace with guesses. Retrieved source snapshots were retained so the result could later be checked against the original.

  2. Read the format

    Structured extractors separate fields. Different source column names map into a shared structure.

  3. Validate and normalise

    Address components are separated, values checked and duplicates identified. A matching column name does not guarantee matching content.

  4. Review exceptions

    Suspicious records enter review. AI can assist, but corrections and quality checks must precede promotion of the data.

  5. Produce a usable output

    The checked structure can feed a directory or an export of selected fields. A public source does not automatically make every subsequent use of personal data appropriate.

What failed, and what changed?

Joining and splitting text corrupted addresses.

One source put state and ZIP into the city field; another put a suite number there. The parser was corrected, regression tests added and affected records re-imported. Fixing code alone was not enough.

A repeated import must not silently add duplicate rows.

A replayed import introduced duplicates. The dataset was rebuilt. A procedural limitation remained: the duplicate guard depended on the projected table, which therefore had to be rebuilt before re-import.

Privacy rules also need exception review.

For providers marked as residential, the program hid street addresses using the recorded provider type. A new or unexpected type description could evade that check; this is not a guarantee that every sensitive field is always detected.

My role and the AI contribution

I set the data’s purpose, target definitions and budget constraints. AI agents I directed helped research sources, write extractors, author tests and implement corrections. Review findings had to become fixes before data was promoted.

Where did the data go?

The database and files were local. Public sources were fetched externally, and AI-assisted review used an external model service. Public data and local storage do not automatically mean unrestricted use or entirely local processing.

CSV · XLSX · JSONDifferent formats, one shared structure
The useful part is not the file count. It is being able to compare the same fields across sources and trace a record back to its origin.

Sources: project source review and my operating experience. Examples are illustrative, not customer data.

The problem

Public sources arrive in varied formats. A column name can match while the content does not: a matching column does not mean matching data.

What was built

  • A normalisation pipeline: fetching the sources, converting them to a shared shape, validation, duplicate checks and human review before use.
  • The review also leans on an external model. That dependency is deliberately described only at a high level and no details are published here.

My role and the AI role

My role

I defined the scope and the quality requirements, and I did the final review.

AI role

AI assisted with the fetcher, the normalisation and the tests.

A selected lesson

An address parsing bug taught that real variations need regression cases. And one precise nuance: the duplicate import recovery and process rule is not permanent idempotency. It is a recovery rule for one situation, not a guarantee that the same import will always be safe to repeat.

Next step

Want to build something similar?

I can help you build with AI while you keep control of the result.

Get in touch