Firefly
“Years of scraped Pakistani court judgments sit in messy databases. Make them something a legal product can actually use.”
23steps, run from one config
955judgments in the first full volume
39fields per case
The situation
At Firefly I work on Pakistani case law. The raw material is years of scraped judgments: long blocks of text, inconsistent names, courts that are really section labels, and fields that are simply empty.
The job: turn that into clean, structured data, without ever losing or hiding anything.
What I shipped
- One pipeline instead of a pile of notebooks. Twenty-three steps, run from a single config file, that pull out judges, parties, outcomes, case types, the laws cited and the courts.
- A local LLM where rules can't decide. Rules first; a model only where the text genuinely needs reading. Then every value is cleaned into one standard form.
- Nothing lost or hidden. The raw data is never touched. Every step works on its own copy and writes a before/after report of what it changed, filled or emptied.
- A hand-check loop. The calls no code can make go to a person once. The answers are saved and applied on every run after that.
Stack
Next
→