Ran the CSV through our deduper first. Categorisation was granular enough to segment by sub-vertical without regex. Genuinely useful.
Appreciate you flagging what worked, we'll keep tightening the column coverage in the next refresh. Reach out if you'd like an early look at the next refresh.