Ran the CSV through our deduper first. Categorisation was granular enough to segment by sub-vertical without regex. Couple of small gaps but workable.
Thanks for the detailed review, the points you raised are on the refresh checklist for the next pull. Open a ticket via the Support tab and we'll prioritise the columns you flagged for the next refresh.