Ran the CSV through our deduper first. Sample matched live data when I spot-checked random rows. Genuinely useful.