DMA15

Direct marketing, 1872 to the cookie

Lists & Segments

The Merge-Purge Problem

When a mailer rented multiple lists for a single campaign, the same name appeared on several of them — the merge-purge operation deduplicated the combined file and, in doing so, revealed how much overlap the list industry was selling.

More in Lists & Segments

A specimen United States Postal Money Order form with a detachable purchaser's receipt stub

Deduplicating rented lists showed mailers how much overlap the list trade had been selling them.Photo: Printed sample of USPS postal order on a punched card · Wikimedia Commons

When Deduplication Exposed the Industry's Secrets

Renting a single mailing list was straightforward. Renting six for the same campaign created a problem that the list industry spent years pretending was manageable: the same person's name appeared on multiple lists, each rented at full price, each promising distinct prospects.

The solution was the merge-purge, a data operation in which all rented files were combined into one master file and duplicate records removed before the lettershop ran a single piece. In mechanical form, before computing made the step routine, the process meant physically sorting address cards and pulling matches by hand — expensive, slow, and imprecise. When mainframe processing became accessible to large mailers through the 1960s and 1970s, the merge-purge became a standard pre-campaign step, and its results were often uncomfortable reading for the list brokers involved.

A mailing-list card index open in its drawer, an adult's hand spread across the cards

Before the file there was the drawer: the broker's card index, rented by the thousand names.

Photo: Tima Miroshnichenko / Pexels

The discomfort was quantitative. A mailer who rented 200,000 names across four lists might discover, after the merge-purge ran, that the deliverable unduplicated universe was closer to 130,000. The remainder were duplicates — names appearing on two, three, or all four of the rented files simultaneously. Since list brokers charged per thousand names regardless of overlap, the mailer had paid for the same household several times over. The overlap figures, once visible, made plain that the "distinct" audiences being sold were frequently not distinct at all.

The commercial embarrassment ran deeper than the billing question. High overlap rates suggested that the underlying compiled lists — assembled from voter registrations, phone directories, warranty cards, and purchase records — drew from the same source pools. A mailer targeting, say, outdoor enthusiasts might rent three separate "outdoor buyer" files only to learn that all three had been compiled substantially from the same mail-order response data. The PRIZM geodemographic clustering that Claritas made available after 1974 gave mailers a different purchase criterion — segment geography rather than category response — but it did not eliminate the overlap problem, it simply reframed it.

Bound archive volumes with typed entries rest open in gray plastic trays

ZIP+4 resolved an address to a block face; postage discounts, not segmentation, drove its adoption.

Sophisticated mailers began requesting RFM-ranked selections and demanding merge-purge reports as a contractual deliverable before finalising campaign quantities. The duplicate analysis also refined a mailer's understanding of its own house file: names appearing on every rented list were almost certainly already customers, and mailing acquisition offers to existing buyers wasted postage and muddied response measurement.

By the 1980s, the merge-purge had evolved from an inconvenient arithmetic exercise into a genuine strategic tool ↗ — one that pushed the list industry toward greater transparency about sourcing and, eventually, toward the data hygiene standards that Acxiom and Experian would later institutionalise across the warehouse era.