Duplicate Management: Frequently Asked Questions
What is duplicate detection?
Duplicate detection is how our system identifies when the same donor or supporter has more than one record in the database. It compares contacts that share a name pattern plus at least one strong signal (email, address, or phone), scores their similarity across eight fields, and flags matches above 70% confidence.
Why does duplicate detection matter?
Duplicate records make donors look less loyal than they really are and fragment your data across multiple entries. Donors end up with duplicates for all kinds of reasons such as signing up multiple times, filling out forms on different websites, or getting entered by different staff members. Catching these keeps your data clean and your donor picture accurate.
How does the system find duplicates?
The process has four steps:
- Step 1: Grouping Rather than comparing every record against every other record, the system first groups records into smaller candidate pools. Two records land in the same group if they share a first-initial + last-name pattern and also match on at least one of: email, address, or phone. This keeps the comparison set small and fast.
- Step 2: Scoring Within each group, the system compares records across eight fields: first name, last name, street address, city, state, zip code, email address, and phone number. Each field gets a closeness score whereas exact matches score high; typos or abbreviations score lower.
- Step 3: Weighting The eight field scores are combined using a weighted formula. First name and last name carry the most weight. Email and phone matter, but less so. Address fields are helpful but count the least.
- Step 4: Threshold If the combined score is 0.70 or higher (70% confident), the pair is flagged as a likely duplicate. Below 0.70, the system moves on.
Why group records first instead of comparing everything?
Comparing every record to every other record in a large database is computationally prohibitive. Smart grouping narrows the field to genuine candidates first, making nightly scoring of millions of records practical.
Why is the confidence threshold set at 70%?
A lower threshold (say, 50%) would generate too many false matches and burden staff with bad suggestions. A higher threshold would miss real duplicates. 70% strikes a balance: be confident before flagging, and accept that some duplicates will be caught in a later review rather than surfaced automatically.
What counts as a "strong signal"?
Email, phone number, or physical address. These are much less likely to be shared coincidentally than a name alone, so matching on any one of them is meaningful evidence that two records belong to the same person.
What if two people just happen to have the same name?
The system doesn't flag a pair based on name alone. Two "John Smiths" with completely different addresses, emails, and phone numbers will score low enough to stay separate. All eight fields contribute to the final score.
Does the system automatically merge duplicate records?
No. The system identifies and scores candidate pairs but does not merge anything on its own. The flagged list goes to staff who will make the final call on whether to merge. If staff determine that two flagged records are actually different people, that decision is recorded and the same pair won't be flagged again.
Can I dispute a duplicate match?
Yes. If you disagree with a flagged pair, you can use the “Mark as not a duplicate” functionality.
How often does duplicate detection run?
Every night. New imports are automatically checked against existing records as part of the nightly process.
Can the system detect duplicates across different organizations?
Not currently. Detection is scoped within each organization. A donor record at Org A and a record for the same person at Org B are treated as entirely separate.
What does the system not do?
-
It does not send donor data to any external service. All processing stays within our data warehouse.
-
It does not merge records automatically.
-
It does not operate across organizational boundaries.
-
It does not catch every duplicate, meaning that the 70% threshold means some lower-confidence matches are intentionally skipped.
What do you want to do next?
Interested in learning how Bonterra keeps your data safe? Read more about Bonterra's AI Responsibility and Data Security Practices here.
Ready to learn more about Duplicate Management? Access these 'How do I' articles:
Want to learn more about Que Skills? Check out the Que Skill Index here.
Not what you are looking for? Return to Understanding Bonterra Que here.
How does EveryAction detect duplicate donor records? | Why does my database have duplicate contacts? | How do I mark a record as not a duplicate in EveryAction? | Why isn't my duplicate being flagged in EveryAction? | How often does duplicate detection run in EveryAction? | How does EveryAction compare records to find duplicates? | Does EveryAction automatically merge duplicate records?
