KK-DATA avatar KK-DATA

Similar Audience Expansion in Practice: Using High-Quality Screened Number Data to Build a Precise Retargeting Seed Pool

受众 扩展 kkdata 种子数据 精准投放

Lookalike Audience Expansion in Practice: Building a High-Quality Seed Pool for Precise Retargeting with Screening Data

As overseas customer acquisition enters the stage of存量 competition, Lookalike Audience expansion has become a key method for many teams to break through cold-start bottlenecks. However, many invest budget and run models only to find frustratingly low conversion rates. The problem usually lies not in the ad platform’s algorithm, but in the quality of the seed data.

This article will provide a complete practical workflow and pitfall avoidance guide, centered on how to use a screening system to obtain high-quality seed numbers (valid, active, with gender labels) and then use them for Lookalike expansion and retargeting on ad platforms. Whether you are operating Telegram communities, WhatsApp marketing, or multi-platform customer acquisition, this approach can help you improve ROI.

The Cold-Start Dilemma: Why Are Your Lookalike Audiences Underperforming?

The core of Lookalike audience expansion is: the ad platform learns the common characteristics (interests, behaviors, devices, etc.) of the users in the seed list you provide, then finds other users matching that profile across the platform. The quality of seed data directly determines the quality of the model. Two common problems occur in practice:

Invalid Numbers Dilute the Lookalike Model

If the seed list contains a large number of empty numbers or numbers not registered on the target platform, the ad platform cannot capture real user features when matching. These invalid numbers act as noise, pulling the model’s “attention” off course, resulting in poorly correlated expanded audiences. For example, if your target users are active cross-border e-commerce buyers on Telegram, but the seed list contains 30% numbers that are not registered or have been deactivated, the model will learn from non-existent user profiles, with predictable results.

Lack of Behavioral Labels Leads to Blurry Profiles

With only a phone number, the ad platform can obtain very limited user signals. Without activity (how often they log in) or gender/identity information, the model can only rely on weak signals like number region or network type, resulting in expansion akin to “the blind men and the elephant.” For example, if your provided numbers include both heavy daily users and silent users who registered but never logged in, the platform cannot differentiate, and the final expanded audience quality will inevitably suffer.

Three Key Dimensions of High-Quality Seed Data

What kind of seed data can support a high-conversion Lookalike audience? At a minimum, it should satisfy the following three dimensions:

  • Number Validity: Ensure the number is registered on the target platform (Telegram/WhatsApp, etc.) and currently reachable.
  • Platform Activity: Be able to distinguish whether a user is “active in the last 7 days,” “active in the last 30 days,” or long-term silent, allowing for stratified targeting based on different conversion intents.
  • Gender/Identity Recognition: In certain scenarios (e.g., beauty, menswear, maternal & baby products) that are sensitive to user gender, seeds with gender labels can significantly improve model accuracy.

When these three dimensions are combined, the seed pool is no longer a simple “collection of phone numbers,” but a set of high-value data with clear behavioral profiles.

Practical Workflow: From Number Screening to Lookalike Ad Placement

The following steps represent a complete operational chain, implemented using the KK-DATA screening system. If you use other tools, the logic is reusable, but specific interfaces may differ.

Step 1: Generate and Screen Global Numbers to Obtain Target Seeds

Open the KK-DATA console. In the number generation module, select the target country/region (supports 240+ countries) based on your target market, and batch generate numbers. Generation is free; you can set the quantity based on your target audience coverage (e.g., Southeast Asia, Latin America, Middle East).

Then go to the screening module and submit a Telegram or WhatsApp screening task. Taking Telegram as an example, you can check:

  • Telegram registration check: Determines if the number is registered on Telegram.
  • Telegram validity check: Confirms the account can currently receive messages (not banned/abnormal).
  • Telegram activity check: Specify activity windows of last 7/15/30 days to filter active users with recent login behavior.
  • Telegram gender recognition: Identify male/female users via avatar recognition (optional).

Before submitting the task, the system will display an estimated cost. Confirm to start screening. After completion, you will receive a Telegram notification.

Step 2: Extract High-Value Screening Results and Export as Seed List

After screening, log in to the KK-DATA app console and view the results. In the export module, you can combine filter criteria:

  • Condition A: Telegram valid + active within 30 days + gender male
  • Condition B: Telegram valid + active within 7 days + gender any
  • Condition C: WhatsApp valid + wsid exported

Export qualifying numbers in CSV or TXT format. At this point, your seed pool has invalid numbers removed and carries activity and gender labels, making it much cleaner than the original number list.

Step 3: Import Seeds into the Ad Platform to Create Lookalike Audiences

Taking Facebook or Google Ads as an example, the general procedure is:

  1. Log in to the ad account, find “Audience Management” or “Customer Match” tools.
  2. Upload your CSV/TXT seed file, map fields (phone number, email, etc.).
  3. The system will match numbers (usually via hash encryption), creating a seed audience.
  4. Based on the seed audience, click “Create Lookalike Audience”: choose the audience ratio (1%-5%, smaller ratio means higher similarity but smaller coverage), confirm and start expansion campaigns.

Different platforms have minimum seed number requirements (usually at least 100 effective matches), but quality is far more important than quantity — 500 high-quality Telegram active user seeds often outperform 50,000 unscreened number lists.

The Hidden Value of Data Deduplication: Avoiding Waste and Improving Model Purity

If you run multiple screening tasks simultaneously (e.g., different countries, different periods), the same number may enter the seed pool multiple times. This creates two issues:

  • Duplicate matching by the ad platform: The same user is affected by multiple seed records, biasing model training.
  • Wasted balance: Repeatedly screening the same number consumes screening balance without adding value.

KK-DATA has a built-in cross-task data deduplication repository that automatically identifies and removes numbers already present in other tasks. Before importing new numbers to generate seeds, it is recommended to run them through the deduplication module first to ensure each seed is unique and fresh. Deduplication may seem minor, but it significantly improves Lookalike model purity.

Retargeting Strategy: Layered Targeting Using Screening Results

Once the seed pool has activity and gender labels, you can move beyond “one-size-fits-all” ad placement and achieve layered precision retargeting.

High-Activity Seeds for “Instant Conversion” Audiences

Select numbers active within the last 7 days as seeds. These users have high device usage frequency and message response tendency. Lookalike audiences built from them are more likely to respond quickly to instant offers or limited-time promotions. Suitable for scenarios like e-commerce flash sales, app download promotions.

Low-Activity Seeds for “Wake-Up and Remarketing”

Users who logged in within 30 days but are not high-frequency can serve as seeds for “dormant users.” These Lookalike audiences cover a broader range but have milder conversion intent. Suitable for strategies like brand exposure, daily content touchpoints, wake-up discounts — which expand reach while controlling ad costs and gradually reactivating users.

Which Pitfalls Could Cause Lookalike Audience Failure?

Here are common mistakes teams make in practice, each capable of directly lowering ROI:

  • Insufficient seed quantity: Fewer than 100 effective matched seeds prevents the platform from building a stable profile. Aim for at least 500-1000.
  • Ignoring country differences: Using numbers from country A to create Lookalike audiences for country B leads to completely mismatched user traits. Always segment seeds by target market.
  • Not performing gender filtering: For gender-specific products (e.g., female cosmetics), without gender labels, the model may expand to include a large number of opposite-gender users, wasting budget.
  • No deduplication before import: Duplicate numbers distort the model; multiple records make the platform think that user’s features are especially important, when it’s just data redundancy.
  • One-time import without updates: Number validity degrades over time; churn increases. It is recommended to rescreen monthly or quarterly and replace aging seeds.

Important Reminder: Seed Data Privacy Compliance

Ensure the numbers you use are obtained legally (e.g., users voluntarily registered, have consented to receiving marketing communications). Before uploading numbers to ad platforms, comply with local privacy regulations (e.g., GDPR, CCPA). KK-DATA only provides number validity detection tools; it does not participate in data collection or usage compliance.

How to Measure Lookalike Audience ROI?

To evaluate Lookalike expansion effectiveness, compare the following metrics before and after:

MetricLow-Quality Seeds (Unscreened)High-Quality Screened Seeds
CPM (Cost Per Mille)Often 20%-50% above averageNear or below average
CTR (Click-Through Rate)Typically below 0.5%Can reach 1%-3% or higher
Conversion Rate (Purchase/Registration)Mainly low efficiency trafficCloser to target audience
CPA (Cost Per Acquisition)Obvious budget wasteCan be reduced by 30%-60%

If your Lookalike audience expansion consistently fails, pause and review the quality of your seed data. An efficient Lookalike model depends 90% on seed data and 10% on the platform algorithm.

Tip: Cost Optimization Advice

KK-DATA charges per call, no subscription plans. It is recommended to first add a small amount of USDT for a small-scale test (e.g., 1000 numbers) to verify screening effectiveness before large-scale generation and screening. The estimated cost is shown before task submission; insufficient balance will prevent submission, avoiding overspending.

Frequently Asked Questions

Q: What is the minimum number of seed data needed for Lookalike audience expansion?

A: Requirements vary by platform. Facebook suggests at least 100 valid numbers in the seed list, and Google is similar. However, seed quality is far more important than quantity; even 500 high-quality seeds can outperform 50,000 low-quality lists.

Q: How can the activity data from KK-DATA screening be used for Lookalike?

A: You can export numbers by activity windows such as “active in the last 7 days” or “active in the last 15 days” as seeds of different time granularities. High-activity seeds target immediate conversion; low-activity seeds cover a broader audience while maintaining relevance.

Q: After uploading seed numbers to the ad platform, how will the platform use them?

A: The ad platform uses these numbers to match against its user database, builds a seed user profile (e.g., interests, behavior, devices), and then finds other users similar to that profile to form a Lookalike audience. The platform does not expose matching details; the seed numbers are only used to train the model.

Q: Do I need to update the seed pool regularly?

A: Strongly recommended. Number validity decays over time (users cancel, change numbers, become dormant). It’s advisable to rescreen monthly or quarterly and use the deduplication repository to exclude duplicates. An updated seed pool keeps the Lookalike model timely.

Q: Is it feasible to run Lookalike expansions for both Telegram and WhatsApp simultaneously?

A: Yes, but they need to be created separately. User groups and behavioral traits differ across platforms. It’s recommended to use screening results from a single platform as seeds and create Lookalike audiences for that platform accordingly. For example, use Telegram active numbers to test Facebook Lookalike, and WhatsApp active numbers to test Google Customer Match.


Action Guide at the End