What data sources does AI Research use?

The short answer: two separate sources, depending on the step. Social Media Scanning pulls from real public posts on Xiaohongshu, Douyin, X/Twitter, Instagram, and TikTok. Interview and Discussion use atypica's pre-built persona library (~5,000 personas), which doesn't depend on social media data — those personas were built from a mix of real interview transcripts, user profiles, and prior Scout runs.


Sources by feature

FeatureData sourcePrivacy scope
Social Media ScanningPublic posts on 5 platformsPublic only — no DMs, no closed groups
InterviewPre-built persona libraryBuilt from anonymized research and Scout outputs
DiscussionPre-built persona librarySame as Interview
Scout-generated personasPublic posts (above)Same as Scanning

You don't pick one source — different features use different sources automatically.


What Social Media Scanning actually pulls from

Domestic (China): Xiaohongshu, Douyin.

Overseas: X/Twitter, Instagram, TikTok.

Coverage is recent public posts only — typically the last few weeks to months. No historical archives. No private accounts. No login-gated content.


What Interview/Discussion personas are built from

The public persona library has ~5,000 personas covering common demographics. Each persona is built from a mix of:

  • Real user interview transcripts (anonymized)
  • Behavioral profiles from research contexts
  • Outputs from prior Scout runs

Personas are stable: once a persona exists, it answers consistently across all interviews and discussions. Personas don't update their behavior based on real-time social media — they speak from their fixed background.


Why Interview doesn't check social media in real-time

If you ask a persona in an Interview "would you buy ¥18 sparkling coffee?", they answer from their persona background — not from what's trending on Xiaohongshu today. This is intentional:

  • Personas are consistent, so research is reproducible
  • Real-time social data wouldn't fit a single persona's perspective
  • Mixing live data with persona responses would muddy what the persona is supposed to represent

For real-time trend data, use Scout. For depth on individual motivations, use Interview. The two complement each other.


Choosing platforms for Scout

Let AI pick by default — it matches platforms to your demographic.

If you want to specify manually:

DemographicBest platform(s)
Young women (18-35), consumer goodsXiaohongshu
Gen Z (18-25), short videoDouyin
Lower-tier marketDouyin
Tech, business, opinion leadersX/Twitter
Fashion, lifestyle, visual brandsInstagram
Young global users, product reviewsTikTok

For cross-validation, scan 2–3 platforms and look for consistency. Contradictions signal you need more digging.


Regional filtering

Scout accepts regional filters:

  • Tier-1 cities (Beijing, Shanghai, Guangzhou, Shenzhen)
  • New tier-1 cities (Chengdu, Hangzhou, Wuhan, etc.)
  • Provinces (Sichuan, Zhejiang, etc.)
  • Overseas regions (USA, Europe, etc.)

Filter too narrow and there's not enough content. Filter too broad and the output gets noisy. Match your filter to your actual research question.


Data freshness

SourceFreshness
ScoutRecent (last few weeks to months)
Persona libraryBuilt periodically, not live

If your research question is about current trends or recent events, run Scout first. If your question is about stable attitudes or decision-making patterns, Interview alone is enough.


What isn't included

  • Historical archives — Scout only pulls recent posts. No "what were people saying in 2023" research.
  • Private DMs and closed groups — only public posts are scanned.
  • Real-time social listening — Interview/Discussion personas don't know what's trending today.
  • Cross-platform identity resolution — Scout can't reliably link the same person's posts across Xiaohongshu and Weibo.

For any of these, supplement with a dedicated social listening tool (e.g., specialized analytics platforms).


Privacy and compliance

What gets scanned: publicly posted content only. No private messages, no closed-group discussions, no login-required pages.

What's extracted: themes, attitudes, language patterns, behavioral observations — not user IDs or identifying details.

What stays out: original post text isn't stored after the run; only summarized insights are kept.


Related: Scout, Interview, Discussion, Persona Library

Last updated: 8/8/2026