Every creator CreatorPal finds gets a grade (Recommend, Human review or Not recommend), a one-line reason and a personal opener. Large language models produce all three. This post explains how that works, where it went wrong, and what we changed, with real numbers from our own testing.
Two questions, not one
"Is this a good creator for me?" hides two different questions:
- Is it on-topic? Does this account actually make content about what you searched for, say home baking?
- Does it fit the brief? Given your goal, your offer and your audience, is this creator worth contacting?
A creator can be perfectly on-topic and a poor fit (a brilliant baker when you are promoting a mobile game), or a good fit and off-topic. We grade them separately: the topic check decides whether a creator belongs in your results at all, and the grade measures fit with your brief.
The pipeline
- Word match (free). Every important word of your search has to appear somewhere in the creator's bio, captions or analysed topics. "Mobile gaming" needs both "mobile" and "gaming", not either one. This throws out the obviously unrelated before anything is paid for.
- A small model's screen. A fast, inexpensive model (Claude Haiku) reads a batch of bios and captions and keeps only individual creators whose own content is about the topic. Shops, brands, manufacturers, restaurants, museums, official pages and fan or repost accounts are dropped. Creators fetched live are screened before anyone is charged a credit for them.
- A stronger model's grade. A more capable model (Claude Sonnet) reads the bio, recent captions and up to two post images against your brief. It returns the grade, a short reason in your interface language and an opener that references something the creator actually posted.
What went wrong first
The first version leaned on a coarse category: "home baking" mapped to food & drink, and the database returned any fresh food account in your follower range. The model then (correctly) rejected them, so a real search for home bakers produced this:
- 29 creators found, all archived as off-topic: mostly mukbang (eating-on-camera) channels and restaurant accounts.
- 104 live-lookup credits charged for creators nobody wanted.
Two subtler failure modes showed up once we tightened things:
- One mention is not a niche. A food reviewer who visited a bakery twice passes a "mentions baking" test. Text matching cannot tell a baker from someone who once ate a croissant; a model reading the whole profile can.
- Broad category is not topic. PC strategy streamers are "gaming" but not "mobile gaming". Collectible shops are "gaming" but not creators. The screening prompt now names these traps explicitly.
Details that mattered
- Missing images are marked, not dropped. If a post image can't be loaded, the model is told "image unavailable" instead of seeing an empty slot. Otherwise it confidently describes nothing.
- Fail open on errors, not on judgments. If the AI service hiccups, creators are kept rather than silently lost. If the model says "off-topic", that is respected.
- Language-aware output. Reasons come back in the user's interface language (we support English, Chinese and Thai), while openers stay in the language of the message being sent. Thai and Chinese use several times more tokens than English, so output budgets have to allow for it.
- Sources matter as much as models. No grader can rescue bad candidates. Searching reels about a topic, then exploring accounts similar to the real topical creators found, produced far better candidates than keyword user search.
The result
The same "home baking" search, 10 creators, 500–30,000 followers, after the changes:
| Before | After | |
|---|---|---|
| Relevant creators kept | 0 of 29 | 10 of 10 |
| Credits charged | 104 | 6 |
| Time to finish | stopped after 4+ minutes | about 2 minutes |
CreatorPal internal test, October 2026. One niche and one run, so treat it as illustrative, not a benchmark.
Why this matters if you're evaluating AI tools
Ask any AI influencer tool two questions: how does it separate "on-topic" from "good fit"? and do you pay for results it then throws away? The answers tell you more than any demo.
Try the grader on your niche What to automate with AI agents