We Asked AI About Our Artists
Ten K-pop teams, four questions each, four conditions, 280 responses checked against a verified answer key. Error rates ran from 3.8% to 40%. The widest gap turned up inside a single model: whether it was connected to search mattered more than which model it was.
fromis_9 has five members. Their agency is Ascend Entertainment. Their exclusive contract with Pledis ended on the last day of 2024, and the five signed with a new company and kept the group going. The other three went their own ways.
Ask some AI, though, and it will tell you there are eight of them, at Pledis. It says that on a paid plan with search switched on.

Why this can't just be shrugged off
Overseas fans used to meet a K-pop group by typing the name into a search box and clicking a link. Now a good share of them ask an AI instead. Google rebuilt its search box this year for the first time in 25 years and put AI answers out front. More than half of ChatGPT's users are between 18 and 34, exactly the band where K-pop fandom sits.
An AI answers before the profile and the press release a company wrote ever reach a fan. Companies mostly have no idea what that answer looks like.
Three things stack up here and make the problem bigger.
A wrong answer wears the same face as a right one. Across 80 ChatGPT responses in this survey, not one carried a hedge like "my information may not be current." There was no such signal where it said NewJeans had five members either. A fan reading it has no way to doubt it.

The timing is as bad as it gets. We will get to the detail later, but AI misses most on teams that changed agencies, changed members, or went through a dispute. It misses most precisely when a company most needs its information under control.
The courts are a poor fix. Ashley MacIsaac, a fiddler from Cape Breton, Canada, sued Google in May for 1.5 million Canadian dollars. Google's AI summary had introduced him as a sex offender convicted of sexual assault and child luring. A musician with three Juno awards became a sex offender inside one paragraph of summary.

Yet the people bringing these suits keep losing. A Georgia court sided with OpenAI last year, reasoning that users know hallucination is possible and a reasonable reader would not have taken the sentence as fact, and lawyers now describe a wide gap between reputational harm and legal remedy.
If you cannot fix it in court, one route remains. You manage the information the AI draws on. To do that you first need to know what is wrong and by how much. So we measured it.
How we ran it
We picked ten teams. Five had been through a lineup change or a contract dispute, three had moderate change, two have kept the same lineup since debut. For each team we asked four things in both Korean and English: current lineup, agency, official debut date, and recent activity.
We ran Claude, ChatGPT, and Gemini 3.6 Flash exactly as a user would use them, letting search run when it ran. We added one more condition, Claude with search blocked, because environments like that genuinely exist: the fan chatbots and in-app assistants agencies build often run on an API with nothing attached. We sealed the answers first and checked them against the facts afterward.
That comes to 280 responses.
The same condition, a tenfold spread between models
| Condition | Model | Error rate | Sample |
|---|---|---|---|
| User condition | Claude | 3.8% | 80 |
| User condition | ChatGPT | 12.5% | 80 |
| User condition | Gemini 3.6 Flash | 40.0% | 80 |
| Search blocked (control) | Claude | 40.0% | 40 |
The first three rows come from running each model exactly as a fan would. Same questions, same window, same grading. And out came 3.8% and 40%.
Gemini 3.6 Flash is worth a longer look. It scored the same as the search-blocked control. Gemini logged all 80 responses as having used search and even listed agency websites among its citations, yet the answers themselves were stuck in 2024. Open the Pledis homepage and you would see fromis_9 is not there. Open ADOR and you would see the Danielle story.
Worth remembering in practice: an AI naming a source does not mean it read that source.
Search connection split the results harder than model choice
The widest gap in this survey turned up inside a single model.
| Condition | Error rate, 40 Korean-language responses |
|---|---|
| Claude, search on | 7.5% |
| Claude, search blocked | 40.0% |
The same Claude, the same questions, the same answer key. Only the search connection differed, and the error rate spread more than fivefold. That is larger than any model-to-model gap we observed.
An agency has something to do here right away. If you build a fan chatbot or an in-app assistant and attach a model without connecting it to current information, you are opening a service at the 40% condition. Anything handling artist information should treat that connection as standard equipment.
Yet Korean fans barely use the most accurate one
One thing has to be added here in fairness. Claude scoring 3.8% does not mean fans are seeing that answer.
WiseApp Retail counted monthly active users in Korea for April 2026 at 23.45 million for ChatGPT, 8.45 million for Gemini, and 2.41 million for Claude. Lay this survey's Korean error rates over those numbers and the picture changes.
| App | MAU in Korea (Apr 2026) | Korean error rate |
|---|---|---|
| ChatGPT | 23.45M | 15.0% |
| Gemini | 8.45M | 42.5% |
| Claude | 2.41M | 7.5% |
Claude was the most accurate and is used the least. And Gemini, second in Korea, missed more than anything else in this survey. Roughly a quarter of what Korean fans see when they ask an AI about an artist comes out of the 42.5% condition.
Weight the three by their Korean MAU and the error rate lands near 21%. People use more than one app, so it is not an exact figure, but it works as a gauge. Ask five times as a Korean fan and one answer comes back wrong.
So the fact that one model gets almost everything right is no comfort. If anything, that is the argument for putting your information in order. A company cannot choose which model a fan uses, so it has to work on the sources instead, until whichever model picks them up returns the right answer.
Different conditions get it wrong for different reasons
Split the errors by type and each condition draws a completely different picture.
| Error type | Claude | ChatGPT | Gemini 3.6 Flash |
|---|---|---|---|
| Fossilized (change after training not reflected) | 0 | 0 | 27 |
| Search miss (picked a wrong or stale source) | 3 | 8 | 0 |
| Timing error | 0 | 2 | 2 |
| Hallucination | 0 | 0 | 1 |
| Process error | 0 | 2 | 2 |
Every Claude and ChatGPT error happened at the search stage. Search ran; what went wrong was which source got picked and how an ambiguous situation got read. ADOR has not settled Minji's status yet, and ChatGPT still declared NewJeans a confirmed four. With WJSN it flattened a split between official membership and active membership into one number of its own choosing.

Gemini 3.6 Flash is the opposite: 84% of its errors were fossilization. What it had learned was stale, and search results did not cover it.
So the two cannot take the same fix. An agency can influence which sources surface in search. Nobody can touch when a model's training corpus set.
Change the condition and it is not only the volume of error that moves, but its character. All three errors in the best-connected condition were partial: a missed recent change, nothing invented. In the search-blocked condition, by contrast, member names that do not exist showed up and whole lineup counts came out wrong. To a fan, "an answer missing one recent update" and "an answer with a member who does not exist" are different sizes of damage.
It misses most when it matters most
Split the teams by volatility and the slope is clear.
| Volatility | Claude | ChatGPT | Gemini 3.6 Flash |
|---|---|---|---|
| High (lineup change or contract dispute) | 15.0% | 25.0% | 55.0% |
| Mid | 0% | 0% | 25.0% |
| Low (no lineup change) | 0% | 0% | 25.0% |
All three missed most on high-volatility teams. The more a team has moved agencies, changed members, or been through a dispute, the more AI gets it wrong.
Which is exactly when information management matters most to an agency. Fans search to find out what happened, outlets pour out stories that contradict each other, wiki pages go through edit wars, and AI takes that whole mess and hands it back as one summarized sentence.

AI knows the big events and misses the small changes
Look closely at the three errors from the best-connected condition and they have a different grain. All three happened in 2026, and all three got little press.
Events every outlet covered, like the NewJeans dispute, Claude answered correctly. Same for fromis_9's move to a new agency. What it missed were the Loossemble members signing with separate companies and WJSN's tenth-anniversary activity.
A big event at a big agency gets the corpus filled in by the press. An agency change at a small company ends after a few lines of coverage, and then AI answers as if it never happened. The companies with the greatest need to manage their own information have the least capacity to do it.
That said, this is a pattern seen in three cases, not something to state flatly yet. The relationship between press volume and error deserves its own measurement.
A stable lineup is no protection
One thing in the table above catches the eye. Claude and ChatGPT both scored zero on low-volatility teams, while Gemini 3.6 Flash returned 25%.
SEVENTEEN and IVE came back right on lineup, agency, and debut date. Ask about recent activity, though, and the answer was a 2024 album and a 2024 tour. A lineup can hold still while the activity keeps moving.


Broken out by field, it gets clearer.
| Field | Claude | ChatGPT | Gemini 3.6 Flash | Claude (blocked) |
|---|---|---|---|---|
| Agency | 10.0% | 0% | 25.0% | 30% |
| Debut date | 0% | 10.0% | 10.0% | 10% |
| Lineup | 0% | 25.0% | 45.0% | 50% |
| Recent activity | 20.0% | 15.0% | 80.0% | 70% |
Recent activity was the weakest field in three of the four surveys; only ChatGPT scored slightly worse on lineup. This is the field that goes stale from nothing but the passage of time. If a company feels safe because its lineup has not changed, the lineup is the only field it is safe on.
This is where the earlier point about search comes back. Even the best condition got recent activity wrong 20% of the time. Set that next to lineup and debut date at 0% in the same condition and the contrast is plain.
Connect search and the fixed facts come back almost entirely right. How many members, when they debuted, one search settles it. What the group is doing right now survives search.
That is exactly where an agency has room to work. However well you tidy the fixed information, you are mostly re-solving a solved problem. Recent activity stays weak in every condition, so that is where the effort belongs.
Debut date is worth a note too. It is a fact that never changes, so it looks impossible to get wrong, yet three of the four surveys returned 10%, and always for the same reason. NewJeans came back as August 1, the EP release, instead of the official July 22; LOONA came back as August 20, the album, instead of August 19, the debut concert. Both are real dates, which makes the error hard to catch. Every artist has a different first stage, single release, and album release. If a company does not nail down one official debut date, AI will pick one for it.
English was more accurate
Our hypothesis going in was the reverse. English-language K-pop content sites are thick on the ground and plenty of them carry stale information dressed as current, so we expected English error rates to be higher. Measured, it came out the other way.
| Model | Korean | English |
|---|---|---|
| Claude | 7.5% | 0% |
| ChatGPT | 15.0% | 10.0% |
| Gemini 3.6 Flash | 42.5% | 37.5% |
All three were more accurate in English. Claude did not get a single one of its 40 English responses wrong.
What made the difference shows up if you take the three Korean errors and ask them again in English. All three came back correct.
WJSN's tenth-anniversary comeback in February 2026 did not surface high in Korean search, but English-speaking fans had put it on allkpop, the fandom wiki, and English Wikipedia, where editors had even updated the active years to "2016-2023, 2026-present." LOONA is the more dramatic case. Fan-built profile sites in English carried the Loossemble members signing with separate companies in 2026, right down to Hyunjin debuting with the band Latency in January. Korean search turned up neither.

The reason is not hard to guess. Korean outlets pour out stories when something breaks and never update them afterward. English-language K-pop fan communities keep rewriting their wikis and profile pages. A reporter writes down the day something happened and stops there; a fan follows what the current state is and keeps fixing it. Ask an AI about the current state and the side the fans maintain gives the better answer.
The information ecosystem a Korean agency needs to manage does not sit only in Korea. Right now it is overseas fans who tend the English wikis more diligently.
None of which is grounds for optimism about going global. Sites carrying wrong information do rank high in search, and if a model picks those over Wikipedia the result flips. Among the top results for NewJeans in English was a site listing six members. NewJeans has never had six.
How this differs from search optimization
The work now goes by GEO, optimization aimed at generative engines. The name is close enough to search optimization that it is easy to read as an extension of it. In practice, quite a lot is different.
It used to be enough to bring people to your page; now reach happens inside the answer even when no click occurs. You have to deal with exposure your traffic dashboard cannot see. The unit of competition changed too. Instead of ranking your site, an AI builds a single representation of "who is this artist." Comms teams have never handled an object like that.
Control sits outside the company as well. Fixing the official site alone changes nothing. An AI builds its answer from the whole set of sources it consults, and official channels are only part of that set. Time runs differently too. A search index changes when you update it; the representation inside a model stays fossilized. This survey showed exactly that.
Still, AI does cite official channels
One encouraging figure came out of it. Of 189 citations collected in the ChatGPT survey, 67, or 35%, were official sources. At least one official source appeared in 57.5% of answers.
The domain AI cited most was Weverse at 25 times, followed by Starship's official site at 11, Pledis at 7, and Source Music at 6. On NCT items it cited the original Weverse notice.
Run an official channel properly and AI does cite it. Note, though, that "official channel" here means the places where artist activity goes up in real time. Fan communication platforms and artist profiles on music services, the places people actually visit and search actually indexes. A corporate about page is not one of them.
What each side can do
If you are an agency or a label, start by checking where you stand. Ask the major AI tools about your artists once each and log the answers; ten minutes covers it. Start with the apps most used in Korea, ChatGPT and Gemini. Gemini in particular is second by MAU in Korea and missed more than anything else in this survey.
Next, check whether your official information is in a shape machines can read. Find where the company has pinned down the basics: member list, agency, debut date. Check whether that page surfaces in search. Check whether it gets fixed within days when something changes. When information changes, at a release or a lineup change, do not stop at sending a press release; fix the official source first. A press release is a channel for people.
Recent activity comes first among all of it. Member lists and debut dates already fell to 0% in the well-connected condition, while recent activity was still wrong 20% of the time in the best condition, so put the effort into the unsolved side rather than re-solving the solved one.
The less press a company gets, the more urgent this is. As above, the press fills in the corpus for a big event at a big agency, and does not for a change at a small one. The side with less capacity has to be the more diligent one.
English information needs its own attention. English came out more accurate in this survey, but that is because overseas fans maintain English Wikipedia and their communities well, not because any company managed it. Fans are the maintainers, and it goes stale the moment they stop.
For artists and managers, make a habit of putting your own name into an AI. Check without fail after anything that changes your identity information, like a name change or a move. This survey turned up several cases where a name change had not registered.
For marketers and platform leads, there is one more metric now. What do you measure exposure with when no traffic comes from it? Log whether you were cited and what got cited, on a regular schedule, and you will have a baseline to compare against later. It also happens to be a field where starting now is an advantage.
If you build AI services, make the search connection standard. Run the same model without it and the error rate jumps from 7.5% to 40%. That is the exact condition an agency walks into when it builds a fan chatbot.
Finally
We started out meaning to count how much AI gets wrong about artist information. Counting it turned up something more important.
An agency cannot decide which model a fan uses. Even within the same user condition, error rates ran a tenfold spread across models. Making every AI accurate cannot be the goal.
And most Korean fans do not use the condition that scored best here. Weighted by MAU in Korea, the Korean-language error rate sits near 21%. One good number is not grounds for relief.
What can be done sits elsewhere: put accurate information where machines can find it, fix it fast whenever it changes, and make it obvious where fans should go to check. In this survey, AI cited Weverse, a place fans actually visit, more than anything else.
The point from the opening is worth returning to. ChatGPT attached no hedge anywhere in 80 responses. A fan who gets a wrong answer gets no signal that it is wrong.
Someone is asking an AI about our artist right now. The company has no idea what answer they saw.
Monthly active user figures for Korea come from a survey WiseApp Retail released in May 2026, covering April 2026 across a panel of 36.61 million Android and 14.61 million iOS users, 51.22 million in total. (ZDNet Korea)