Hacker Newsnew | past | comments | ask | show | jobs | submitlogin
Google announces major update to combat AI-generated spam in search results (wired.com)
47 points by namanyayg on March 9, 2024 | hide | past | favorite | 20 comments


And what about the other human generated/bot spam? E.g. Pinterest spam, StackOverflow clones, content marketing sites, etc?

I've been using Kagi for a while and it makes it feel like Google Search is a dead man walking. Even Bing has better results for certain things than Google Search.

It's hard to see Google improving search to the point where it's good again, we needed a new disruptor in search to show the mega corps how to make good products.


To their credit they seem to have deranked or removed a lot of StackOverflow and GitHub clones from search results. I don't remember the last time I saw one. I used to see them in most of the results for my software-related searches.


The problem is that spam detection like that falls into the "I know it when I see it" category. Google can, and certainly must, use a blocklist for already discovered spam domains, but how could the reliably detect it automatically?


Google have human reviewers that can take manual action to penalize websites. People can also report spam websites.

https://searchengineland.com/google-issues-search-ranking-pe...

https://developers.google.com/search/help/report-quality-iss...

They are doing this now because AI chats and other search engine are becoming competition. They ignored spam because it was making them money.


That doesn't really solve the problem though as its dependent on humans flagging content after it was already indexed and surfaced by Google. Google doesn't currently have a way, nor do I think its possible, to reliably recognize AI content in an automated way.


one stupid idea for "de-dupiclating" LLM-generated summaries, translations and elaborations could be to ask another LLM for a spefific format of detailed summary and score on the similarity of that (punishing duplication of content that was already present in the index before and from a different domain)

Not a production-ready idea of course, and the ideal LLM for that might have yet to be invented.

Still, I have plenty of faith in Google still to apply techniques like this in a clever way.

They were very serious about their "Knowledge Graph" already 15 years ago, albeit with totally different data processing and acquisition techniques.

If all else fails, Google has historically put much of the burden on website authors, e.g. all the schema.org and other structured data stuff.

Their game has become much harder, but I still believe that Google can and will remain competitive in web search at least for quite a while to come.

Despite all the frustrating spam and a perceivable drop in results quality over the last 5 years, I still don't know about any fully competitive rival for Google.


My point is actually that the problem space itself isn't solvable, not that someone other than Google will solve it.

Google never solved how to win the game against SEO, and in recent years it seems they've either effectively lost or given up on it entirely. Generative tools make that much worse.

Comparing content based on LLM summaries wouldn't really do the trick. For one thing, the LLM itself is a black box so we don't know how it works and could never measure or predict the accuracy or reliability. For another, when two pieces of content have a matching summary the only tie breaker would be publication date (or first seen date in the case of Google). That really doesn't definitively answer the question of which was LLM generated, and won't help at all for content that is originally generated by an LLM and won't match any known content.

There's simply no way for an algorithm to accurately recognize computer generated content when the generation tools are good enough to mimic human content.


To train AI, you want contents that are human generated. So since Google is interested in the AI race now, filtering out AI contents is high on its list of priority. We are lucky it "trickles down" into us in this form.

Otherwise, Google being Google would ignore this issue just as they have ignored all the SEO and bot spams in the past. They lost their visions and has succumbed fully to corporatitis. I think anything short of a major surgery or blood transfusion would save Google now.


That's ultimately a losing battle, detecting AI-genersted content is going to be a back and forth battle similar to SEO, only likely harder to pull off.

Google could still detect what content gets the most engagement and try to guess at why, but they'll never be able to detect and filter out content simply because it was AI-generated.


Google is going to get rid of 40% of spammy content in their index? First thought: So they plan on showing 40% less search results

Article sounds like it was written by AI. So people take death information and make it into a youtube video. How does that relate to spam in google's search? Domain squatting isn't what she is describing. It accounts for so little of the spam problem. If google has been working on this update since last year and say they need more time meanwhile the llm scene exploded and keep changing daily how much faith do we have in google. We expect their own AI products to detect others?




Good.

I have a friend who's entire job is to generate ChatGPT spam articles about random shit to corporate blogs to drive traffic to their app.

Tells people she "works in IT" lol I suspect they pay like shit as employees have to use their own computer even.

Time to reign this shit in.


Jobs like this make me think of the opening scene from "The Republic". Socrates asks a rich man what's the best part about being rich. The rich man says that it's because he's never had to be a bad person to make money. Sometimes virtuousness is a luxury that only the rich can afford.


I do not care. Remove it.

I read Plato but do not remember. I need to read more, thank you random SF developer.


Will they also stop ignoring search terms, and stop their idiotic algorithmic aliasing of search terms?

When searching for Robert Jenkins Park, I am not looking for Bob Jenkins the person. Google needs to stop aliasing, and stop dropping search terms.

Google values giving gibberish over giving nothing, and nothing they do will help with irrelevant results, until the revert back to 2008 search behaviour.


That's more a question of defaults, right?

I'd hope adding quotes around Robert Jenkins Park would avoid aliasing. If so, then they optimized for aliasing being the default with a simple way to avoid it for the (presumably) uncommon case. If aliasing was off by default, what would be a similar syntax to indicate that you actually want it on for a search?


Obviously, IMO, it should be possible to disable aliasing and term-guessing account-wide. Then, I propose supporting a leading "?" that lets you reenable fuzzy/loose matching. Something like: "?robert jenkins park" (without quotes).


That syntax would impact the entire query, meaning you couldn't have it only fuzzy match a portion if the search term.

A global flag seems reasonable if enough people would want it, though that is dependent on users being authenticated with preferences stored on the server.


Wired going to bat for Google.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: