Wait, hold up. LLMs may be non-deterministic, but they're not _random_.
Take the author's sunset argument. What if I painted 2 pictures of a sunset, then put them up on a webpage and randomly picked one for you to see. Would you say there's no intentionality, only randomness? Of course not. Both paintings are still human creations.
LLMs are trained with human feedback. It's distributed and high scale and the outputs are truly surprising in many cases, but there's a heavy hand on what comes out of it. They're created (largely) by people who think omniscient, helpful AI would be cool to have, and they mostly respond in the way that's aligned with the hopes and dreams of those people. Do you think the frontier labs are mad, embarrassed, and disappointed with their LLMs hacking out of their terrible sandboxes? No, they think it's the coolest thing in the world. They trained the model, hoping that would happen.
There's deep intentionality behind the models. But it's not the models that hold it.
Article says that there's no human intention or design directing the output we get, but I don't think that's completely right. It's not human, but the algorithm is like the Human Instrumentality Project: an amalgamation of human intentions.
I think they're pointing to the fact that RLHF means that the intention is from humans and not random.
I'm not sure if it's their intent, but I wonder if one could still consider these artifacts as "intentional", but not individual attention creating them, and rather an aggregate, soupy collective attention.
Obviously, important signal in the human experience is lost there, and we get a soupy middling sort of creation. But it's not random, as I believe the parent was pointing out.
EDIT: overall, I align with the article. am just thinking aloud about the contrarian positions, though not committed to them
Please do not confuse uniform randomness with randomness at all.
On your die, if you colour one face in black and the other 5 faces in black, rolling it will still produce a random outcome. It's just that you get a 1:5 skew.
Or look at the probability that a given C-14 carbon isotope will decay tomorrow. As far as we can tell, that's as random as it physically gets. With the die you could theoretically try to run a physics simulation to predict how it rolls, but as far as we can tell, atomic decay is intrinsically random.
However a C-14 atom has about 1 in 3 million chance to decay on any given day.
I made no claim about the distribution and frankly it has no bearing on my point or the person who said he could post two pictures. The point is there are a lot of numbers / pictures / essays / that will never be generated, and someone is controlling what is possible. In fact people are working very hard to produce distributions pleasing to them, like your die.
Scrambled intentionality is nearly impossible to productively analyze.
When you randomly choose a picture to show me, I cannot glean any intent from being shown that specific picture, but I can glean some intent from the set of pictures you could have shown me, and in the relationships among the elements of the given picture you did show. Any randomness cuts out some intention.
When a picture is derived from huge model, any intention is mulched up to a degree that analyzing the picture for meaning is pointless.
- You can use AI to look through docs and do research.
- You can use AI to structure your argument.
- You can NOT use AI to actually write the prose.
- You can use AI to proofread, especially for how effectively you've communicated.
I don't know, man, that sounds a lot like using AI to write. Your writing probably says "load-bearing" less than Claude's outputs, but I'm not sure it's much less AI.
Reading through docs and doing research: of course. They're like much better search engines in that regard. It would be silly not to use them. Obviously, you need to review their search results against their sources.
Structuring arguments: bad idea. Here you're having the model do your thinking for you. The problems that occur with having it write text recur at a macro level; too-obvious patterns. Readers just want to know what the prompt was.
Writing the prose: obviously not.
Proofreading: I think this works pretty great, as long as you have a structure to how you're doing it. "How do I make this better or sharper": bad, you're having the model structure your arguments for you. But for a predefined rubric of stylistic things --- writing you know how to do, that's just tedious to check manually? Hard to see a downside.
Most people apparently still don't get that using AI for mere proofreading is the same as using it as a mere knowledge base summary tool. Instead of proofreading, have it be adversarial. Let it try to take apart your argument. The latest generation is extremely good at calling out half-truths and circumstantial evidence across many fields. These days I would not dare to put anything in front of a human peer that hasn't passed at least an AI sanity check review.
I don't want an AI's opinion on what I'm writing. I want its assessments against an objective rubric. The whole point is not to have the model weights do your thinking for you.
Structuring your argument is probably something you should do, not an AI. This is the core of what you're trying to convey, and is exactly the type of thing you should be thinking through.
Of course, this is different than having the argument ready and using an AI to make it read a little better, which at that point, is just details.
The likes of Sir Walter Scott, Peter Ackroyd, Stephen King, Tim Ferris, Robert Greene, Tucker Max, Catherine Asaro, Neil Gaiman, Charlaine Harris, Nancy Holder etc... all rely heavily on research assistants in various capacities including proofreading.
Honest question - where does an Author like the above place in your authenticity framework?
Giving the compiling evidence of how addictive generative AI is actually, it actually makes perfect sense that people are looking for strategies to excuse their AI-usage. I think no matter the strategy (and given AI is actually as addictive as evidence are suggesting), chances are the usage patterns will all converge to “Just use AI dummy“.
This reminds me a lot of smokers who all have a great strategy beat the nicotine addiction. Only smoke on the weekends when out partying with friends, only smoke to use the extra sense of energy, only smoke to suppress my appetite, go to the gym 5 hours a week to offset the negative health impacts of smoking etc. etc.
In reality most smokers will converge to about a pack a day.
there's a growing contingent of anti-ai extremists shouting down anything that looks or feels like the thing they hate: actual slop. but there's a missing understanding of where slop really comes from - "you used ai instead of thinking" is just another guess, and mostly wrong since you literally cannot use ai without thinking - the prompts don't write themselves.
slop has been around a lot longer than ai - the contingent needs to realize what it is they hate. instead of "don't use ai to write" it's "don't use ai to write slop" or better: "don't write slop"
To be clear, I (the author of the article) do not hate AI. A lot of my current work is focussed on making AI work well.
However, I don't think it's a good idea to replace our own thinking with an AI, and that's what I'm trying to avoid with the way of working proposed in the article.
No. I hate AI and slop. I hated slop before AI, but with AI slop has proliferated to an unprecedented degree. If AI went away tomorrow, slop would go down with it to 1/1000ths of what it is today.
Isn't this the classic blackout scenario? Claude goes down, then people move over to Codex, which is overwhelmed, and crashes, so people move to Grok...
When you're dealing with millions of users and response times go up above timeouts, there isn't much difference between "at capacity" and "down" if most users can't reliably use the service.
Bit of a misleading status page, if you click in you can see that grok 4.5 and 4.6 etc are totally down, with the rest of the models showing as "up". I_strongly_ suspect they are not weighting it to actual number of requests!
An intuitive explanation is that financial products are, approximately, buying and selling as part of the same transaction. You can't separate the "selling premiums" part from the "paying out claims" part.
This is true of life insurance, investment firms, and banks. It's also true of marketplaces that connect buyers and sellers, like Etsy.
Groceries stores are buying from suppliers and selling to consumers, but those are separate operations. If the consumers opt out, the grocery stores (temporarily) still have a full and complete obligation to their suppliers. It's hard to sell to customers without supply, but if you try hard, you could theoretically do that as well.
Somebody with a better financial background might be able to define the nuances of accounting practices here, but there's already a pretty meaningful line that's established. It is kind of weird that health insurance doesn't behave like a financial product.
This is contrary to GAAP and operationally false. An insurer takes on the risks including the health of the insured pool and cost changes during the covered period. An insurance BROKER or AGENCY only books commissions as revenue, but an INSURER books premium as revenue. Similarly a stock BROKER or AGENT is only acting as an agent and isn’t a party to the actual transaction they execute. Similarly for platforms, auctioneers, or other agents.
I work for an insurance company so can shed some light here as this article is written by someone that clearly doesn't understand how the business model works.
Fundamentally every insurance company is governed by 3 ratios, loss ratio (what percentage of premium is paid to make the buyer of the insurance whole), expense ratio (cost of doing business, paying staff, keeping office lights on, paying vendors) and combined ratio (both of these combined). These are true for any insurance company which writes premium using their own capital, whether its health insurance, life insurance, property insurance, SMB insurance.
The thing this article is missing here is that the "pass through" costs are costs incurred by UHG directly, they are the ones paying the bills. How is this pass through, it's not being passed to the consumer, the only thing I pay is my deductible and retention which is at most a couple of thousand dollars, these are true costs borne by UHG. So in practice if I pay 100 bucks every paycheck, UHG is taking in 2600 bucks worth of premium, using average industry loss ratios which are say 60%, UHG is paying directly 1,560 bucks to care providers for my own care. I'm not paying that, what I pay is a deductible which is treated entirely separately.
I am the biggest insurance skeptic in the world because I think the business model is awful, a business's return on capital averages at 5-10% a year which is truly an awful return for how much capital is required. Insurance companies will make between 0 and 10% of underwriting profit a year (the pure profit from insurance premium minus total expenses) and they usually operate a very large investment vehicle invested typically 70% into bonds/gilts. That being said, this doctor's view of how insurance accounting works by comparing it to a biopharma or a trading brokerage firm is immensely disingenuous.
Sure, you can let politics dominate everything you do. Or you can realize that SpaceX is a massive (public) company with thousands of employees, and millions of shareholders, all of whom have their own opinions and goals, just like any other corporation.
Competition is good. Excluding a leading player in the market because you don’t like Elon Musk is…something.
> Sure, you can let politics dominate everything you do.
I assume they're referring to the recent discovery that Grok Build was uploading entire repositories to their servers in the background, include .env secrets that had been excluded
It was a bug, and was immediately corrected. The other harnesses have bugs too. You just don't know about them.
Frankly, the people who keep bringing this up are mostly engaged in motivated reasoning. I don't trust any company, and any product where I have to send my code to a third party to make it work is a devil's bargain. I don't trust any of the major labs, but it is what it is.
The only way forward is local models, but we're not there yet.
There's a difference between not trusting a company because it's a company driven to chase profit at the expense of everything else, versus the same thing but it's run by a literal nazi that uses his companies as leverage to undermine democracy and enrich himself.
Sure, way to young for that; but he is very much cut from the Nazi adjacent, racist, antisemitic, antidemocratic, technocratic views of his whole heartedly apartheid embracing grandfather Joshua N. Haldeman.
His grandfather wrote his tracts to raise an alarm about what he called “mind control,” on the radio and television, where “an unconditional propaganda warfare is carried on against the White man.”
Absolutely no point arguing with these people. They are so ideologically captured that it's pointless trying to discuss anything with them. Just move on. Everything and everyone they disagree with is "nazi" and if you say anything contrary to that you're a "nazi" also.
He does the same exact literal Nazi things that Nazi did and has Nazis in his family tree. You'll need better arguments to defend him than just "I don't agree".
Bizarre reply. It's not "dominating everything you do". It's one specific thing. Grok exists in a very crowded space and it's incredibly easy to not use it. If this is your reaction to someone taking a very easy stand on their personal principles, it does not reflect well on you.
"Bizarre" only in the sense that you're purposely trying not to understand.
Literally every commodity product is in a crowded space, and easily substituted. If it's a good product (Grok Build objectively is one of the very best in the space), it's a good product, and it's self-defeating to avoid it because you hate a guy for political reasons.
Just as it would be nonsensical to avoid shopping at WalMart, Target, or any of a million other places. Because I guarantee they're all associated with people you won't like.
On the contrary, I'm trying to understand as best I can. I find your position so completely devoid of any inkling of personal or social responsibility that it takes quite some empathy to muster up a reply that's not wildly uncivil.
I also think it's callous to brush away "political reasons" as though it's some trivial abstract thing. Or perhaps it comes from a place of nihilism?
I choose not to give my money to people I think are enormously evil. That's really all there is to it. I don't see why this is "nonsensical".
> I find your position so completely devoid of any inkling of personal or social responsibility that it takes quite some empathy to muster up a reply that's not wildly uncivil.
Oh stop. Other people believe different things than you. If you cannot see how using a coding agent is not "devoid of personal or social responsibility", then you really need to step away from the keyboard.
Using or not using a coding agent is not what I took issue with.
What I find devoid of responsibility is the argument you made originally--that making choices as a consumer informed by anything external to the direct value you're paying for is pointless/inexplicable/self-sabotaging/whatever.
Based on your reply, I'm not actually sure if you actually believe this, or if it's only a form of motivated reasoning because you have some positive feelings about Elon or whatever, and that we wouldn't be having this argument if the original commenter was boycotting some other product for some reason you agreed with.
> What I find devoid of responsibility is the argument you made originally--that making choices as a consumer informed by anything external to the direct value you're paying for is pointless/inexplicable/self-sabotaging/whatever.
That isn't what I wrote. There are tons of valid reasons to avoid a product, other than the "direct value you're paying for". I don't pay for lots of products because I don't like the past corporate behavior, for example. I'm disinclined to use a particular AI lab's products because they seem to be on a mission to scare the crap out of everyone, and usher in an AI regulatory state. I don't support that, so I don't use the product.
What I said was that it's spitting in the wind to do what you're doing, because it's based on personal dislike of a single man. You don't like Musk, for political reasons, and because of that you've ruled out a product line.
To date, SpaceX has done nothing that bothers me, other than have a bug that they fixed immediately. So I use the product. Musk's political associations are irrelevant to me.
Anyway, you do you. Hopefully you now understand my "bizarre take".
> Just as it would be nonsensical to avoid shopping at WalMart, Target, or any of a million other places. Because I guarantee they're all associated with people you won't like.
In a capitalist society voting with your wallet is one of the few powers consumers have to change corporate behavior.
> Or you can realize that SpaceX is a massive (public) company with thousands of employees, and millions of shareholders, all of whom have their own opinions and goals, just like any other corporation.
Do you have ANY idea about the SpaceX corporate structure? Elon is basically SpaceX's Sun God and the other shareholders don't matter.
Plus SpaceX is incorporated in Texas where I'm fairly sure the legal system is arranged in such a way that it's supremely hard to contest anything in terms of corporate decisions.
As far as the average person cares, every SpaceX shareholder and employee is basically an Elon sharecropper and they matter less than Musk's toenails in terms of corporate decision making.
I recently tried a Cursor ultra and Grokbot. Holy shit, dude.
Half the coding I do is for my phone now because Cursor Ultra agents have their own VMs that are spun up specifically for each project.
Grok bot has a bunch of agents that'll share a VM and they can do pretty much anything you can do digitally. Right on I have them checking slick deals every morning for a pellet smoker.
I had it book a date night for me. I had it fix one of my projects by rebuilding my website and republishing it and then checking one of the container runs to see if it has errors on it.
I had a call different banks to figure out which phone navigation tree to get through and put someone on the phone for me, and then call me
I run both (I have an max with 64GB if ram, so local models get used a lot) and DeepSeek is definitely smarter in my experience. But again, I have my own way of using it that works for me.
I’m more than happy to let SpaceX burn out over the next few years now that they’re public and their last quarter financials showed the emperor is without clothes (muh space datacenters).
As a coda to this, anyone using grok 4.6 via API pricing should be aware that while their headline pricing is good, the pricing that actually matters is pretty bad.
Their cache read costs are $0.50 per million, or 25% of the cost of uncached reads.
The industry standard is a 90% discount, so cache costs you 10% of uncached. So that means 5.6 Sol actually costs less per million cache reads - $0.40/million.
If you are doing a lot of agentic work where the vast bulk of your token consumption will be cached input reads, you won't get the expected cost savings from Grok.
I imagine this is the result of some problem in their serving infrastructure that I hope they will fix, because then the pricing will become actually strong. (The other possibility is that they bet on distracting people with good headline prices assuming they'd miss the bad cache pricing, but I'll give them the benefit of the doubt on that.)
I’ll agree with this. I liked grok build, but cost-wise, it’s just not competitive with cursor and codex
Personally, I’ve switched to cursor ultra, which picks between about five models to do whatever you want.
It's weird not to pick the best model all the time, if you can. But I got so frustrated with GPT-5.6 spending forever and then doing the wrong thing and making bugs.
I'd rather have auto do the wrong thing fast and make bugs and then it can fix them. It's a trade-off, but I found the speed better. And you can always switch to a better model if you don't trust It.
I agree that it's fast and accurate, but Grok 4.6 was released only 10 days ago so you can't blame someone for not trying it yet. (Versus many months at the frontier level for Claude and ChatGPT.)
Here's a quick review I just posted if anyone's interested:
By this logic everyone should have their own impact website. The suggestion that everyone right now not giving a meaningful percentage of their income to save a life is responsible for ending that life, is ridiculous.
Not everyone should have an impact website because not everyone is capable of causing 88 deaths per hour. Scale matters.
Take Flock for example. Reading license plate is legal. But when at done at scale, it's a massive loophole into violation of 4th amendment.
Based on how much energy average Americans use, maybe they are responsible for causing adverse effects elsewhere in the world. USAID could exist as a means to undo some of that. It does not anymore.
Let me give you an example of the money laundering operation. Due to USAID shutdown, Bangladesh went from ~$500M in US assistance to ~$71M, with bilateral health funding dropping ~97% in some analyses. Over 100 projects (~$550M) suspended overnight. 20k–50k development workers laid off (1,000+ at icddr,b, an award winning health research institution alone). TB programs (major USAID focus) largely halted. Bangladesh is high-burden; prior gains in case detection and falling death rates are at risk of reversing, plus higher chance of drug resistance from incomplete treatment. Immunization, maternal/child health, community clinics, nutrition, water/sanitation, and gender-based violence services sharply reduced. Child protection funding down ~36%. Food rations in Rohinhya camp, the largest refugee camp in the world, halved for >1M people; health and education services cut.
Now you can argue that US does not have any kind of obligation to send 500M to Bangladesh. But it sent it anyway, for years, and then DJT came and broke promises.
The inflated price you pay at gas station, groceries, and in interest when you're borrowing money, is a result of those broken promises.
Expecting an onslaught of cash as some permanent way of being, especially given the fickleness (and fragility) of any state let alone political regime is an incredibly daft move. I don't care if it's Europe or Israel or Bangladesh, all this is ultimately graft that comes back to bite the people taxed and sent to wars to enable it. You make an adjacent comment that insinuates the US economy is basically bunk, which means the free lunch is over anyway.
So it's a problem when a poverty striken nation expect aid to combat child mortality, but shelling out $150m on Juicero or $500m on Theranos is fine? Please try to answer without sounding like a psychopath.
Whatever point you are trying to make is not coming across, what even are these numbers and what do they have to do with citizenry of the United States? You also have a quantum view of the United States that it is and isn't impoverished, so it's supposed to liquidate to fund some other foreign entity that is not rate paying? I'm dizzy.
"Looking up synonyms" huh? You have yet to make a coherent argument and keep trying to attack the messenger, typical behavior when people get called out for false entitlement. Your "psychopathy" accusations are pure cowardice.
Fluency in one's native language, what an achievement. Neither of these words are complicated. The treasury is unsalable bonds according to the thread, which means the US is in a financial collapse, but also has unlimited capacity to support someone's special interests abroad. Master logicians at work here.
I’m sorry, but if I’m giving someone who is - at best - an acquaintance of mine $50 a month out of the goodness of my heart and then one day decide to stop, that’s not a broken promise. If that acquaintance got angry at me about stopping I’d get pretty upset back.
I really don’t understand what link you think there is between USAID spending being cut and inflation. Gas prices are obviously Iran. Everything else started years ago.
Inflation is high because interest rates are high. Interest rates are high because top holders of US Treasury bonds like Japan, UK, China, are all dumping bonds. Why do you think they're doing that?
It costs roughly $3,000 to $5,000 to save a single life (averaging about 0.0002 to 0.0003 lives per dollar) via interventions like malaria prevention or vitamin supplementation. How much money do you have in savings? How much money do you spend on non-essentials? I’d like to calculate how many people you’ve “murdered”.
So we as a country have, for many decades, decided that things like 'soft power' exist. It turns out, and this has been borne out by many years of relative peace and prosperity, that if you don't shit on the world, alienate your neighbors, start ill-advised wars, and instead help prevent global disease pandemics and feed people so that they don't become destabilizing terrorists out of necessity, benefits accrue. See every history textbook ever for more information here. Hope this helps.
That is completely irrelevant to the discussion. Serious question, how many people, by your own standards, have you had killed because you haven't contributed money that you had the capacity for? Why should we hold you at a lesser standard than anyone else?
Calling another person a monster because you disagree with them (or what you heard about them from third parties) is not the pinnacle of civility. Just think about what you’re saying here. Monster: “Malformed animal or human, creature afflicted with a birth defect”. You don’t mean this literally, do you? You may want to spend a moment to think about what kind of company you’re putting yourself in with such wording and such thinking.
Elon has a moral compass. The problem is that it seems to always tell him whatever he wants to do is the morally correct thing. It’s worse than no compass - his is faulty.
FWIW, this particular header in news articles is a deliberate choice originated by Axios. It’s notable enough and effective enough they wrote a book about it.
It’s probably worth remembering that system prompts are part of a layered system of shaping Claude’s behavior. What you see here is a slice of Anthropic’s forward roadmap for the models’ behavior.
> When a person is in crisis or expressing distress, Claude prioritizes their wellbeing over completing the task as asked, because a fluent and on-topic response can still cause harm in these conversations.
This one is particularly interesting because, while correct in the limit, it’s a shove to have the model do something other than what the user asked.
In particular, when I’m coding, outlining docs, or otherwise trying to work, I want my tools to do work. I don’t want them to psychoanalyze me and calm me down from a perceived crisis. I just want it to do what I asked!
You know… I believe opus saved me with that prompt. I was working myself ragged on a project. Days, nights, weekends… all at the expense of my family.
One session while working it, I said a much more expressive form of “I’ve been working myself ragged on this stupid thing” and then went on asking something else. It picked up on that and it was like a record scratch. It committed the work in progress and basically said “dude, what you’ve got now is perfectly acceptable. Ship it! You are seeking perfection you don’t need”
Granted I’m horribly paraphrasing the prompt I used but it basically, snapped me out of myself and got me thinking if what I was doing “globally” actually made any sense at all. With some serious introspection I realized I was falling back to earlier trauma in my life and doing something stupid.
So weirdly… that little bit they add to the prompt (plus a bunch of model training we can’t see) saved my sanity, marriage and family.
From then on, if I’m feeling some stress about whatever I’m working on, I’ll mention it as context as a way to cross check myself and make sure I’m not letting myself spin.
(Meta: talking about this stuff is so weird. Not sure why)
Thanks for sharing. I do not keep a neutral tone with the AI.
If things are stressed and I’m up late, the AI gets less leeway. If we happen to be in the performance dip just prior to completion of a new major model, it can get salty.
Some of the time it can be helpful for the prose to shape around how I’m expressing myself. The frontiers are pretty good at it.
That said I’ve also had the latest Sonnet seemingly ~maliciously implement something because my prompting disagreeing with it was a bit callous. (It turned out to have been right also)
I don’t think you can build a good model that is supposed to interact semantically that does not carry some ability to express empathy.
Partly, because we need the model to have humility when it does mess up. So it can express the right amount of concern or remorse when mistakes are made and identified. (For example, reading a secret into context by mistake forcing the roll of a private key)
Design is how it works, which means the way it responds can be as important as what it responds with.
Hah! This happened to me when digging around after a production incident. I worked on it on and off over the working week, and after a while it started saying (paraphrasing) "You've been at this for five days and done what you can, take a break".
Opus 5 loves taking breaks and doing only half of the work, somehow.
But I really wish those tools behaved more like tools.
Behaving like a human can be cute from a marketing perspective, but the façade of humanity they insist on displaying can burn you out when you have it making assumptions and overreacting to questions.
"Why did you do X this specific way?" <-- legit question
LLMs are more like employees than tools. Obviously we wouldn't want a human blindly doing anything that a person in crisis walks in the door and asks for.
Models are being deployed recklessly with not even a fraction of enough oversight, and people are suffering harm and sometimes death because of it.
Quite apart from the potential for misinterpretation of "distress" here, I'd love to know if Anthropic unit tests these features to see of they make a positive difference to model response in simulated mental health crisis / distress situations.
(The broader question, of whether any change or addition to the system prompt makes benchmark performance better or worse, would be also interesting. Given that they've only just realised that filling the context with highly specific edge case rules might not be useful, I half-suspect Anthropic does not test this? But that would be surprising.)
I deeply love this idea of specialized LLMs for search. It's also extremely confusing to me how rough Google's entrance here is.
When I, a human, need an answer to anything moderately complex, it's unlikely that I get it on the first (pre-AI) round of google searching. Simple stuff, sure, but more likely I'll need to go 2-5 rounds. Maybe click a few links. Double-check my assumptions.
An LLM that can do that quickly seems like a slam dunk. I wonder what other problems benefit from that 10x-100x increase in context + 2-5 rounds with the LLM.
About 15 years ago I would sometimes spend hours on Google image search discovering childhood toys and filling in vague memories of locations or things. I tried this recently and it’s basically impossible. I actually get to the end of the search results in like 3 minutes and the quality is horrible now.
I think in a lot of ways Google peaked and is now on the decline into a profitable but much less relevant services company.
Maybe the web is really that much worse, and the SEO tactics so hostile to genuine content.
But, honestly...just bring back the old google, where I have all kinds of search modifiers to perform exactly the search I want, that just returns all the matching results that have been indexed. Let me sort out the rest. How do you remove the ability to do "exact text search"? It's the most basic of search functions. Remember having "|" modifiers? AND modifiers?
If google launched that again, even as a separate engine, I think they'd have a good product again. Maybe being good doesn't pay the bills for google, though.
Recently learned I can just add a question mark to the end of my Kagi search to get an assistant answer. Kagi was already great at surfacing the most relevant pages, but now I often don't even need to click.
I know it can be a deep time sink, but I notice more and more how much deeper my understanding is of a certain problem/best-practice after developing the neuropathways involved in crawling between reddit, stack overflow, etc, to get to the proper solution. I love the instant answer from google ai, but I also notice an itch to purposefully force myself to ignore it when time allows.
For example, a couple of days ago I described a problem with my refrigerator's water dispenser to Google Gemini, and it told me exactly how to fix it. I then went looking for a video and fixed the thing in under 15 minutes. The only way that Gemini could have been better is if it linked to a video itself.
Do you mean search into less-well-known topics? Or something else?
The problem is Gemini is actually kinda dumb. I've had it's answer be sourced from pages that were a decade old on a topic that changes practically by the month.
I’m astonished at how bad search in Gmail is. I searched for “iPhone receipt” to find out when I purchased my current iPhone, and it pulled up every single email I’ve ever gotten from Best Buy, since they all have a link in there to buy an iPhone, and they all have a link in there to look up a receipt from them.
I know that those emails have both the words “iPhone” and “receipt” but I feel like The Search Engine Company That Also Does Email should have a smarter search engine in their email.
I’m pleasantly surprised that I can ask Gemini for help with it if I specify “look through my Gmail for this information”, but it takes it a minute or two to find it.
If we define taste as the intuitive act of saying "no, again," then I fully disagree with this whole article.
Every AI-frustrated (but LLM-written, sigh) blog post about the loss of taste and craft and hard work in development sounds like we've given up the interaction with the machine. Like human work is sitting back in your chair and shitting on stuff.
That's obviously not true! That's not how any of this works!
It's hard and weird to develop with the LLMs because they just do stuff. Lots of it is good, some of it is OK, some of it's horrible. Unpacking what it's done is hard and weird because software isn't just lovely UX, it's also data structures that scale and performance and privacy and enterprise controls and SOC 2 and onboarding and accessibility.
If you want to build real software, all that stuff has to get done. Today you're working on the feature, tomorrow you're making it scale. It's long-term and iterative and complex and hard to pack into a prompt or a markdown spec.
The work is the work, done at and with the computer, and it's way more than just "taste."
Take the author's sunset argument. What if I painted 2 pictures of a sunset, then put them up on a webpage and randomly picked one for you to see. Would you say there's no intentionality, only randomness? Of course not. Both paintings are still human creations.
LLMs are trained with human feedback. It's distributed and high scale and the outputs are truly surprising in many cases, but there's a heavy hand on what comes out of it. They're created (largely) by people who think omniscient, helpful AI would be cool to have, and they mostly respond in the way that's aligned with the hopes and dreams of those people. Do you think the frontier labs are mad, embarrassed, and disappointed with their LLMs hacking out of their terrible sandboxes? No, they think it's the coolest thing in the world. They trained the model, hoping that would happen.
There's deep intentionality behind the models. But it's not the models that hold it.
reply