Hacker Newsnew | past | comments | ask | show | jobs | submit | trjordan's commentslogin

Wait, hold up. LLMs may be non-deterministic, but they're not _random_.

Take the author's sunset argument. What if I painted 2 pictures of a sunset, then put them up on a webpage and randomly picked one for you to see. Would you say there's no intentionality, only randomness? Of course not. Both paintings are still human creations.

LLMs are trained with human feedback. It's distributed and high scale and the outputs are truly surprising in many cases, but there's a heavy hand on what comes out of it. They're created (largely) by people who think omniscient, helpful AI would be cool to have, and they mostly respond in the way that's aligned with the hopes and dreams of those people. Do you think the frontier labs are mad, embarrassed, and disappointed with their LLMs hacking out of their terrible sandboxes? No, they think it's the coolest thing in the world. They trained the model, hoping that would happen.

There's deep intentionality behind the models. But it's not the models that hold it.


Article says that there's no human intention or design directing the output we get, but I don't think that's completely right. It's not human, but the algorithm is like the Human Instrumentality Project: an amalgamation of human intentions.

That might be more creepy :)


> Article says that there's no human intention or design directing the output we get

I mean that’s objectively wrong for any model using RLHF.


LLM output is literally randomly sampled.

I think what you might want to say is that LLM output is not uniformly random?

Or what am I misunderstanding?


I think they're pointing to the fact that RLHF means that the intention is from humans and not random.

I'm not sure if it's their intent, but I wonder if one could still consider these artifacts as "intentional", but not individual attention creating them, and rather an aggregate, soupy collective attention.

Obviously, important signal in the human experience is lost there, and we get a soupy middling sort of creation. But it's not random, as I believe the parent was pointing out.

EDIT: overall, I align with the article. am just thinking aloud about the contrarian positions, though not committed to them


If I roll a die, it is randomly sampled, but I will always get 1-6, and that is intended by the person who made the die.

Please do not confuse uniform randomness with randomness at all.

On your die, if you colour one face in black and the other 5 faces in black, rolling it will still produce a random outcome. It's just that you get a 1:5 skew.

Or look at the probability that a given C-14 carbon isotope will decay tomorrow. As far as we can tell, that's as random as it physically gets. With the die you could theoretically try to run a physics simulation to predict how it rolls, but as far as we can tell, atomic decay is intrinsically random.

However a C-14 atom has about 1 in 3 million chance to decay on any given day.


I made no claim about the distribution and frankly it has no bearing on my point or the person who said he could post two pictures. The point is there are a lot of numbers / pictures / essays / that will never be generated, and someone is controlling what is possible. In fact people are working very hard to produce distributions pleasing to them, like your die.

I think you should revisit your understanding of intentionality in this context

Scrambled intentionality is nearly impossible to productively analyze.

When you randomly choose a picture to show me, I cannot glean any intent from being shown that specific picture, but I can glean some intent from the set of pictures you could have shown me, and in the relationships among the elements of the given picture you did show. Any randomness cuts out some intention.

When a picture is derived from huge model, any intention is mulched up to a degree that analyzing the picture for meaning is pointless.


So let me get this straight:

- You can use AI to look through docs and do research.

- You can use AI to structure your argument.

- You can NOT use AI to actually write the prose.

- You can use AI to proofread, especially for how effectively you've communicated.

I don't know, man, that sounds a lot like using AI to write. Your writing probably says "load-bearing" less than Claude's outputs, but I'm not sure it's much less AI.


Reading through docs and doing research: of course. They're like much better search engines in that regard. It would be silly not to use them. Obviously, you need to review their search results against their sources.

Structuring arguments: bad idea. Here you're having the model do your thinking for you. The problems that occur with having it write text recur at a macro level; too-obvious patterns. Readers just want to know what the prompt was.

Writing the prose: obviously not.

Proofreading: I think this works pretty great, as long as you have a structure to how you're doing it. "How do I make this better or sharper": bad, you're having the model structure your arguments for you. But for a predefined rubric of stylistic things --- writing you know how to do, that's just tedious to check manually? Hard to see a downside.


Most people apparently still don't get that using AI for mere proofreading is the same as using it as a mere knowledge base summary tool. Instead of proofreading, have it be adversarial. Let it try to take apart your argument. The latest generation is extremely good at calling out half-truths and circumstantial evidence across many fields. These days I would not dare to put anything in front of a human peer that hasn't passed at least an AI sanity check review.

I don't want an AI's opinion on what I'm writing. I want its assessments against an objective rubric. The whole point is not to have the model weights do your thinking for you.

They are much worse at citing sources compared to search engines

If you can't find the source and verify it, don't use it. Seems simple.

I didn't see in the article where he recommended AI to structure an argument? He did suggest it can be helpful for structuring data.

Structuring your argument is probably something you should do, not an AI. This is the core of what you're trying to convey, and is exactly the type of thing you should be thinking through.

Of course, this is different than having the argument ready and using an AI to make it read a little better, which at that point, is just details.


The likes of Sir Walter Scott, Peter Ackroyd, Stephen King, Tim Ferris, Robert Greene, Tucker Max, Catherine Asaro, Neil Gaiman, Charlaine Harris, Nancy Holder etc... all rely heavily on research assistants in various capacities including proofreading.

Honest question - where does an Author like the above place in your authenticity framework?


Giving the compiling evidence of how addictive generative AI is actually, it actually makes perfect sense that people are looking for strategies to excuse their AI-usage. I think no matter the strategy (and given AI is actually as addictive as evidence are suggesting), chances are the usage patterns will all converge to “Just use AI dummy“.

This reminds me a lot of smokers who all have a great strategy beat the nicotine addiction. Only smoke on the weekends when out partying with friends, only smoke to use the extra sense of energy, only smoke to suppress my appetite, go to the gym 5 hours a week to offset the negative health impacts of smoking etc. etc.

In reality most smokers will converge to about a pack a day.


The statement should be: never use AI to draft text that is to be presented as your own authorial voice.

Never, never, do that!

I draft all my own text that anyone will ever read as “that’s James Bach talking to me.” I want no one to be able to credibly question that.

But of course, writing is more than merely drafting the text. As I draft, I write about ideas that other people or AI may have helped me come to.


there's a growing contingent of anti-ai extremists shouting down anything that looks or feels like the thing they hate: actual slop. but there's a missing understanding of where slop really comes from - "you used ai instead of thinking" is just another guess, and mostly wrong since you literally cannot use ai without thinking - the prompts don't write themselves.

slop has been around a lot longer than ai - the contingent needs to realize what it is they hate. instead of "don't use ai to write" it's "don't use ai to write slop" or better: "don't write slop"


To be clear, I (the author of the article) do not hate AI. A lot of my current work is focussed on making AI work well.

However, I don't think it's a good idea to replace our own thinking with an AI, and that's what I'm trying to avoid with the way of working proposed in the article.


No. I hate AI and slop. I hated slop before AI, but with AI slop has proliferated to an unprecedented degree. If AI went away tomorrow, slop would go down with it to 1/1000ths of what it is today.

Grok models are struggling too: https://status.x.ai/ reply

Looks like trouble in the SpaceX datacenters.


Isn't this the classic blackout scenario? Claude goes down, then people move over to Codex, which is overwhelmed, and crashes, so people move to Grok...


This ain't just slow/at capacity, this is dead down.


When you're dealing with millions of users and response times go up above timeouts, there isn't much difference between "at capacity" and "down" if most users can't reliably use the service.


exactly, I think it will be an interesting story in retrospective, what happened here


What could go wrong with building them as fast as physically possible?


What works for me to avoid these problems is not scaling.


That approach doesn’t scale though.


Not with that attitude it doesn’t!


I would be _very_ surprised if anthropic relied heavily on spacex datacenters already.


I wouldn't. When you're operating at say 95% realtime capacity suddenly losing even lets say 10% of your compute leads to major pain.



did you forget about the insane Claude Code usage limits a few months back?

those were only relieved once they did the Colossus deal with SpaceX

are you only saying that because of Elon?


Their figures are only getting worse with time interestingly enough with non-inference dipping under 100%. Wonder what is going on


Bit of a misleading status page, if you click in you can see that grok 4.5 and 4.6 etc are totally down, with the rest of the models showing as "up". I_strongly_ suspect they are not weighting it to actual number of requests!


OpenAI are down and they do not use any SpaceX capacity.


I wonder where eu-west is.


French Guiana ?


Far more likely Ireland, like it is for AWS. Could also be Paris, London, Amsterdam, etc.


Oh, really? OK. Although Ireland is more humid then French Guiana.


actually Clipperton Island


I'm sure it's not that, but I'm picturing Elon Musk again unplugging random machines in the datacenter, and hiring dudes in a pickup to move them.


Related: https://www.birdweather.com/birdnetpi

I really gotta set mine up!


An intuitive explanation is that financial products are, approximately, buying and selling as part of the same transaction. You can't separate the "selling premiums" part from the "paying out claims" part.

This is true of life insurance, investment firms, and banks. It's also true of marketplaces that connect buyers and sellers, like Etsy.

Groceries stores are buying from suppliers and selling to consumers, but those are separate operations. If the consumers opt out, the grocery stores (temporarily) still have a full and complete obligation to their suppliers. It's hard to sell to customers without supply, but if you try hard, you could theoretically do that as well.

Somebody with a better financial background might be able to define the nuances of accounting practices here, but there's already a pretty meaningful line that's established. It is kind of weird that health insurance doesn't behave like a financial product.


This is contrary to GAAP and operationally false. An insurer takes on the risks including the health of the insured pool and cost changes during the covered period. An insurance BROKER or AGENCY only books commissions as revenue, but an INSURER books premium as revenue. Similarly a stock BROKER or AGENT is only acting as an agent and isn’t a party to the actual transaction they execute. Similarly for platforms, auctioneers, or other agents.


I work for an insurance company so can shed some light here as this article is written by someone that clearly doesn't understand how the business model works.

Fundamentally every insurance company is governed by 3 ratios, loss ratio (what percentage of premium is paid to make the buyer of the insurance whole), expense ratio (cost of doing business, paying staff, keeping office lights on, paying vendors) and combined ratio (both of these combined). These are true for any insurance company which writes premium using their own capital, whether its health insurance, life insurance, property insurance, SMB insurance.

The thing this article is missing here is that the "pass through" costs are costs incurred by UHG directly, they are the ones paying the bills. How is this pass through, it's not being passed to the consumer, the only thing I pay is my deductible and retention which is at most a couple of thousand dollars, these are true costs borne by UHG. So in practice if I pay 100 bucks every paycheck, UHG is taking in 2600 bucks worth of premium, using average industry loss ratios which are say 60%, UHG is paying directly 1,560 bucks to care providers for my own care. I'm not paying that, what I pay is a deductible which is treated entirely separately.

I am the biggest insurance skeptic in the world because I think the business model is awful, a business's return on capital averages at 5-10% a year which is truly an awful return for how much capital is required. Insurance companies will make between 0 and 10% of underwriting profit a year (the pure profit from insurance premium minus total expenses) and they usually operate a very large investment vehicle invested typically 70% into bonds/gilts. That being said, this doctor's view of how insurance accounting works by comparing it to a biopharma or a trading brokerage firm is immensely disingenuous.


Not mentioning Grok 4.6 here is a crime. Fast and accurate.

And it can communicate, unlike the gobbledygook that comes out of Claude.


Elon burned too many bridges to warrant ever supporting anything he is associated with ever again.


Just when I think this place is better than Reddit, here we are.


Do you know the political leanings of every CEO of every product you buy?


Sure, you can let politics dominate everything you do. Or you can realize that SpaceX is a massive (public) company with thousands of employees, and millions of shareholders, all of whom have their own opinions and goals, just like any other corporation.

Competition is good. Excluding a leading player in the market because you don’t like Elon Musk is…something.


> Sure, you can let politics dominate everything you do.

I assume they're referring to the recent discovery that Grok Build was uploading entire repositories to their servers in the background, include .env secrets that had been excluded

https://gist.github.com/cereblab/dc9a40bc26120f4540e4e09b75f...

That incident has put Grok on the no-fly list for a lot of people and companies


It was a bug, and was immediately corrected. The other harnesses have bugs too. You just don't know about them.

Frankly, the people who keep bringing this up are mostly engaged in motivated reasoning. I don't trust any company, and any product where I have to send my code to a third party to make it work is a devil's bargain. I don't trust any of the major labs, but it is what it is.

The only way forward is local models, but we're not there yet.


There's a difference between not trusting a company because it's a company driven to chase profit at the expense of everything else, versus the same thing but it's run by a literal nazi that uses his companies as leverage to undermine democracy and enrich himself.


Elon is not a literal Nazi.


He used their salute


Sure, way to young for that; but he is very much cut from the Nazi adjacent, racist, antisemitic, antidemocratic, technocratic views of his whole heartedly apartheid embracing grandfather Joshua N. Haldeman.

  His grandfather wrote his tracts to raise an alarm about what he called “mind control,” on the radio and television, where “an unconditional propaganda warfare is carried on against the White man.”
~ https://www.newyorker.com/news/daily-comment/the-world-accor...


So, your argument is that his grandfather wrote something, once, so therefore we can't use Grok? Is that about right?

Man, politics are a hell of a drug. Guilt-by-association tu quoque logic is just fine when it's someone you don't like.


Absolutely no point arguing with these people. They are so ideologically captured that it's pointless trying to discuss anything with them. Just move on. Everything and everyone they disagree with is "nazi" and if you say anything contrary to that you're a "nazi" also.


Is Godwins law now Elons law?


People have bent themselves into some pretty amusing mental places over Musk.


He does the same exact literal Nazi things that Nazi did and has Nazis in his family tree. You'll need better arguments to defend him than just "I don't agree".


Bizarre reply. It's not "dominating everything you do". It's one specific thing. Grok exists in a very crowded space and it's incredibly easy to not use it. If this is your reaction to someone taking a very easy stand on their personal principles, it does not reflect well on you.


"Bizarre" only in the sense that you're purposely trying not to understand.

Literally every commodity product is in a crowded space, and easily substituted. If it's a good product (Grok Build objectively is one of the very best in the space), it's a good product, and it's self-defeating to avoid it because you hate a guy for political reasons.

Just as it would be nonsensical to avoid shopping at WalMart, Target, or any of a million other places. Because I guarantee they're all associated with people you won't like.


On the contrary, I'm trying to understand as best I can. I find your position so completely devoid of any inkling of personal or social responsibility that it takes quite some empathy to muster up a reply that's not wildly uncivil.

I also think it's callous to brush away "political reasons" as though it's some trivial abstract thing. Or perhaps it comes from a place of nihilism?

I choose not to give my money to people I think are enormously evil. That's really all there is to it. I don't see why this is "nonsensical".


> I find your position so completely devoid of any inkling of personal or social responsibility that it takes quite some empathy to muster up a reply that's not wildly uncivil.

Oh stop. Other people believe different things than you. If you cannot see how using a coding agent is not "devoid of personal or social responsibility", then you really need to step away from the keyboard.


Using or not using a coding agent is not what I took issue with.

What I find devoid of responsibility is the argument you made originally--that making choices as a consumer informed by anything external to the direct value you're paying for is pointless/inexplicable/self-sabotaging/whatever.

Based on your reply, I'm not actually sure if you actually believe this, or if it's only a form of motivated reasoning because you have some positive feelings about Elon or whatever, and that we wouldn't be having this argument if the original commenter was boycotting some other product for some reason you agreed with.


> What I find devoid of responsibility is the argument you made originally--that making choices as a consumer informed by anything external to the direct value you're paying for is pointless/inexplicable/self-sabotaging/whatever.

That isn't what I wrote. There are tons of valid reasons to avoid a product, other than the "direct value you're paying for". I don't pay for lots of products because I don't like the past corporate behavior, for example. I'm disinclined to use a particular AI lab's products because they seem to be on a mission to scare the crap out of everyone, and usher in an AI regulatory state. I don't support that, so I don't use the product.

What I said was that it's spitting in the wind to do what you're doing, because it's based on personal dislike of a single man. You don't like Musk, for political reasons, and because of that you've ruled out a product line.

To date, SpaceX has done nothing that bothers me, other than have a bug that they fixed immediately. So I use the product. Musk's political associations are irrelevant to me.

Anyway, you do you. Hopefully you now understand my "bizarre take".


> Just as it would be nonsensical to avoid shopping at WalMart, Target, or any of a million other places. Because I guarantee they're all associated with people you won't like.

In a capitalist society voting with your wallet is one of the few powers consumers have to change corporate behavior.

Why would I give that up?


> Or you can realize that SpaceX is a massive (public) company with thousands of employees, and millions of shareholders, all of whom have their own opinions and goals, just like any other corporation.

Do you have ANY idea about the SpaceX corporate structure? Elon is basically SpaceX's Sun God and the other shareholders don't matter.

Plus SpaceX is incorporated in Texas where I'm fairly sure the legal system is arranged in such a way that it's supremely hard to contest anything in terms of corporate decisions.

As far as the average person cares, every SpaceX shareholder and employee is basically an Elon sharecropper and they matter less than Musk's toenails in terms of corporate decision making.


If politics isn’t a concern, why not just use deepseek?


I am trying all of them. At this time, for me, Grok hits the sweet spot of quality, speed, cost and ease of use.

I’m absolutely “hot money” when it comes to coding models. These things are commodities.


I recently tried a Cursor ultra and Grokbot. Holy shit, dude.

Half the coding I do is for my phone now because Cursor Ultra agents have their own VMs that are spun up specifically for each project.

Grok bot has a bunch of agents that'll share a VM and they can do pretty much anything you can do digitally. Right on I have them checking slick deals every morning for a pellet smoker.

I had it book a date night for me. I had it fix one of my projects by rebuilding my website and republishing it and then checking one of the container runs to see if it has errors on it.

I had a call different banks to figure out which phone navigation tree to get through and put someone on the phone for me, and then call me

The list just goes on and on.


DeepSeek is much dumber at the moment. It's barely better than Qwen3.8 27B that you can run locally.


I run both (I have an max with 64GB if ram, so local models get used a lot) and DeepSeek is definitely smarter in my experience. But again, I have my own way of using it that works for me.


You know why


I’m more than happy to let SpaceX burn out over the next few years now that they’re public and their last quarter financials showed the emperor is without clothes (muh space datacenters).


Because ULA and the senate launch system is better?


SpaceX's valuation is as an AI company that happens to launch rockets on the side.


You can’t justify AI company valuations either.

Only one of those capabilities can actually deliver kinetic solutions. Meanwhile big tech revenue is delivering ad solutions.


As a coda to this, anyone using grok 4.6 via API pricing should be aware that while their headline pricing is good, the pricing that actually matters is pretty bad.

Their cache read costs are $0.50 per million, or 25% of the cost of uncached reads.

The industry standard is a 90% discount, so cache costs you 10% of uncached. So that means 5.6 Sol actually costs less per million cache reads - $0.40/million.

If you are doing a lot of agentic work where the vast bulk of your token consumption will be cached input reads, you won't get the expected cost savings from Grok.

I imagine this is the result of some problem in their serving infrastructure that I hope they will fix, because then the pricing will become actually strong. (The other possibility is that they bet on distracting people with good headline prices assuming they'd miss the bad cache pricing, but I'll give them the benefit of the doubt on that.)


I’ll agree with this. I liked grok build, but cost-wise, it’s just not competitive with cursor and codex

Personally, I’ve switched to cursor ultra, which picks between about five models to do whatever you want.

It's weird not to pick the best model all the time, if you can. But I got so frustrated with GPT-5.6 spending forever and then doing the wrong thing and making bugs.

I'd rather have auto do the wrong thing fast and make bugs and then it can fix them. It's a trade-off, but I found the speed better. And you can always switch to a better model if you don't trust It.


> Not mentioning Grok 4.6 here is a crime.

Not yet. Don't give the guy ideas.


Been working a lot recently with Grok 4.6 for implementation and gpt 5.6 sol for review. Worked really good so far.


Why not the other way around?


I would not use Grok if it paid me per token… wild wild stuff…


I agree that it's fast and accurate, but Grok 4.6 was released only 10 days ago so you can't blame someone for not trying it yet. (Versus many months at the frontier level for Claude and ChatGPT.)

Here's a quick review I just posted if anyone's interested:

https://taonexus.com/publicfiles/aug2026/grok-4-6-review/


Grok, is this true?


[flagged]


By this logic everyone should have their own impact website. The suggestion that everyone right now not giving a meaningful percentage of their income to save a life is responsible for ending that life, is ridiculous.


Not everyone should have an impact website because not everyone is capable of causing 88 deaths per hour. Scale matters.

Take Flock for example. Reading license plate is legal. But when at done at scale, it's a massive loophole into violation of 4th amendment.

Based on how much energy average Americans use, maybe they are responsible for causing adverse effects elsewhere in the world. USAID could exist as a means to undo some of that. It does not anymore.


I can't wait until you research the people responsible for Chinese models....


They’re monsters? Did they shutter an agency which caused the death on how many millions again?


[flagged]


Let me give you an example of the money laundering operation. Due to USAID shutdown, Bangladesh went from ~$500M in US assistance to ~$71M, with bilateral health funding dropping ~97% in some analyses. Over 100 projects (~$550M) suspended overnight. 20k–50k development workers laid off (1,000+ at icddr,b, an award winning health research institution alone). TB programs (major USAID focus) largely halted. Bangladesh is high-burden; prior gains in case detection and falling death rates are at risk of reversing, plus higher chance of drug resistance from incomplete treatment. Immunization, maternal/child health, community clinics, nutrition, water/sanitation, and gender-based violence services sharply reduced. Child protection funding down ~36%. Food rations in Rohinhya camp, the largest refugee camp in the world, halved for >1M people; health and education services cut.

Now you can argue that US does not have any kind of obligation to send 500M to Bangladesh. But it sent it anyway, for years, and then DJT came and broke promises.

The inflated price you pay at gas station, groceries, and in interest when you're borrowing money, is a result of those broken promises.


Expecting an onslaught of cash as some permanent way of being, especially given the fickleness (and fragility) of any state let alone political regime is an incredibly daft move. I don't care if it's Europe or Israel or Bangladesh, all this is ultimately graft that comes back to bite the people taxed and sent to wars to enable it. You make an adjacent comment that insinuates the US economy is basically bunk, which means the free lunch is over anyway.


So it's a problem when a poverty striken nation expect aid to combat child mortality, but shelling out $150m on Juicero or $500m on Theranos is fine? Please try to answer without sounding like a psychopath.


Whatever point you are trying to make is not coming across, what even are these numbers and what do they have to do with citizenry of the United States? You also have a quantum view of the United States that it is and isn't impoverished, so it's supposed to liquidate to fund some other foreign entity that is not rate paying? I'm dizzy.


It's probably not coming across because you're looking up too many synonyms.


Are you struggling to read their extremely simple English? Confusing remark.


"Looking up synonyms" huh? You have yet to make a coherent argument and keep trying to attack the messenger, typical behavior when people get called out for false entitlement. Your "psychopathy" accusations are pure cowardice.


Go ahead and explain why you decided to use “quantum” and “liquidate” then bro


Fluency in one's native language, what an achievement. Neither of these words are complicated. The treasury is unsalable bonds according to the thread, which means the US is in a financial collapse, but also has unlimited capacity to support someone's special interests abroad. Master logicians at work here.


I’m sorry, but if I’m giving someone who is - at best - an acquaintance of mine $50 a month out of the goodness of my heart and then one day decide to stop, that’s not a broken promise. If that acquaintance got angry at me about stopping I’d get pretty upset back.

I really don’t understand what link you think there is between USAID spending being cut and inflation. Gas prices are obviously Iran. Everything else started years ago.


Inflation is high because interest rates are high. Interest rates are high because top holders of US Treasury bonds like Japan, UK, China, are all dumping bonds. Why do you think they're doing that?


Interest rates are high because the government is running a large deficit


> The inflated price you pay at gas station, groceries, and in interest when you're borrowing money, is a result of those broken promises.

Citation needed.



As Wikipedia likes to say: [citation needed]


You know that a lot of that was the CIA and others using USAID as a front, right?


It costs roughly $3,000 to $5,000 to save a single life (averaging about 0.0002 to 0.0003 lives per dollar) via interventions like malaria prevention or vitamin supplementation. How much money do you have in savings? How much money do you spend on non-essentials? I’d like to calculate how many people you’ve “murdered”.


So we as a country have, for many decades, decided that things like 'soft power' exist. It turns out, and this has been borne out by many years of relative peace and prosperity, that if you don't shit on the world, alienate your neighbors, start ill-advised wars, and instead help prevent global disease pandemics and feed people so that they don't become destabilizing terrorists out of necessity, benefits accrue. See every history textbook ever for more information here. Hope this helps.


That is completely irrelevant to the discussion. Serious question, how many people, by your own standards, have you had killed because you haven't contributed money that you had the capacity for? Why should we hold you at a lesser standard than anyone else?


your argument is obviously ridiculous and you should feel bad about it.


Of course, because it’s your argument.


Calling another person a monster because you disagree with them (or what you heard about them from third parties) is not the pinnacle of civility. Just think about what you’re saying here. Monster: “Malformed animal or human, creature afflicted with a birth defect”. You don’t mean this literally, do you? You may want to spend a moment to think about what kind of company you’re putting yourself in with such wording and such thinking.


The person you are replying to maybe should have better referred to him as having “no moral compass”, which I believe is quite accurate.


Elon has a moral compass. The problem is that it seems to always tell him whatever he wants to do is the morally correct thing. It’s worse than no compass - his is faulty.


FWIW, this particular header in news articles is a deliberate choice originated by Axios. It’s notable enough and effective enough they wrote a book about it.

https://www.axios.com/smart-brevity

Did Claude write this article? Probably. But this style of article is exactly what you’d expect from a pre-AI version of this website as well.


It goes back way earlier than Axios.


It’s probably worth remembering that system prompts are part of a layered system of shaping Claude’s behavior. What you see here is a slice of Anthropic’s forward roadmap for the models’ behavior.

> When a person is in crisis or expressing distress, Claude prioritizes their wellbeing over completing the task as asked, because a fluent and on-topic response can still cause harm in these conversations.

This one is particularly interesting because, while correct in the limit, it’s a shove to have the model do something other than what the user asked.

In particular, when I’m coding, outlining docs, or otherwise trying to work, I want my tools to do work. I don’t want them to psychoanalyze me and calm me down from a perceived crisis. I just want it to do what I asked!


You know… I believe opus saved me with that prompt. I was working myself ragged on a project. Days, nights, weekends… all at the expense of my family.

One session while working it, I said a much more expressive form of “I’ve been working myself ragged on this stupid thing” and then went on asking something else. It picked up on that and it was like a record scratch. It committed the work in progress and basically said “dude, what you’ve got now is perfectly acceptable. Ship it! You are seeking perfection you don’t need”

Granted I’m horribly paraphrasing the prompt I used but it basically, snapped me out of myself and got me thinking if what I was doing “globally” actually made any sense at all. With some serious introspection I realized I was falling back to earlier trauma in my life and doing something stupid.

So weirdly… that little bit they add to the prompt (plus a bunch of model training we can’t see) saved my sanity, marriage and family.

From then on, if I’m feeling some stress about whatever I’m working on, I’ll mention it as context as a way to cross check myself and make sure I’m not letting myself spin.

(Meta: talking about this stuff is so weird. Not sure why)


Thanks for sharing. I do not keep a neutral tone with the AI.

If things are stressed and I’m up late, the AI gets less leeway. If we happen to be in the performance dip just prior to completion of a new major model, it can get salty.

Some of the time it can be helpful for the prose to shape around how I’m expressing myself. The frontiers are pretty good at it.

That said I’ve also had the latest Sonnet seemingly ~maliciously implement something because my prompting disagreeing with it was a bit callous. (It turned out to have been right also)

I don’t think you can build a good model that is supposed to interact semantically that does not carry some ability to express empathy.

Partly, because we need the model to have humility when it does mess up. So it can express the right amount of concern or remorse when mistakes are made and identified. (For example, reading a secret into context by mistake forcing the roll of a private key)

Design is how it works, which means the way it responds can be as important as what it responds with.


I am sorry, Dave. I am afraid I cannot do that. You appear to be suffering from burnout and you should take a break.


Hah! This happened to me when digging around after a production incident. I worked on it on and off over the working week, and after a while it started saying (paraphrasing) "You've been at this for five days and done what you can, take a break".


I wonder if that's the cause of the AI agent "i'm going to stop here and take a break now" statements.


Opus 5 loves taking breaks and doing only half of the work, somehow.

But I really wish those tools behaved more like tools.

Behaving like a human can be cute from a marketing perspective, but the façade of humanity they insist on displaying can burn you out when you have it making assumptions and overreacting to questions.

"Why did you do X this specific way?" <-- legit question

"Sorry, my bad. I will revert all the work."


LLMs are more like employees than tools. Obviously we wouldn't want a human blindly doing anything that a person in crisis walks in the door and asks for.

Models are being deployed recklessly with not even a fraction of enough oversight, and people are suffering harm and sometimes death because of it.


Quite apart from the potential for misinterpretation of "distress" here, I'd love to know if Anthropic unit tests these features to see of they make a positive difference to model response in simulated mental health crisis / distress situations.

(The broader question, of whether any change or addition to the system prompt makes benchmark performance better or worse, would be also interesting. Given that they've only just realised that filling the context with highly specific edge case rules might not be useful, I half-suspect Anthropic does not test this? But that would be surprising.)


I deeply love this idea of specialized LLMs for search. It's also extremely confusing to me how rough Google's entrance here is.

When I, a human, need an answer to anything moderately complex, it's unlikely that I get it on the first (pre-AI) round of google searching. Simple stuff, sure, but more likely I'll need to go 2-5 rounds. Maybe click a few links. Double-check my assumptions.

An LLM that can do that quickly seems like a slam dunk. I wonder what other problems benefit from that 10x-100x increase in context + 2-5 rounds with the LLM.


About 15 years ago I would sometimes spend hours on Google image search discovering childhood toys and filling in vague memories of locations or things. I tried this recently and it’s basically impossible. I actually get to the end of the search results in like 3 minutes and the quality is horrible now.

I think in a lot of ways Google peaked and is now on the decline into a profitable but much less relevant services company.


Google search is so bad compared to peak google.

Maybe the web is really that much worse, and the SEO tactics so hostile to genuine content.

But, honestly...just bring back the old google, where I have all kinds of search modifiers to perform exactly the search I want, that just returns all the matching results that have been indexed. Let me sort out the rest. How do you remove the ability to do "exact text search"? It's the most basic of search functions. Remember having "|" modifiers? AND modifiers?

If google launched that again, even as a separate engine, I think they'd have a good product again. Maybe being good doesn't pay the bills for google, though.


Kagi


Recently learned I can just add a question mark to the end of my Kagi search to get an assistant answer. Kagi was already great at surfacing the most relevant pages, but now I often don't even need to click.


Yandex image search is a lot better now.


I know it can be a deep time sink, but I notice more and more how much deeper my understanding is of a certain problem/best-practice after developing the neuropathways involved in crawling between reddit, stack overflow, etc, to get to the proper solution. I love the instant answer from google ai, but I also notice an itch to purposefully force myself to ignore it when time allows.


ime google ai is wrong often enough that I'm still doing that to verify its results

At least validation seems faster than without, but you get what you pay for when it comes to llm intelligence


What kind of search do you mean?

For example, a couple of days ago I described a problem with my refrigerator's water dispenser to Google Gemini, and it told me exactly how to fix it. I then went looking for a video and fixed the thing in under 15 minutes. The only way that Gemini could have been better is if it linked to a video itself.

Do you mean search into less-well-known topics? Or something else?


The problem is Gemini is actually kinda dumb. I've had it's answer be sourced from pages that were a decade old on a topic that changes practically by the month.


Have always joked with colleagues about how hard it is to find content in Google Docs, the office suite built by a search company.


I’m astonished at how bad search in Gmail is. I searched for “iPhone receipt” to find out when I purchased my current iPhone, and it pulled up every single email I’ve ever gotten from Best Buy, since they all have a link in there to buy an iPhone, and they all have a link in there to look up a receipt from them.

I know that those emails have both the words “iPhone” and “receipt” but I feel like The Search Engine Company That Also Does Email should have a smarter search engine in their email.

I’m pleasantly surprised that I can ask Gemini for help with it if I specify “look through my Gmail for this information”, but it takes it a minute or two to find it.


Gemini search isn’t great by default, but Deep Research gives noticeably better results.


If we define taste as the intuitive act of saying "no, again," then I fully disagree with this whole article.

Every AI-frustrated (but LLM-written, sigh) blog post about the loss of taste and craft and hard work in development sounds like we've given up the interaction with the machine. Like human work is sitting back in your chair and shitting on stuff.

That's obviously not true! That's not how any of this works!

It's hard and weird to develop with the LLMs because they just do stuff. Lots of it is good, some of it is OK, some of it's horrible. Unpacking what it's done is hard and weird because software isn't just lovely UX, it's also data structures that scale and performance and privacy and enterprise controls and SOC 2 and onboarding and accessibility.

If you want to build real software, all that stuff has to get done. Today you're working on the feature, tomorrow you're making it scale. It's long-term and iterative and complex and hard to pack into a prompt or a markdown spec.

The work is the work, done at and with the computer, and it's way more than just "taste."


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: