I see this sentiment pretty regularly, and I don’t get it. Variable rewards is not sufficient to establish that it is “ basically gambling”.
Everything in life is variable reward. You invite a friend over, they might accept or they might not. Drive to work, traffic might be good or might be bad. You ask a colleague to finish a task, they might do it or might not or might do a good job or might not.
Everything is variable reward. Is everything gambling?
> "Everything is variable reward. Is everything gambling?"
well, no. If you work overtime and get paid overtime, you are not gambling and that is not a variable reward.
Humans engage more with rewards that are intermittent and variable. Like Futurama's scene from 'The Scary Door' where the character says "A casino where I'm winning, I must be in heaven! A casino where I always win, that's boring, I must really be IN HELL!". A constant predictable reward is boring, less engaging. So if you know you get no overtime, but sometimes your boss rewards you with $5 coffee voucher, sometimes a free pizza dinner, sometimes double-time pay for the time worked or a half-day off, now you might be gambling 1hr overtime for an intermittent variable reward.
> "Drive to work, traffic might be good or might be bad."
Good traffic is not a "reward" for driving to work(!) and you have to drive to work regardless so you are not risking anything [you might be risking your life, but you are not making a choice which can reward you with good traffic]. You might say that going a different route is a choice and a gamble which could reward you with good traffic, but traffic engineering does not work that way because if there was a consistently low-traffic route, everyone else would take that route until it was no faster than any other route. Traffic will generally be the predictable and similar every day, plus 'arriving at work early' is not much of a reward.
> A constant predictable reward is boring, less engaging.
Perhaps but predictable outcome is a very desirable quality. No one wants a hammer that sometimes drives nails and sometimes doesn’t. All of the current harness engineering work is about squeezing predictability out of the LLM.
> Good traffic is not a "reward" for driving to work(!)
Like hell it’s not. I drove into work last Friday and there was no traffic because of the holiday weekend. It was amazing. Had me considering whether Friday should be one of my standard RTO days.
I think you're missing the point. Predictable outcomes are desirable, but they aren't addictive or gambling. People quickly get used to opening the faucet and seeing water come out and stop doing it, whereas people scroll TikTok or channel surf for hours at a time.
In what way was amazing no-traffic "a reward"? What system was rewarding you for what change in behaviour?
You went off on a tangent and I responded. None of this is actually relevant to the topic of whether LLMs are “basically gambling”.
> In what way was amazing no-traffic "a reward"? What system was rewarding you for what change in behaviour?
What does this mean? Are you asking me to describe the dopamine system or are you implying that rewards have to be driven by some external system’s goal?
A reward implies it was given by something outside of your control. The lack of traffic wasn't a reward given to you, it's just a state of the highway.
You invite a friend over, but raccoon appears. Then pigeon appears. Then friend appears but at the last second suddenly becomes a banana. You remember you are out of bananas so you order more and also some cola zero cans on your local grocery delivery app. You are back to the party, but now you have 5 friends in the room, and you run de-duplication query. Now half of your friend is sitting at the sofa, and another half becomes a quarter of banana. Suddenly bananas arrive so you need to open the door. Once you are back there are no friends, pigeons or raccoons but also no bananas and no cola - all the delivery results are gone. This seems to be urgent and important, gotta fix this first before going back to that friend invitation...
This is not my experience with current AI models at all. But regardless you are not describing anything that sounds like gambling. You are describing a weird hallucinogenic experience.
i’m not sure “variable rewards” is the right term, but i do agree with the op that it is very similar to gambling.
regarding your examples, i think the difference is that with ai, you’re literally sitting in front of a machine, pressing a button, and (almost instantly) getting a result that, if not desired, can immediately be tried for again. you even spend “tokens” to do this, and at least in my native language, “token” brings to mind the coins you’d stick in a slot machine
I don’t see much similarity beyond the most superficial.
If you sit at a slot machine and pump quarters into it, each “turn” is independent. You spin and you win or lose. It’s pure chance and there is no destination. You execute the exact same action over and over and hope random chance brings you more money.
If you sit down in front of a coding harness, the progress is incremental and directed. You ask for a thing, the LLM produces something that is hopefully close to what you wanted. You give it more direction to prod it closer to the end state you want. You are not executing the same action, but incrementally nudging it in the right direction. I’ve literally never restarted from the same initial state with the same prompt and hoped for a different result and I don’t know why anyone would. Rarely I’ve thrown away the progress made and started over but always with a very different prompt that includes learnings from the failed attempt.
So maybe it's more like poker than a slot machine? Still intermittent rewards with a lot of random chance. The nuance of the analogy isn't that important to the general idea
The LLM could one shot a brilliant solution, or it could lead you down a days long path to nowhere. It could give you accurate useful information or it could completely make up something that isn't at all correct. The fact that each usage is a dice roll where you have a desired outcome that will be fulfilled at a variable level makes it very similar to gambling. And in fact the part that makes it addictive is the near hits where the LLM comes very close to giving you what you want, but not quite there. That keeps you coming back.
i agree with you that it’s principally different from a slot machine, and that it’s possible to use it in a way (like you describe) that is much more focused, for lack of a better term, to great effect
most people don’t use ai this way though, and i still feel like the end-psychological reward mechanism is very, very similar to gambling regardless of how well one utilizes it (and this is even more obvious with image generation as you chase that perfect output)
perhaps it’s better to compare it to gacha than slots?
How do they use it? Surely no one is just repeating the same prompt over and over (except as a Ralph loop perhaps, which is automated). I’m really struggling with the notion that most people just throw the same prompt repeatedly hoping it eventually works. Because that doesn’t sound like gambling. It sounds crazy (and frustrating).
> and i still feel like the end-psychological reward mechanism is very, very similar to gambling regardless of how well one utilizes it
In the sense that you get a dopamine reward when you succeed, sure, but I get the same reward when I code by hand and achieve a successful result.
> and this is even more obvious with image generation as you chase that perfect output
This is fair, because sometimes with image generation the same exact prompt will produce very different output. This is becoming less true as the models get better and it becomes more effective to direct image generation iteratively than to keep starting from scratch with a barely tweaked prompt.
When working on a problem models will walk you down a garden path requiring only yes/no answers or very brief clarifications for a long time. It's always proposing its own workarounds/suggestions/etc. Especially if trying to debug something where it has more understanding than the user so really all it needs from you is "uh sure try that too I guess" from time to time.
And sometimes the debugging has already veered completely off course at the beginning so it's futile, but each time the fleeting hope that just a few thousand more tokens will magically fix it tempts you to keep going a little longer.
as an example, i was using chatgpt a few weeks back to help me remember the name of a painting i’d seen about a decade ago. i could recall the general shape of the subject and that it was europeanish, but nothing else. after seven turns or so it finally got it, and honestly, the relief of finally remembering the name felt like, well, hitting a jackpot
i’ve had a similar feeling of success after trying to get it to give a comprehensible answer when asking it for a solid counter-argument to philosophical questions. it is indeed often crazy and frustrating
> but I get the same reward when I code by hand
i have only done very simple coding work with llms, so that may be why our ideas differ about the feeling of reward. this is where the comparison to gacha makes more sense than slots, since when you’re coding, you still get a reward each turn whilst chasing the final/desired result
I have this thing with a wallpaper with a small island, a (tiny?) house, a pier and a boat. Probably somewhere in Canada or Scandinavia. Had it on my PC in the 90s and no model was able to help me with this.
> the relief of finally remembering the name felt like, well, hitting a jackpot
I understand the joy of success but I fail to see how this is gambling. I could have an equivalent conversation with a friend (more likely about a movie in trying to remember than a painting, but still) and get the exact same type of iterative “no, not that one, it was more like X” and feel elated when my friend finally realizes I’m taking about a scene from Hot Tub Time Machine.
sorry, maybe my english is failing me here. what i was getting at is that using ai feels like gambling to me, where tokens, time, etc. are wagered against the chance for desired output. the risk is that you waste your tokens/time, and the prize is a useful answer (if not the exact code/solution/whatever one was hoping for)
i wasn’t trying to convince you, i’m just explaining that this is gambling in a meaningful sense to some people, especially with how turn-based and unpredictable the whole system is. whether it is actually gambling (semantically, legally, ontologically?) isn’t really interesting imo. what’s interesting is that it’s structured similarly and feels nearly identical to some people
but then, i also feel like there’s an element of gambling in the example of you talking with your friend, though i think it would be better illustrated if it was a conversation with a random person
I just wanted to chime in that there is nothing wrong with your English. I'm a native English speaker and I, too, find using an LLM to accomplish something to be very gambling-like. You're wagering something of value (your allotted tokens, and your time) on an unpredictable outcome that you hope benefits you.
thanks, i really appreciate the kind words. it’s just such a precise language that i worry i might be accidentally adding (or not including) important subtext in discussions like this
I agree with the other commenter. Your English is perfectly good. Your subjective experience is your own and I can’t disagree with that.
I’m working from a place where my employer pays for my tokens so I’m also not spending anything except my time. Maybe if I were, it would feel more like gambling.
I would agree that 1-2 years ago models were more "slot machine"-esque - sometimes the output was good, sometimes the output was bad. And as a result, I primarily used them for auto-complete functionality and bouncing ideas around. In those workflows, you can easily ignore it if the spin is wrong.
Not everyone has the desire to work around the system, and many are diametrically opposed to the concept of AI. They get this perception that it's a slot machine because of that inconsistency, and then do the human thing of assuming that other people must just be flawed if they're different from them. They're "addicted to gambling".
Obviously, things have changed. Open models can still be like that, but are often so fast and cheap at iterating it doesn't matter. SOTA models aren't perfect, but are to the point that they're generally much better than the average developer.
But once that perception set in and the meme spreads, it's really hard for some to break out of it. Especially at the pace AI development has been moving. It's just that simple.
1000%. My social community if reference is nomads of varying skills. The lower-skilled among them (the ones who were on the digital marketing, crypto,, now into building sales CRM tool) are manifestly demonstrating addictive tendencies around vibecoding) - the chain-smoking ADD guy pulling all-nighters still sticks in my mind. I myself (also low sw-skilled, just not the digital marketing / crypto hustling type) am exhibiting these tendencies in the form of all-nighters that in practice produce very little leverageable value-add. To remedy I try to be conservative in my goals - like only produce a comprehensive english-language PRD instead of any code till the tools get better and stroke the gambling instinct a bit less.
(By the way this is not to say that all nomads are low-skilled or even low-sw-skill, not by a long shot. I’ve met plenty of CS-degreed nomads who I treat as the wizened experts that they are, from whom deferentially elicit their pearls of hard-won software project wisdom. It was one of these even who, non-derisively, recognized the ‘pulling the slot machine lever’ reflex among vibecoders after I posed the apparent addictive tendency. On later describing this to a longtime friend deeply conversant in social science, he immediately responded with a nod and the phrase ‘variable reward response’.)
I really do think its tge case. Ever notice how juuust when you run out of usage its allllmost correct? then the handy helpful "buy more credits!" link pops up. Meanwhile local models are finishing the task without fucking around and without alterting shit it wasnt supposed to touch.
How is it a environmental catastrophe level, datacentre requiring, bullshit model is so much worse than something that runs on my workstation and doesn't cost us a ha itable planet?
Either they're fucking with us serving 8b models at scale or china really has the AI race in the bag so much so that they can openly release what the USA jealously guards.
They moan and complain about china copying from them (while doing the same), but if that's the case in full, then why are the chinese models better? you dont copy bad work and come out ahead.
Modern LLM services are engineer's pipe dream that was heavily shaped by the shadiest product management dark patterns you can find: applying gambling-style engagement tactics, exploiting cognitive biases, exploiting users' lack of technical understanding to inflate product expectations, using fear mongering in external and investor communications. And that's not even a complete list.
Programming before AI was always variable reward. It was a gamble against your own time and patience. Maybe I'd waste hours down the wrong rabbit holes trying to find a library that worked for my use case. Maybe I'd waste a day trying to get an API to do something it turned out it couldn't do. Maybe I'd have to redo my entire approach because of some factor I hadn't considered. Something I wrote could have worked on the first try or I could have had to spend the day chasing logic errors (or multiple days chasing memory errors if it was C or C++). Maybe I would just get bored of the project, especially if I realized there were 20 layers of yaks I needed to shave first, and Visual Studio got stuck updating again, and before I could even start actually coding I had to spend the entire evening on an exhausting merge conflict. My entire weekend could be gone with nothing to actually show for it.
I got so sick of all this at some point that I slowly stopped doing anything that wasn't my job. But then AI got better and better and I realized it was the ultimate unblocker. When that dreaded malaise started creeping in signaling it was a project's end because I didn't want to waste any more of my life dealing with bullshit orthogonal to what I was trying to do, I'd give it to the AI. It felt like a miracle the first time this worked, and it still does. If we were previously equipped with shovels to dig through bullshit, we now have a fully automated Bagger 288.
The reward schedule now isn't variable anymore; the chance that I finish something in a good state is 100%. I can focus on the parts I actually enjoy - architecting the broader system, making the parts mesh together in a sensible way that's easy to work with and has some mathematical elegance to it, hand coding the bits I want to be really specific about (but now without the endless frustration of bugfixing or import errors and edgecases being immediately discovered, thanks to the AI).
>Maybe I'd waste a day trying to get an API to do something it turned out it couldn't do.
I was working on a side project recently. I had spent months designing the data model in my spare time, thinking through how to make it as elegant and durable to change as possible in the long term, since (if I launched it) the repercussions for getting it wrong would be significant.
Once I had a working design, it probably would have been several more months to build a working prototype and start testing it.
Instead, Claude knocked out the prototype for me in an afternoon. And it immediately became clear that it didn't work: not because the data model didn't solve all the problems I wanted it to solve, but because it didn't fit the shape of how I quickly learned a normal person would need/want to interact with the product. I was so focused on the long term, that I never thought about what the first five minutes of a user with hands on the thing would need. And the changes needed would be significant.
Maybe there's some variable reward mechanism. But I sure was glad to be able to pull that particular slot machine handle and learn that than waste even more of my time on what was a dead end.
(Appreciating this entire thread) - side question: you used the word ‘shape’ in the abstract sense, something I never came across until Claude vibecoding came along. Was it common / did you use ‘shape’ in the abstract sense before say 2025?
Also, my coming from being a non-coder, I have a lot of appreciation for the possibilities for project failure one way or another due to a data model or project schema being wrong, even though I still only have a superficial understanding of what either of those concepts even are. . My question is, does your conception of data models in the abstract come from a formal academic course, like an algo’s & data structures course, or from trade-knowledge acquired through the practitioner grapevine?
> It was a gamble against your own time and patience.
At this point, what do the words even mean? Your own patience and available time are always completely random and fairly distributed across a large enough data set?
> Maybe I'd waste hours down the wrong rabbit holes trying to find a library that worked for my use case. Maybe I'd waste a day trying to get an API to do something it turned out it couldn't do. Maybe I'd have to redo my entire approach because of some factor I hadn't considered.
Our ignorance isn't random chance. As we research and experiment, we reduce the problem area.
> At this point, what do the words even mean? Your own patience and available time are always completely random and fairly distributed across a large enough data set?
Predicting the time a task will take is impossible. Something that sounds like a 5 minute script can turn into a month of banging your head against unknown unknowns. I lose my patience when the afternoon I allocated is getting overrun by nonsense and I'm missing out on other things I wanted to do or household maintenance.
> Our ignorance isn't random chance. As we research and experiment, we reduce the problem area.
Every thought we have has random chance to be wrong despite our conviction that it's correct. Descartes' Evil Demon plays his tricks on all of us. How many times have you typed some line of code only to realize it was obviously wrong afterwards? Even for simpler matters we "hallucinate" all the time. I was deep in thought trying to help someone come up with an acronym the other day and felt convicted that "Goal Oriented Augmented Retrieval" worked for GOAL until I said it aloud.
Our thoughts and actions are consistently wrong some portion of the time because our meat computers are not perfect positronic brains running prolog. We put cereal in the fridge and say "you too" to the waiter. Every thought we put down or action we take is a gamble on the soundness of the thought or action.