Stories
Slash Boxes
Comments

SoylentNews is people

SoylentNews is powered by your submissions, so send in your scoop. Only 9 submissions in the queue.
posted by jelizondo on Saturday July 26 2025, @04:32PM   Printer-friendly
from the he-gets-it! dept.

Arthur T Knackerbracket has processed the following story:

Enterprise CIOs have been mesmerized by GenAI claims of autonomous agents and systems that can figure anything out. But the complexity that such large models deliver is also fueling errors, hallucinations, and spiraling bills.

All of the major model makers – OpenAI, Microsoft, Google, Amazon, Anthropic, Perplexity, etc. - are singing from the same hymnal book, the one that says that the bigger the model, the more magical it is. 

But much smaller models might do a better job with controllability and reliability. 

Utkarsh Kanwat is an AI engineer with ANZ, a financial institution headquartered in Australia. In a blog post, he broke down the numbers showing that large GenAI models become mathematically unsustainable at scale.

"Here's the uncomfortable truth that every AI agent company is dancing around: error compounding makes autonomous multi-step workflows mathematically impossible at production scale," Kanwat wrote . "Let's do the math. If each step in an agent workflow has 95 percent reliability, which is optimistic for current LLMs," then five steps equal a 77 percent success rate, ten steps is a 59 percent success rate, and 20 steps is a 36 percent success rate.

What does that all mean? "Production systems need 99.9%+ reliability. Even if you magically achieve 99% per-step reliability (which no one has), you still only get 82% success over 20 steps. This isn't a prompt engineering problem. This isn't a model capability problem. This is mathematical reality."

Several analysts and GenAI specialists back Kanwat's view.

Jason Andersen, a principal analyst for Moor Insights & Strategy, said that enterprises often opt for the path of least resistance. If the large model maker is promising to solve all of their problems, they want to believe that. But it is often the much smaller and more-focused strategies that deliver better results.  Small, tight and well-scoped is good. Loosey goosey is bad. There is a lot of wisdom in going small

"This points out that the real value of an agent in an enterprise sense is to put boundaries around the model so you can get a certain degree of purpose out of it," Andersen said. "When you have a well-crafted and well-scoped (GenAI) strategy, you are likely to have more success."

The larger the model, "the further away you get the accuracy line, further away from reliability," Andersen said. "Small, tight and well-scoped is good. Loosey goosey is bad. There is a lot of wisdom in going small."

Andersen said that he asks CIOs whether they want the AI model "to be the pilot or the navigator?" 

A good example of this, he said, is GenAI-powered vibe coding. Should AI be helping the coder or replacing the coder?

"Both have humans in the loop but what role is the human providing? Is the human running the show or is GenAI running the show?" Andersen asked. 

Justin St-Maurice, technical counselor at Info-Tech Research Group, agreed that many enterprises are not doing themselves any favors by focusing overwhelmingly on the largest models.

"We are putting agents into complex sociotechnical systems. Agent systems run the risk of causing feedback loops and going off the rails, and the inherent nature of LLMs is randomness," St-Maurice said. "There is a real balance between taking advantage of the generative nature of GenAI and putting rules around it to make it behave deterministically."

Andersen offered an analogy of hiring a new employee and instead of training that new worker on how the team does things, the executive told the new employee to figure it out on their own. And when that new employee's work is not what the executive wanted, the company blames the employee rather than the executive who didn't want to spend the time or money on training new talent.

Kanwat also argued that the smaller models – even when deployed in massive numbers – can be far more cost-effective and often deliver an outright lower price.

"Context windows create quadratic cost scaling that makes conversational agents economically impossible," Kanwat said, and then he offered what he said was his own financial experience.

"Each new interaction requires processing all previous context. Token costs scale quadratically with conversation length. A 100-turn conversation costs $50-100 in tokens alone," Kanwat said. "Multiply by thousands of users and you're looking at unsustainable economics. I learned this the hard way when prototyping a conversational database agent. The first few interactions were cheap. By the 50th query in a session, each response was costing multiple dollars more than the value it provided. The economics simply don't work for most scenarios."

Kanwat said that many autonomous agent companies are going to have severe economic issues.

"Venture-funded fully autonomous agent startups will hit the economics wall first. Their demos work great with 5-step workflows, but customers will demand 20+ step processes that break down mathematically," Kanwat said. "Burn rates will spike as they try to solve unsolvable reliability problems."

Andersen agreed with the pricing concerns.

"The more context you have to give every step, the more the price goes up. It is a logarithmic pricing model," Andersen said, stressing that the model makers are going to soon be forced to sharply increase what they charge enterprises. 

A chorus of AI insiders chimed in. Himanshu Tyagi is the co-founder of AI vendor Sentient and he argued that "there's a trade-off between deep reasoning and streamlined reliability. Both should coexist, not compete. Big Tech isn't going to build this. They'll optimize for lock-in." Robin Brattel, CEO of AI vendor Lab 1, agreed that many enterprises are not sufficiently focusing on the benefits of smaller models.

"AI agents that focus on specific, small-scale applications will have reduced error rates and be far more successful in production," Brattel said. "Multi-step AI agents in production will find data inconsistency and integrations incredibly challenging to resolve, causing costs and error rates to spiral." 

Brattel had specific suggestions for what IT should look for when assessing various model and agent options.

Consider the "Low precision requirement. Can the solution be approximately right? Illustrations are easier than code because the illustration can be 20 percent off the ideal and still work," Brattel said. Another factor is "low risk. Generating a poem for a custom birthday card is low risk compared to a self-driving car."

One security executive who also agreed that small can often be better is Chester Wisniewski, director of global field CISO at security vendor Sophos. When Wisniewski read Kanwat's post, he said his first reaction was "Hallelujah!" 

"This general LLM experiment that Meta and Google and OpenAI are pushing is all just showoff (that they are offering this) Godlike presence in our lives," Wisniewski said. "If you hypertrain a neural network to do one thing, it will do it better, faster and cheaper. If you train a very small model, it is far more efficient."

The problem, he said, is that creating a large number of smaller models requires more work from IT and it's simply easier to accept a large model that claims to do it all. 

Creating those small models "requires a lot of data scientists that know how to do that training," Wisniewski said.

Even Microsoft conceded that small models can often work far better than large models. But one of their AI execs said small only works well for enterprises if the CIO's team has put in the time and thinking to map out a precise AI strategy. For those IT leaders who have yet to figure out exactly what they want AI to do, there is a reason to still embrace the largest of models.

"Large models are still the fastest way to turn an ambiguous business problem into working software. Once you know the shape of the task, smaller custom models can be cheaper and faster," said Asha Sharma, the corporate VP for AI at Microsoft. "Smart companies don't pick a side. They standardize on a common safety and observability stack, then mix and match models to meet quality, cost, and latency goals." (Note: Microsoft declined an interview request from The Register. We reached out to just about every major model maker and they either declined or ignored our request. The Microsoft comment above came from an emailed statement sent after publication.)

Not all enterprises have focused solely on large models. Capital One, for example, has focused on GenAI efforts that limit themselves to their internal data and they also severely limit what can be queried to what the database knows.

Kanwat said most enterprises are not the ideal clean environments for GenAI experiments. 

"Enterprise systems aren't clean APIs waiting for AI agents to orchestrate them. They're legacy systems with quirks, partial failure modes, authentication flows that change without notice, rate limits that vary by time of day, and compliance requirements that don't fit neatly into prompt templates," Kanwat said. "Enterprise software companies that bolted AI agents onto existing products will see adoption stagnate. Their agents can't integrate deeply enough to handle real workflows."

The better enterprise approach, Kanwat said, "is not a 'chat with your code' experience. It's a focused tool that solves a specific problem efficiently."


Original Submission

This discussion was created by jelizondo (653) for logged-in users only, but now has been archived. No new comments can be posted.
Display Options Threshold/Breakthrough Mark All as Read Mark All as Unread
The Fine Print: The following comments are owned by whoever posted them. We are not responsible for them in any way.
(1)
  • (Score: 5, Insightful) by JoeMerchant on Saturday July 26 2025, @04:39PM (4 children)

    by JoeMerchant (3937) on Saturday July 26 2025, @04:39PM (#1411594) Journal

    > smaller models might do a better job with controllability and reliability.

    So, 1) demonstrate that, 2) profit.

    --
    🌻🌻🌻🌻✌️ [google.com]
    • (Score: 3, Interesting) by Mojibake Tengu on Saturday July 26 2025, @07:37PM (1 child)

      by Mojibake Tengu (8598) on Saturday July 26 2025, @07:37PM (#1411637) Journal

      You can run DeepSeek R1 locally, even on Raspberry Pi.

      That's what makes Murricans so freaky about that one, any drone can think now. Or at least make some unassisted decisions.

      DeepSeek is already forbidden in Germany (by law) and in Czech Republic (by government edict).

      If you ask me, DS/R1 is pretty obsolete (half a year old), we can have better models now, Kimi K2 is my idol of the month.

      --
      The kinder you are, the easier it is for wicked people to morally coerce you.
      • (Score: 3, Informative) by Anonymous Coward on Sunday July 27 2025, @01:18AM

        by Anonymous Coward on Sunday July 27 2025, @01:18AM (#1411665)
        DeepSeek the company is banned because of alleged illegal data transfers.

        The DeepSeek the open sourced AI stuff, should still be able to be used. And that should be the main point of DeekSeek anyway - you run your own without sending your secrets etc to others.
    • (Score: 2) by stormwyrm on Saturday July 26 2025, @07:49PM (1 child)

      by stormwyrm (717) on Saturday July 26 2025, @07:49PM (#1411638) Journal
      It looks like it's not the sort of thing that a third-party provider can provide at scale, but perhaps as a bespoke service to a client. Perhaps a company like the one I work for (a very large multinational) could take this approach in-house, hire specialists in data science and neural network models or engage the services of a third party that can do it, with highly restricted custom internal models focused on very specific things that are needed by the enterprise.
      --
      Numquam ponenda est pluralitas sine necessitate.
      • (Score: 3, Informative) by JoeMerchant on Saturday July 26 2025, @11:02PM

        by JoeMerchant (3937) on Saturday July 26 2025, @11:02PM (#1411646) Journal

        The company I work for (a very large multinational) has taken this approach in-house, hired specialists in data science and neural network models to do it with restricted custom internal models focused on specific, and broad things that are needed by the enterprise.

        Results... vary by instance.

        --
        🌻🌻🌻🌻✌️ [google.com]
  • (Score: 5, Interesting) by Dr Spin on Saturday July 26 2025, @06:29PM

    by Dr Spin (5239) on Saturday July 26 2025, @06:29PM (#1411629)

    For those IT leaders who have yet to figure out exactly what they want AI to do, there is a reason to still embrace the largest of models.

    If you don't know what you want, you won't know when you haven't got it.

    In the words of W.C. Fields Never Give a Sucker an Even Break

    --
    Warning: Opening your mouth may invalidate your brain!
  • (Score: 5, Interesting) by anubi on Sunday July 27 2025, @01:24AM (2 children)

    by anubi (2828) on Sunday July 27 2025, @01:24AM (#1411666) Journal

    As I get older, I find I have more and more experience, more and more things that didn't work as planned, and why. All these things to consider makes me slower and slower. I noticed this a lot in myself, as younger colleagues would beat me to timely completion quite consistently, but they would consistently overlook the same attention to detail that had derailed a lot of my previous work. You know those little things like designing feedback loops for what appears to be a simple DC signal output buffer. Without resistors and capacitors in the right place, that thing will oscillate if it's load is capacitive ( like the roll of wire the customer connected it to ). The thing works great on the CAD design platform, works great on the assembly bench, oscillates when installed in the plane.

    The AI is doing the same thing. Its not a hand calculator anymore. There are several metric shittons of things to consider as the number of parameters that influence the design increase. What is the "big-O" notation for this? O = N^N?

    Seems everything I see around me eventually grows too big, becomes unsustainable, and topples over.

    Prime example: Empires of Man. Individuals survive - the Empire did not. Even the Mighty Sears Roebuck and Company ( Something that I could not even imagine as a kid - even Grandpa shopped there when He was a kid! ) disintegrated before me . Then I was introduced to History. All these empires that were and have left only artifacts and.knowledge. Now, I've always liked to take things apart to find out what made them tick. And boy, does this world have a lot of stuff that nobody seems to have much of an idea of how it works.

    Admittedly, we can now build Dubai Towers Burj Khalifa instead of The Tower of Babel, but the underlying paradigms are the same. Trees grow then topple.

    Its a constant tradeoff against the economies of scale, and the inefficiencies of complexity.

    Its been an interesting ride, but I guess since getting older, I have to be content being a passenger on the train, not the engineer. And quit getting worked up over trivial things that I have seen result in train wrecks. Its someone else's problem now. Learn to enjoy retirement.

    --
    "Prove all things; hold fast that which is good." [KJV: I Thessalonians 5:21]
    • (Score: 0) by Anonymous Coward on Sunday July 27 2025, @06:23AM

      by Anonymous Coward on Sunday July 27 2025, @06:23AM (#1411686)

      Sears didn't disintegrate. Equity fund managers saw a loophole in the pension plan, and destroyed the company to steal it. Nothing passive about it at all.

    • (Score: 1) by anubi on Monday July 28 2025, @11:15AM

      by anubi (2828) on Monday July 28 2025, @11:15AM (#1411795) Journal

      I finally remember what one of my supervisors called this phenomena of detail clutter : "Paralysis by Analysis". He had mixed emotions about it. While it lead to high design quality, that is, if it was ever delivered .

      I have had the exact same phenomena play out on my computers. I keep loading more and more stuff in it, until it gets so bad I can't find anything anymore...but it's all good stuff! I just can't stand deleting any.

      So, I do the next best thing....go get another computer and reload stuff from the old machine to the new as I need it, knowing the old machine is still there, all files still intact and searchable should I need them, but now my active work machine is now relieved of all the clutter that I was so loathe to toss.

      I now run a variant of that paradigm, having bought a dozen identical legacy DOS, W95, XP, Machines off eBay, and use of CloneZilla to provide unlimited copies of HDD with my core applications ready to go.

      I can always buy more blank HDD, and keep the old ones intact should I need to resurrect an ancient project. I do not abandon my old clients. That is MY work on those disks and I am leaving myself open to use it for other things. All of my machines can share files through my private LAN or LapLink 3 for the DOS machine that isn't FTP compatible.

      These are locked into place, frozen in time. No more upgrades. They are my design tools. They should perform the same as they do now for the rest of my life. My biggest fear is Microsoft may have left a time bomb in their products that may do me in, despite all my efforts to ensure the perpetual license these were sold with remain as such. Although I think Microsoft used to make a fine product, it seems to me XP was the pinnacle product, with WIN7 offering 64 bit processing. Seems subsequent products are so laden with telemetry, DRM enforcement, ad services, remote system administration ( whether I want it or not ), and compliance enforcement , that I find in me no desire to have anything to do with it. I find it useless.

      If, for some reason, someone demands I use a modern system, I will have to go get one. It will probably take me forever and a day to get anything done with it. The golf pro is a stickler for having his own golf clubs as he is ranked by whether or not he gets the ball in the hole, as I am ranked by those who would hire me by whether or not my stuff works. I simply don't have the time to keep learning all the new ways of getting schematic capture, PCB layout, or getting my data digitized, processed, and formatted as I want. So far, between my assembler and C++ compiler, I have been able to construct anything I need. I still often use discrete analog/power components like we did in the sixties. Those were the days I could look at a schematic and tell you exactly what each part did, why it was there, and what one would do if they wanted the circuit to do something else. I am so old I am still on a first-name relationship with most common vacuum tubes.

      If my employer is looking for theater, he best go hire an MBA and get executive level presentation. I am not good at doing that. I make a lousy salesman. It has taken me decades to feel comfortable knowing what I am doing and avoid most stupid stuff. Don't get me wrong, I can screw up just as good ( likely better ) than most, as I am prone to " think outside the box", only to discover I barked up the wrong tree. All I can really say is at least I know why it failed, and learned from that. On someone else's dime.

      --
      "Prove all things; hold fast that which is good." [KJV: I Thessalonians 5:21]
  • (Score: 2) by VLM on Sunday July 27 2025, @03:16PM

    by VLM (445) on Sunday July 27 2025, @03:16PM (#1411718)

    Production systems need 99.9%+ reliability

    Seems rather optimistic. HR is never that effective at hiring, or anything really, its kind of a make work department. Accounting/Finance might come close when they calculate payroll, even then I donno (consider complicated jobs like sales where there's multiple compounded commissions schemes and overlapping bonus factors). Customer support is generally competitive in certain sectors when it is well under 50% effective. For something like the cancellation department, they want 0% reliability.

    They're thinking 99.9% required if you replace your hard disk backup policy with asking a LLM nicely to remember a large string of hex digits. But most business processes are not reliable and its OK.

  • (Score: 2) by VLM on Sunday July 27 2025, @03:20PM (1 child)

    by VLM (445) on Sunday July 27 2025, @03:20PM (#1411720)

    complexity that such large models deliver

    It is, in a sense, a surface area to volume problem where making it bigger is not necessarily going to accomplish anything useful or even result in negative progress or retrograde advance.

    Just like the problem of software development where even in the 1960s it was "well known" that adding more programmers to a late project just makes it later, adding more LLM-slop to a process that already doesn't work at a small scale will just make it fail even bigger.

    This may apply to liberal arts in general; Someone writing a term paper at uni who has no idea what they're doing can't "fix it" merely by making a longer more complicated paper. This also applies to people who can't do math proofs. If you can't figure out the proof of the quadratic formula redoing it in an even longer and more complicated fashion will not have positive results...

    • (Score: 1) by khallow on Sunday July 27 2025, @04:28PM

      by khallow (3766) Subscriber Badge on Sunday July 27 2025, @04:28PM (#1411741) Journal

      This also applies to people who can't do math proofs. If you can't figure out the proof of the quadratic formula redoing it in an even longer and more complicated fashion will not have positive results...

      I've seen this in person. Someone purported to have a problem of Fermat's Last Theorem using basic algebra. The paper wasn't that long (10-15 pages maybe), but it was filled with algebraic expressions of growing complexity until somehow the desired final expression fell out by magic. There are a few basic errors that will do that easily such as divide by zero (1*0 = 2*0, thus 1=2) and who knows what was hiding in that paper? I sure didn't. But I do know that if there was a simple proof of the FLT, then someone would have an easy to understand proof of it by now, not the math equivalent of shell games.

(1)