Stories
Slash Boxes
Comments

SoylentNews is people

SoylentNews is powered by your submissions, so send in your scoop. Only 17 submissions in the queue.
posted by hubie on Wednesday August 05, @08:47AM   Printer-friendly

Government agency will use Google Cloud H4D VMs to replace HPE Cray machines:

Uncle Sam will no longer be hosting his own supercomputers to predict the weather. The U.S. National Oceanic and Atmospheric Administration has picked Google Cloud to provide the infrastructure for its weather forecasting operations.

In an announcement, NOAA boasted that it will be the first national weather prediction center to run on the commercial cloud, though the UK's Met Office is also in the process of moving its own weather prediction system to Microsoft Azure in a hybrid setup. Weather operations are typically run on in-house or government-funded supercomputer systems, which helps drive the HPC (high performance computing) market.

[...] The plan is to move NOAA's Weather and Climate Operational Supercomputing System, run by the National Weather Service (NWS) division, over to the cloud by December 2027, along with the software that generates NWS weather data for analysis.

The agency is hoping that the cloud will make model forecasting more nimble, resulting in earlier predictions and better warnings for all the extreme weather events that seem to keep occurring these days. It was the in-house systems that were holding things back, evidently. 

"Cloud-based high-performance computing will accelerate the transition of research into operations by eliminating traditional bottlenecks of on-premise systems," said NOAA Administrator Neil Jacobs in a statement

Jacobs noted that the cloud's flexibility for providing large amounts of compute is advantageous: the agency can ramp up cycles during tropical storm season, then wind them down during calmer periods.

[...] For the job, Google plans to use Google Cloud H4D VMs, built on AMD Epyc processors. Google labels these instances as "virtual machines" because they run under a hypervisor that integrates Google's networking and orchestration tools. As a result, they can be synchronized to run large jobs the same way supercomputers do.  

According to Google, customers can access H4Ds for as low as 3 cents per core-hour without long-term commitments. For supercomputing jobs, they can also use Cluster Toolkit to deploy clusters and Cluster Director to maintain them. Google Cloud's Batch can handle the queuing, scheduling, and resource provisioning.


Original Submission

 
This discussion was created by hubie (1068) for logged-in users only. Log in and try again!
Display Options Threshold/Breakthrough Mark All as Read Mark All as Unread
The Fine Print: The following comments are owned by whoever posted them. We are not responsible for them in any way.
(1)
  • (Score: 5, Interesting) by turgid on Wednesday August 05, @09:27AM (14 children)

    by turgid (4318) Subscriber Badge on Wednesday August 05, @09:27AM (#1450465) Journal

    They haven't thought this through, have they? What's happening to "the cloud" right now?

    AI Mania.

    All of that cloud capacity is getting gobbled up by AI mania. It's only going to get more scarce and more expensive until the bubble bursts. We know that there isn't enough cooling water and electricity to power it all. There's also not enough land (real estate) on which to build it.

    Plus why would you hand over your critical business infrastructure to a third party? It just doesn't make sense.

    If I were them, I'd be doing lots of very in-dept studies into the pros and cons of it all over many more months and taking my time to write up the reports, about a thousand powerpoint slides and many hours worth of presentations for everyone and their dog to see before making any rash decisions.

    • (Score: 5, Insightful) by shrewdsheep on Wednesday August 05, @11:30AM (2 children)

      by shrewdsheep (5215) on Wednesday August 05, @11:30AM (#1450472) Journal

      All of that cloud capacity is getting gobbled up by AI mania. It's only going to get more scarce and more expensive until the bubble bursts. We know that there isn't enough cooling water and electricity to power it all. There's also not enough land (real estate) on which to build it.

      AI and server workflows are still distinct and so is the current capacity. I believe that current weather models are still non-AI.

      The more pressing questions are: What guarantees can the weather service still give? Do they have sufficient documentation and strong enough contracts to guarantee the same level of reliability as they could using own servers? Do they have migration plans to other cloud services in place and can they guarantee transition times/cost? Can they guarantee a migration back to own servers and keep relevant data exfiltrated at all time? Do they have a smaller server still online inhouse, in case of emergency?

    • (Score: 4, Informative) by JoeMerchant on Wednesday August 05, @11:55AM

      by JoeMerchant (3937) on Wednesday August 05, @11:55AM (#1450477) Journal

      If the NOAA system is anything like other government supercomputer systems I have known about, it is 30 years overdue for modernization and dangling by budget shoestrings. The move to the cloud may be an opportunity for some overdue overhaul, and who knows, maybe next cycle the hardware vendors will be standing on the inaugural podium instead of the cloud providers.

      --
      🌻🌻🌻🌻✌️ [google.com]
    • (Score: 3, Interesting) by Anonymous Coward on Wednesday August 05, @12:08PM (6 children)

      by Anonymous Coward on Wednesday August 05, @12:08PM (#1450483)

      They haven't thought this through, have they?

      That presupposes they actually think, decision makers these days are depressingly so 'look a squirrel'

      What's happening to "the cloud" right now?

      AI Mania.

      Precisely the squirrel that's currently the focus of their broken little minds, 'we must have cloud as that's where AI lives, AI will save us, all hail AI...' or something akin to that runs through their heads like wee motorcars.

      All of that cloud capacity is getting gobbled up by AI mania. It's only going to get more scarce and more expensive until the bubble bursts.

      Ah, but what happens to the infrastructure paid for by the schmucks when the bubble is deliberately burst?

      We know that there isn't enough cooling water and electricity to power it all.

      Oh there is, unfortunately there are currently pesky humans competing for it

      There's also not enough land (real estate) on which to build it.

      Cf. the last response

      Plus why would you hand over your critical business infrastructure to a third party? It just doesn't make sense.

      Wheels within wheels, it might not make sense to us, but then we're not privy to their plans.
      Based on a number of seemingly insane 'to the cloud and AI' decisions, I'd make a guess that they're trying to consolidate into the 'cloud' and put under increasing AI control a number of human managed services that they deem are critical to their future plans.

      These services, factories, utilities etc are currently running on hardware they don't own, control or want to allocate any future human manpower to maintain, so they're persuading, by the art of expedient bullshittery, the current owners and operators of said services to move them and their control of them into the cloud and hand them over to their ever so obedient 'djinn' AIs to run them.

      If I were them, I'd be doing lots of very in-dept studies into the pros and cons of it all over many more months and taking my time to write up the reports, about a thousand powerpoint slides and many hours worth of presentations for everyone and their dog to see before making any rash decisions.

      Oh, I've worked in departments where decisions (rash or otherwise) were made down the pub, everything else (studies, reports etc) were mere window dressing. I've even sat on a technical committee where our sole function was to exist as a sop, a rubber stamp, to give the impression that the sometimes insane unilateral decisions of our lords and masters were taken on technical merit, and to provide them with a convenient whipping boy when they invariably went tits up.

      • (Score: 2) by JoeMerchant on Wednesday August 05, @03:32PM (4 children)

        by JoeMerchant (3937) on Wednesday August 05, @03:32PM (#1450506) Journal

        >There's also not enough land (real estate) on which to build it.

        Such a well informed observation deserves a little thought:

        Not enough land? What stops a datacenter from being built vertically? They need cooling, being up in the air is all the better. They choose not to do it because it is a little more expensive and a little slower to build. Right now, some datacenters are being situated on bare earth under tents - this only works in very specific locations, but they are doing it there because: it's fast and cheap to build (and the hardware is likely obsolete in a few months anyway...)

        Have you even a glimmer of a concept of places like Iceland? Not enough land, you say? Maybe they'd need to drop some more fiber connectivity from Iceland to the rest of the world, but there's geothermal energy, there's plenty of cooling (water and elsewise), and there's plenty of material to build concrete towers from. Australia may be a little more challenging from a cooling standpoint, but solar energy is pretty abundant there, and I'm having trouble wrapping my head around this "not enough land" statement when looking around outside of Alice Springs... Even the U.S. southwest would seem to have quite enough land for every datacenter on the planet, if only the companies doing the building would go to the trouble to locate there, instead of some cheap land near people where they can get cheap water for cooling.

        --
        🌻🌻🌻🌻✌️ [google.com]
        • (Score: 3, Informative) by canopic jug on Wednesday August 05, @04:01PM (3 children)

          by canopic jug (3949) Subscriber Badge on Wednesday August 05, @04:01PM (#1450514) Journal

          Not enough land? What stops a datacenter from being built vertically?

          In a word: cost. They are built as absolutely shoddily as they can legally get away with. Walls and ceilings are sheet metal with minimal insulation, if even that. The floors are poured concrete. You're basically dealing with a multi-acre pole barn. Doing anything vertically would not just increase the cost, but increase the costs by orders of magnitude. Keep in mind that building the centers is an end in itself, really it is the only goal. As such, shoddy materials help 100% — both by maximizing the profit margin and by allowing them to be built with such speed that local communities and governments do not have time to even react let alone work up a means to block the construction.

          The good news, I hear, is that most of the material can be recycled when they are torn back down. The only loss is the permanent destruction of the topsoil which was destroyed to make place for the buildings.

          --
          Money is not free speech. Elections should not be auctions.
          • (Score: 2) by JoeMerchant on Wednesday August 05, @04:30PM (1 child)

            by JoeMerchant (3937) on Wednesday August 05, @04:30PM (#1450516) Journal

            > increase the costs by orders of magnitude.

            Also increase the potential future utility of the structure for something other than a datacenter... vertical farms come to mind, particularly if there's abundant water available for evaporative uses.

            > The only loss is the permanent destruction of the topsoil which was destroyed to make place for the buildings.

            Typically they'll scrape the organics layers off when pouring a slab, but the slab itself isn't really permanent. Even the thickest concrete can be removed if you are determined enough. https://world-nuclear-news.org/articles/in-pictures-demolition-of-stade-reactor-building-progresses [world-nuclear-news.org]

            --
            🌻🌻🌻🌻✌️ [google.com]
            • (Score: 3, Informative) by canopic jug on Wednesday August 05, @06:45PM

              by canopic jug (3949) Subscriber Badge on Wednesday August 05, @06:45PM (#1450518) Journal

              Also increase the potential future utility of the structure for something other than a datacenter...

              That's unfortunately not part of their profit equation. The scams (and they're all scams) made around data center construction have the construction itself as an end goal. Anything after that, from their perspective, is irrelevant. Thus the costs are minimized. They'd use cardboard if they could get away with it. Electricity supplies are addressed with a bit of mumbling and hand waving as they're externalities, from the perspective of the construction investor.

              The company hiring the construction company and, to a much lesser extent, the construction company itself make a profit. Everyone else generally left holding the bag. If the community 'leaders' were crooked and/or naive enough to offer tax breaks, then they even end up losing a lot of money in the process.

              --
              Money is not free speech. Elections should not be auctions.
          • (Score: 2) by Reziac on Thursday August 06, @05:32AM

            by Reziac (2489) on Thursday August 06, @05:32AM (#1450548) Homepage

            There's a good point. No matter how badly the AI/datacenter craze goes tits-up, contractors still made a shitload of dough building the damn things.

            --
            And there is no Alkibiades to come back and save us from ourselves.
      • (Score: 2) by epitaxial on Wednesday August 05, @03:51PM

        by epitaxial (3165) on Wednesday August 05, @03:51PM (#1450511)

        This post could be boiled down to saying "ask the party currently controlling all three branches of government"

    • (Score: 5, Insightful) by datapharmer on Wednesday August 05, @02:57PM (2 children)

      by datapharmer (2702) on Wednesday August 05, @02:57PM (#1450503)

      They did think this through. Why build and own when the taxpayers can pay their buddies indefinitely at whatever rate they want to charge. "Sorry, no weather forecast for you as you were preempted by another customer... unless you want to pay a premium rate perhaps for guaranteed capacity?"

      • (Score: 3, Informative) by JoeMerchant on Wednesday August 05, @03:38PM (1 child)

        by JoeMerchant (3937) on Wednesday August 05, @03:38PM (#1450507) Journal

        There was a time (maybe even still today) when the US military paid commercial airlines for "as needed" use of their jumbo jets. They did things like outfit the floors for heavy cargo use (think: quick pull out all the seats from a 747 and use it to transport tanks...) and compensated the airlines for the privilege.

        It _is possible_ for government (taxpayers) and industry to closely cooperate and provide the necessary services without either party being overly abused. Of course, this requires a certain amount of integrity in government...

        --
        🌻🌻🌻🌻✌️ [google.com]
        • (Score: 2) by jb on Thursday August 06, @07:04AM

          by jb (338) on Thursday August 06, @07:04AM (#1450552)

          Of course, this requires a certain amount of integrity in government...

          True. But it also requires a certain amount of integrity in the vendor.

          That's probably a bigger sticking point. Governments change every few years, so there's at least a chance that some day your country might get one with integrity (after all, it has happened before).

          On the other hand, Google only gets worse and worse every year: no chance at all of that mob ever being anything vaguely approaching trustworthy.

  • (Score: 3, Interesting) by day of the dalek on Thursday August 06, @09:06AM

    by day of the dalek (45994) Subscriber Badge on Thursday August 06, @09:06AM (#1450553) Journal

    Unlike what other comments have said, this predates "AI mania". NOAA has been trying to modernize their numerical weather prediction (NWP) systems for quite awhile. A lot of the comments dumping on NOAA really don't understand what's going on and why this is happening. It's an overly simplistic analysis. I don't like NOAA switching to cloud computing and some of the other changes going on, but I do understand it. And it's not so easy as dumping on NOAA for buying into hype about cloud computing without considering the consequences.

    NOAA's NWP is built around fluid dynamics models like various configurations of WRF [github.com] and several other models with their own codebases. Take a look at the WRF code for yourself, if you'd like. There's a whole lot of Fortran in there, and there's not a whole lot of documentation about how it works. It's a massive code base, with a lot of optional modules that can be enabled or disabled at runtime, allowing the model to run in different configurations. It's also actually two different dynamical cores, the Advanced Research WRF (ARW) core and the Non-hydrostatic Mesoscale Model (NMM) core, which are part of WRF. When you're effectively bundling two models into one, it naturally makes that model more complex. Other models are more streamlined in what they do, but there's still have a huge amount of Fortran code powering even more modern models like anything that uses the FV3 [github.com] dynamical core. Like I said, take a look at FV3 and you'll find there's a whole lot of Fortran despite it being much newer than WRF.

    Meteorology programs don't require students to learn Fortran, and the language is being phased out in many parts of the field. A lot of modern tools use Python, sometimes with wrappers that have Fortran or C code underneath. Current students are learning Python. NOAA already has huge amounts of technical debt with maintaining all their existing NWP, so they're trying to replace WRF-based NWP with stuff built around the newer FV3 code, and have most NWP models use FV3. It's something called the Unified Forecast System [noaa.gov] (UFS). I don't really like the idea of building everything around FV3, and it turns out this didn't actually work so well. So NOAA built another dynamical core called the Model for Prediction Across Scales [github.com] (MPAS) to drive some of their NWP models. So we're right back to having two dynamical cores, a bit like the ARW and NMM cores in WRF. But perhaps we'll stop at two cores this time.

    Basically, NOAA thought they could conserve resources and make up technical debt by having all their NWP models be different configurations built around the FV3 core. That way they could turn off a lot of legacy models that actually have been tuned to work pretty well, but haven't been getting a lot of updates. There are other limitations like the difficulties in making WRF run on GPUs, so they believe they can get more performance gains by switching away from WRF. If they only have to support FV3, then they don't have to split personnel across maintaining a lot of different NWP models. But this didn't work, so now they have a second dynamical core anyway, which is MPAS.

    Many of the newer NWP models built around FV3 and MPAS actually seem to be regressions, in that they don't forecast the weather with as much skill as their legacy counterparts. Of course, NOAA will iterate and create new versions of the FV3-based and MPAS-based NWP models, and they'll eventually get that forecast skill back. It's not a permanent regression, but it will still be a drop in forecast skill for at least a couple of years. Basically, NOAA believes it's better to accept a degradation in NWP forecasting skill if it can erase some of the technical debt.

    As for the hardware, it's also really expensive to maintain the high-performance computing (HPC) systems that run these models. The High-Resolution Rapid Refresh (HRRR) model is a 3km mesoscale model built around the WRF-ARW core. That's a legacy model at this point, and I don't believe HRRR has received any updates for a few years now. My understanding is that the current configuration of the HRRR runs hourly on 1,728 cores, or 72 nodes that have 24 cores each, and obviously using MPI. This is one of the smaller NWP models, actually. There's also a HRRR ensemble that requires twice as many cores. These are the relatively small ones, with the FV3-based global forecast system running on over 6,000 cores every six hours. And NOAA has a suite of many other models that also require thousands of cores to run every few hours. There are also data assimilation systems that prepare the grids that the NWP models use. The data assimilation is also complex and can require a few thousand cores of its own to run. So you're talking about some pretty serious hardware requirements for the HPC systems.

    Some models also don't need to run all the time. When there is a tropical cyclone or a potential one in the Atlantic or Pacific, NOAA runs high-resolution nested models that follow the tropical cyclone. Each tropical cyclone had a Hurricane WRF (HWRF) nest that ran every six hours on 1,920 cores to forecast that storm. The HWRF is being replaced with a new FV3-based system called the Hurricane Analysis and Forecast System (HAFS) that runs on a few thousand cores. But like HWRF, HAFS only runs when there's a tropical cyclone or one is expected to develop. But even when HWRF or HAFS aren't running, NOAA would still need to maintain the few hundred nodes that would be used to run them. That's where the cloud is helpful, because those instances are spawned on demand, only when there's a need to run some of these models. And the same thing goes for experimental models during tests like the Spring Forecast Experiment.

    The bottom line is that NOAA just doesn't have the funding to purchase and maintain this hardware, and to pay the scientists to develop the NWP software. It's an absolute mess.

    And the situation wasn't made any better when DOGE went on a rampage and encouraged people with experience, as in people who are more likely to understand the legacy systems, to accept early retirement. Or people got tired of the uncertainty about whether DOGE was going to cut them, and they decided they could get better jobs elsewhere. Or Russell Vought decided to target the National Center for Atmospheric Research (NCAR), which is the lab that does a lot of work developing and maintaining WRF. NCAR didn't get cut, but the uncertainty about their jobs due to political incompetence and malice is another thing that might encourage scientists and engineers to pursue jobs elsewhere. So on top of the limited funding and technical debt, there's political malpractice in managing the weather forecasting enterprise. And like I said, you don't have a supply of new meteorologists being trained in universities to learn Fortran, who could take over the maintenance of the legacy models.

    So, yeah, I don't really like replacing NOAA's HPC systems by moving the NWP into the cloud. But this has been in progress for many years, and I remember discussions before the "AI mania" about trying to use Azure for some on-demand NWP for forecasting severe storms. The story here is that they've finally settled on Google's cloud, so I assume Azure is out. I understand why people want to dump on NOAA for replacing their own HPC systems and moving their NWP suite into the cloud. I don't really like UFS, turning off legacy models that work well, pushing everything into the cloud. But there's a whole lot more here than NOAA being swayed by hype around cloud computing.

  • (Score: 2) by VLM on Thursday August 06, @01:00PM (1 child)

    by VLM (445) on Thursday August 06, @01:00PM (#1450566)

    There's surprisingly little technical detail available and network capability.

    I would assume the code relies on high BW and low latency and the old cluster is modern-ish Infiniband or at least some exotic form of multiple 10G ethernet, but casual research found nothing.

    The google cloud info seems nearly content free its only buzzwords and scale free "line goes up and to the right" graphs.

    I honestly don't know if any serious "cloud" provides high BW low latency connections (stuff way beyond, uh here's a gigabit ethernet, maybe two of them bonded). I'm told Azure has some ridiculously expensive option but Azure is pretty much a non-starter for anyone who's not a MS kool aide drinker with infinite pockets of cash.

    So you can buy COTS 800G infiniband (which despite the name is more like 1.6 TB not 800 GB) and if you have to ask the price, you can't afford it (lets say it makes 10G look cheap). Its about $1K per switch port, $2K per card/module and cables are quite expensive also. But if you need terabyte speed instead of gigabyte speed, its COTS.

    Note that you can't lower latency by bonding 100 10G ethernet ports together, you'll just be able to slowly send 100 packets mostly in parallel. Infiniband latency made the news when it got under 100 ns packet hop, assuming your cables are short enough. It's quite fast; accessing another node's memory is slowly approaching being about as bad as accessing your own memory. Local memory is still "like twenty times faster" when you account for all the overhead but its better than at least 1000 times slower with gigabit ethernet.

    • (Score: 3, Interesting) by day of the dalek on Thursday August 06, @02:46PM

      by day of the dalek (45994) Subscriber Badge on Thursday August 06, @02:46PM (#1450578) Journal

      I've read that the latency of the connections in Google's H4D cloud is around 1.5-5 microseconds. This is significantly higher than the 90-200 nanoseconds that I found quoted for Infiniband. Even if those numbers aren't quite accurate, the latency is an order of magnitude greater. I believe H4D has 200 Gbps of bandwidth.

      Whether or not this is a problem depends on the application. When a NOAA model is run, it splits a domain up into blocks. Communication is needed between nodes to send lateral boundary conditions to blocks of the domain that reside on other nodes. It shouldn't be too hard to figure out if this is a problem.

      The models are fluid dynamics models, numerically integrating a set of partial differential equations. At each timestep, the lateral boundary conditions must be communicated with other nodes. Now, the HRRR model runs on 1,920 cores in its present configuration. With 24 cores per node on NOAA's supercomputers, this requires 80 nodes. H4D has 192 cores per node, meaning the same model can run on 10 nodes, which should mean there's less data to be computed between nodes. But the latency is still potentially an issue.

      The dt for the HRRR is 20 seconds. A standard off-hour run of the HRRR takes about 30 minutes of computing time to integrate 18 hours. So that's roughly 100 seconds to integrate 1 hour of time. It means that during that 100 seconds of compute time, it's necessary to compute lateral boundary conditions 180 times, or roughly every 0.55 seconds. If we assume that it's actually a bit faster because radiation code is very slow but doesn't run every time step, and if we assume that there's also some time needed to gather the data and periodically write output, we might generously bring it down to 0.25 seconds per time step. That would mean lateral boundary conditions need to be communicated four times per second. The latency just isn't large enough for this to be a problem, even if it's 5 microseconds.

      I believe the 200 Gbps of bandwidth in H4D is the same as Infiniband, but having 8x as many cores per node will mean there's less data that has to be sent between nodes. I suspect that H4D actually wins out just because there's less data that has to be sent between nodes. Basically, for numerical weather prediction, lateral boundary conditions just aren't sent between MPI processes and to other nodes often enough for the latency to be a problem. If you really turned up the model resolution to sub-kilometer horizontal grid spacing and needed to bring down dt to avoid violating CFL conditions, then this might be a problem. But for operational numerical weather prediction, it shouldn't really be a problem.

(1)