Rendered at 18:52:53 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
Animats 23 hours ago [-]
"Secondary market data shows moderately-used 2 to 3 year old GPUs trading at 50% to 70% of new pricing under normal conditions."
That's for good NVidia H100 units.[1] There's a shortage of those. That seems to be the price after removal, cleaning, testing and refurbishing. Raw units removed from a shutdown will not be as valuable.
H100 units are available on eBay, but multiple sellers are using the same picture of a new unit in its original packaging, a bad sign.[2] Some even have pictures with the logos of a competitor.
re: your last paragraph, there's probably only about 30 to 40 (at max) reputable relatively high volume dealers of used/refurb ex datacenter server equipment dealers on ebay that are located in the US48 states. It would be very risky in my opinion to buy a used GPU or multiples of GPU from some rando who has 14 feedback.
If you search ebay for server equipment like a Dell R840 with 768GB RAM, the same sort of dealers who are selling that and have thousands of feedback (at 98.5% of greater rating) are the ones I would consider much less risk.
lightbendover 23 hours ago [-]
Who is buying these at 70% of new pricing given the sky high likelihood of them being shot? Maybe it's safe to buy from small labs that went under quickly, but I can't imagine a cluster that has been operating near its thermal limits for a couple years fetching that kind of resale.
christina97 20 hours ago [-]
Why would they be shot? Unlike the consumer cards that are basically factory overclocked to look good on benchmarks, the datacenter GPUs are designed to run at full tilt 24/7 and survive for years.
bigbuppo 19 hours ago [-]
According to the article these cards have a 9% annual failure rate.
htrp 18 hours ago [-]
>The number traces to Meta’s Llama 3 technical report, which documented 419 unforeseen disruptions across 16,384 H100s over 54 days of training, of which 148 were GPU failures and 72 were HBM3 memory failures.
From an annualized number on the llama 3 training report. would be interesting to see if we have a better idea given that we're already on rubin.
christina97 18 hours ago [-]
It depends entirely on the time-to-failure distribution though whether used cards are a good deal or not. Often this kind of hardware has a bathtub shaped hazard rate, actually getting burned in cards may mean you get the weeded out solid specimens, and forgo the lemons.
lmm 16 hours ago [-]
Most hardware has an exponential (memoryless) failure distribution in practice. The bathtub curve is a myth.
fc417fc802 11 hours ago [-]
A bathtub curve is also memoryless though, and the back portion can be exponential. Consider two sub-populations of the same hardware. The first is small, evaded QA, and will fail early. Failures follow an exponential fall off. The second is much larger and failures rise exponentially with age.
Also consider mixing in "was dropped during shipping" or "was stored improperly".
eru 10 hours ago [-]
How is a bathtub curve memoryless?
fc417fc802 2 hours ago [-]
My mistake, it seems I misunderstood what OP meant by that. Ignore that semantic detail and my position still stands. You've got a mix of subpopulations and many of the failure modes are definitely time dependent in some manner at least in the real world.
Xalutiono 10 hours ago [-]
Where is this exponential thing coming from?
Why would hardware failure raise exponential with linear scale? Doesn't make sense
eru 10 hours ago [-]
> Where is this exponential thing coming from?
That's what you get from a constant failure rate per unit of time.
It's saying that hardware fails in the same way that Carbon-14 dating works.
Xalutiono 6 hours ago [-]
Lets hope they add more context to it.
Most if not all hardware failure on my side is def not memoryless besides the one lightning damage thingy but thats not hardwares fault.
onlyrealcuzzo 14 hours ago [-]
A 9% (static) failure rate, implies an average lifespan of 11 years...
camkego 13 hours ago [-]
With a median lifespan of 7.35 years, although I doubt that 9% per year remains constant as the number of years increases.
aaron695 9 hours ago [-]
[dead]
Chaosvex 19 hours ago [-]
Given the decades of consumer and enterprise GPUs often being identical or near identical hardware, it'd be interesting to see if there's any evidence of this actually being true.
eru 10 hours ago [-]
Binning can take in a supposedly uniform stream of chips and produce different tiers on the output.
threetonesun 22 hours ago [-]
It was true of crypto GPUs too, although mostly people picking them up for gaming. Always seems high to me too but if you can get any guarantee of them not being on fire when they were pulled the bathtub curve keeps you pretty safe, thermal limits are limits for a reason.
rxyz 22 hours ago [-]
crypto gpus didnt run at 100% power draw so the strain wasn't that massive
azeemba 22 hours ago [-]
Why weren't crypto gpus running at 100%?
3eb7988a1663 15 hours ago [-]
To advertise the maximum possible performance numbers, CPUs and GPUs massively shoot up the power draw for that final few percentage points. If you are worried about actual calculations/$, cutting back on the power will save your energy bill for practically no loss.
When I tested "ECO" mode on my AMD CPU, performance was ~97% of regular mode, and temperatures dropped 5-10C. I think ECO mode drops TDP from something like 100W->65W. Cheaper and cooler to run for basically no cost.
If you are bitcoin mining or selling GPU capacity, you are incredibly conscious of your energy bill, so it only makes sense to optimize the power draw.
skeptic_ai 14 hours ago [-]
Last point should also be for llm. Not just crypto
fc417fc802 10 hours ago [-]
Not true. A B200 "only" draws 1 kW. At 16¢ USD / kWh that's $1400 / year. When the hardware costs multiple tens of thousands you might be looking at a yearly electricity bill of ~3% of that. And that's before accounting for the datacenter itself.
The difference is that much of the hardware crypto was running on was dirt cheap in comparison. Add to that a lot of crypto was being run on discount gear in cobbled together setups. Some people were even stealing electricity which often involved installing large clusters of relatively cheap devices in places without proper cooling.
jpc0 10 hours ago [-]
Capex != Opex
Your are measuring two very different things.
And you are only considering electricity bills for the GPU. Higher TDP also means more heat which means more cooling.
It's not really easy to make an off hand comment about costs. A 30-40% decreaae in electricity usage and heat output for a fleet of GPUs definitely makes a massive difference to your opex bottom line.
fc417fc802 48 minutes ago [-]
The datacenter will already have been built with an explicit cooling budget. Moving 1 watt of heat requires substantially less than 1 watt to do the work. So you've got a constant multiplier against something like a 9% annual failure rate and an order of magnitude difference in cost.
dapperdrake 8 hours ago [-]
And the change in failure rate and technical advances are a de-facto change in actual depreciation vs. accounting depreciation.
dapperdrake 8 hours ago [-]
Electricity gets converted to heat which has to be removed through the cooling system.
Performance per Watt is the key metric for data centers full of GPUs.
fc417fc802 2 hours ago [-]
That just shows up as a constant factor, accounted for at build out. Consider if each watt dissipated required an additional watt. That's clearly an absurd overestimate yet it would "only" double power costs making the example $2800 / year. Underestimate the B200 at $30k USD and you've got 10.5 years of operation before you achieve parity as a fairly unrealistic worst case estimate. Now consider that at a 9% failure rate the average lifespan is only ~11 years. So electricity (both operation and cooling) is pretty much guaranteed to be less than half your total operating cost no matter how you slice it.
dantillberg 21 hours ago [-]
Most crypto mining on GPUs would use 100% of memory bandwidth, but only a fraction of the compute available. This is a consequence of ASIC resistance of their mining algorithms -- custom silicon can only offer a modest benefit over GPUs if the hard part is memory bandwidth.
noosphr 19 hours ago [-]
So does llm inference. You're lucky if you hit 40% of the advertised flops.
fwipsy 15 hours ago [-]
Right, but datacenter GPUs optimized for LLM training/inference would have a bandwidth:compute ratio scaled to that workload.
noosphr 11 hours ago [-]
No they don't.
dapperdrake 8 hours ago [-]
Then where are the HDMI ports on Nvidia's current data center GPU product lines?
eru 9 hours ago [-]
Why not? Seems like they would be poorly optimised?
noosphr 9 hours ago [-]
Because llm inference is not the only workload a GPU can do and custom silicon cost $10b a chip.
NuclearPM 20 hours ago [-]
Why is that true? Can’t you just make more stuff parallel and shrink the ASIC chips accordingly?
jmalicki 18 hours ago [-]
No - inability to do so is part of the design of a good cryptographic hash, quite explicitly.
Most were powerlimited to a degree to get the most hash/watt out of them.
metalliqaz 22 hours ago [-]
I bought a RTX 3070 off a miner when Eth went to proof of stake.
It was clean, cheap, and is still going strong for daily gaming.
In its working life it was undervolted and probably cooled better than in my rig.
BizarroLand 20 hours ago [-]
I got an entire Prebuilt PC when POW ended from a miner. AMD 5950x, 64gb ram, 3090, 2tb SSD, all for $1700. This was 4 years ago when the 3090 was $1500 by itself, and with minor upgrades it's still going strong today.
It's wild to think that the system now is worth at least as much as I paid for it then if not much more than that. I saw a similar one going for $2500.
fwipsy 15 hours ago [-]
Honest question. Does silicon wear out due to high temperatures, or is it more like lightbulbs where it wears out from thermal cycles? Does it really wear out at all?
blobcode 15 hours ago [-]
It’s mostly due to higher temps resulting in faster ion migration, which can cause a breakdown in the structure of transistors, as well as increased wear on the conductors (though this is rarely a dominating factor). Thermal cycling can also cause cracking, which is also a problem. Though there are lots of reason chips fail due to heat, from the wire bonds on pads getting too hot to increased leakage current at high temps.
Xalutiono 10 hours ago [-]
For a long time it was assumed it wouldn't wear out and tbh if you look how long it takes, how rare it is, its not a real issue...
besides what intel did with Raptor Lake. This series had massive issues with oxidiation. Was the first time ever i became aware of this issue on scale.
Nontheless there are papers out there that silicon can degenerate and does.
I feel like this is not as important as people make it out to be.
tyfon 1 days ago [-]
I half expect Nvidia to have buyback contacts like Ferrari with the larger customers to prevent a price crash when they all upgrade and to keep them scarce.
I hope not though, perhaps I can pick up a H100 in a few years if they get sold on the open market.
rickypp 1 days ago [-]
I worked at an org that had a substantial on-prem GPU datacenter. We transitioned to <Big Cloud Provider> with a substantial negotiated discount rate, with part of the contract being we would sell them all of our hardware and not purchase any more.
hamandcheese 23 hours ago [-]
> and not purchase any more.
why would anyone sign such a contract?
toast0 23 hours ago [-]
If you're planning to run in clouds, committing to not buy hardware (during the contract term, presumably) isn't a big imposition. Maybe you switch to a different cloud, and you wouldn't buy hardware for that.
If you want to switch back to on prem, there's probably a way to structure acquiring hardware so it doesn't break the contract. Maybe you lease it, maybe the purchase happens through a related company, maybe there was no way for the contracted cloud to find out...
collabs 23 hours ago [-]
I don't know if this was IT shrugging me off or if it was something real but some IT person at this big ISP I worked at told me that they cannot just buy an SSD — my windows box at work was running off of a hard disk in 2019 — and that there was some contract that said any computer hardware we bought had to be through HP or something like that and it takes many months it something like that.
lelanthran 10 hours ago [-]
> I don't know if this was IT shrugging me off or if it was something real but some IT person at this big ISP I worked at told me that they cannot just buy an SSD — my windows box at work was running off of a hard disk in 2019 — and that there was some contract that said any computer hardware we bought had to be through HP or something like that and it takes many months it something like that.
Plausible - the hardware might be leased, and so you have no right to modify it. You have to pay them to modify it.
Same as if you leased a car, you cannot do the services yourself, you have to pay them (and an approved agent of theirs) to do it.
thijson 5 hours ago [-]
We could only book our business travel through American Express. It must have made economic sense at the c suite level, I wasn't seeing it at my level, the flights were consistently more expensive.
protocolture 19 hours ago [-]
Was it a HP Finance thing? Because there are a shit ton of people who can transact through them now. We had HP finance, and it became a game of sorts to see what crazy shit we could get with it. Theres a local mob who sell computer parts who were happy to use HP finance. That said it definitely wasnt everyone and if you didnt do the legwork you could definitely be trapped.
Shorel 23 hours ago [-]
The usual very short term corporate thinking that maximizes quarter profits while bankrupting the company in the long term.
IMO a very shortsighted decision, if not downright stupid.
brookst 7 hours ago [-]
It’s also possible to bankrupt a company with excessive focus on the long term.
vasco 13 hours ago [-]
What is short sighted is parroting this over and over. It's not insightful, interesting or adding anything else to the discussion than "mean people are bad".
tryagainian 15 hours ago [-]
[dead]
fsuts 14 hours ago [-]
Because the discount would be sufficiently large.
lithos 5 hours ago [-]
Short term thoughtset MBAs.
23 hours ago [-]
throwaway85825 23 hours ago [-]
How does this work? Is the OEM giving the cloud provider a big discount or is the cloud provider giving you a teaser rate to lock up your business.
wmf 23 hours ago [-]
The list price of cloud is 10x higher than the price of hardware itself so that leaves room for some discounting.
jurgenburgen 12 hours ago [-]
What I’ve learned is that the list price is just for small companies. When you’re big enough you ask for a better contract.
to11mtm 20 hours ago [-]
Yeah....
Where folks (often) get lazy is the resulting math over what the real bean-counters care about (but are too lazy to check often).
In a past life, I worked on costing models for a Cable/Fiber contract house, to help the company decide 'what was profitable to keep in house' versus 'what do we subcontract' (sometimes that could even mean we just 'rented' a machine and had a qualified operator using it, based on that employee's hourly rate and expected L2R for taxes... so many spreadsheets...)
And from from my 'I don't know all the factors for this but I've seen how people screw up the big ones' view
(and frankly, I'm guessing a lot of us have seen and dealt with the same category of 'bad math' around outsourcing IT work...)
An on-prem data center means:
- You need to account for electricity costs
- i.e. CA vs midwest electric rates.
- cooling and power backup capability
- Smaller factor but real
- personnel cost
- e.x. there's probably cases where a smaller org could be better off with 'on-site' server admins that have other roles based on local wages. Kinda case specfic but it's a case.
- whatever the 'space' holding the stuff costs
- Sardonic take :Hey, let's have another unused meeting room instead! (e.x. In the case of on-prem shops that simply fled to AWS in their migration from VMware)
- the cost of licensing whatever is running
- In defense of this, In one of my earliest IT lives, AWS handling the Oracle licensing for a DB was a *huge* win as far as making it as easy as possible to ensure whatever was going on we couldn't have the Oracle licensing folks 'ding' us on whatever infraction occurred between reviews (that could not be understood by the majority of the company, often including the accused. I was never guilty but I saw it happen to others.)
- OTOH I know lots of folks who just want to be lazy about what they have to document.
Still, IMO a lot of orgs don't do the right math around these decisions, or just buy into the 'Well trends can change' as though they can decide as an org they need to suddenly triple capacity in a month and it would be able to organically happen in the first place.
Frankly, the orgs that 'might' need that either have their arch set up where they are in cloud, or they are onprem but can scale to cloud if needed in interim.
lmm 16 hours ago [-]
What are these orgs that have unused meeting rooms? Everywhere I've ever worked has had a chronic meeting room shortage.
nylonstrung 23 hours ago [-]
There's going to be a golden age of GPGPU compute in the next few years once A100/H100 are fully obsolete for running frontier models efficiently and the price plummets
It will be perfect for stuff like GPU-accelerated query engines, "classical ML" and every other CPU-based workload that could conceivably be offloaded to GPU
nl 18 hours ago [-]
There's nothing stopping you doing this now.
You can get used 16GB P100s on AliExpress for ~$100 if you want obsolete GPUs. Allegedly new AMD BC 250s are only slightly more.
I've looked at this some but I already have a GTX1070 which is only supported upto CUDA 11.9.
That's precludes some interesting modern optimizations out of the box. I've spend a lot of LLM tokens backporting some things, but I'm really not sure the hassle is worth it.
New hardware is just better. I think in maybe 5 years when supply and demand are back in equilibrium we are going to have some killer technology for decent prices, and 15yo H100s won't look attractive.
throw0101d 6 hours ago [-]
> You can get used 16GB P100s on AliExpress for ~$100 if you want obsolete GPUs. Allegedly new AMD BC 250s are only slightly more.
In a recent Gamer Nexus video with Level 1 Tech, they mention V100s are also quite useful for many applications that use FP64: stuff four in a workstation, and many PhD candidates would be quite happy with the throughput they can get for certain scenarios.
Schlagbohrer 7 hours ago [-]
Linux hackers will be finding all kinds of crazy uses for hardware that now costs $100,000 and in 5 years will be available as scrap.
This is assuming that there is no big, big disruption to the semiconductor industry (e.g. TSMC getting attacked), in which case... well, I am gonna treat each stick of RAM I currently have like it's a faberge egg.
throwaway894345 23 hours ago [-]
GPGPU? General Purpose GPU? If embarrassingly parallel CPU algorithms weren't offloaded to the GPU previously, why would the A100/H100 price drop make a difference? We had cheap GPU in the past and we still left plenty of performance on the table with CPU programs because they were easier to build.
Is the idea that previously maintaining GPU programs was expensive whereas now AI makes it cheap? If so, I could buy that line of reasoning.
Maybe relatedly, I expect (hope) the hardware manufacturers will ramp up supply in the meanwhile which would also put downward pressure on GPUs. Right now though this hardware crunch is making me sad, not even because of GPUs but also because of general memory / disk.
christina97 22 hours ago [-]
GPGPU programming has become significantly easier now and the payoff is bigger (better hardware), due to the immense investment in this due to ML/AI.
NuclearPM 20 hours ago [-]
What are the best tools for this?
girvo 20 hours ago [-]
We've not had cheap GPUs with this much VRAM before, though. Might be an interesting change, though I also doubt it personally.
nl 18 hours ago [-]
As noted in my other comment you can get obsolete GPUs (P100s, BC250s) with lots of RAM on AliExpress now. It hasn't proven revolutionary.
fc417fc802 2 hours ago [-]
BC250 shares memory with system. 16 GB GPUs have long been available to consumers for a small premium. What's never been readily available before is 40+ GB of HBM.
swiftcoder 13 hours ago [-]
is 16Gb "lots of RAM" in an LLM world? How many of these would you need to stack on a motherboard to inference a decent size model?
oblio 10 hours ago [-]
At least 4, probably 6+.
girvo 10 hours ago [-]
Those don't have a lot of RAM though. Not like these 40, 80GB ones.
> I, for one, welcome the coming age of the post-LLM-datacenter-overinvestment-bust-fueled backyard GPU supercomputer revolution.
> The Big Question is…
> Who is cultivating the option to snap up and repurpose vapourised datacenter investments at fire sale prices, soon as the "datacenter debt" cometh calling?
moffkalast 23 hours ago [-]
The only reason a datacenter would ditch their H100 is if it becomes uneconomical to run them, with newer silicon providing much more power efficiency. When that happens they'll look like a used V100 looks today: horribly inefficient, lacking modern data types and engines, requiring screaming server fans with weird adapters to not melt, way beyond end of life in terms of cuda support. Almost completely damn useless unless you really have no other alternative.
imtringued 10 hours ago [-]
Honestly, the primary reason you would ditch a H100 is because the ML models have moved onto a different data type, e.g. nobody uses fp8 anymore and everyone jumped onto fp4 or some custom block float thing.
mertleee 20 hours ago [-]
[dead]
someguyiguess 1 days ago [-]
That should be illegal. Sounds like a very fraudulent business tactic.
rcxdude 1 days ago [-]
Fraudulent, I wouldn't say so. Anti-consumer or anti-competitive? Sounds like.
KaiserPro 10 hours ago [-]
Its restrictive, but not that different from other types of deals.
For example, when I worked at a VFX software company, we were exclusively with one hardware partner. This unlocked something like a 50% discount across all our infra needs (they were big enough to provide switches, racks, servers, storage)
another company it unlocked a 75% discount.
rbanffy 21 hours ago [-]
It’s common for car companies when they enter a new market. It removes uncertainty from the second hand market.
By doing that, you know upfront what the value of your used hardware will be at the time you decommission it. It removes a lot of the risk for buyers in a volatile market.
lelanthran 6 hours ago [-]
> That should be illegal. Sounds like a very fraudulent business tactic.
a) It's not fraudulent, and
b) Nvidia has already signed on to buyback any unused capacity from the DCs it is selling to.
nl 18 hours ago [-]
The hyperscalers signing these contracts have decent legal departments. Think about Oracle for example - I'm pretty sure they know every trick there is about beneficial contract drafting.
I don't think they need some special protection against this kind of contract.
throw1234567891 23 hours ago [-]
Would you call a trade-in a fraudulent tactic?
lmm 16 hours ago [-]
Trade ins are fine. A contractual commitment to not buy hardware sounds like illegal restraint of trade.
brookst 7 hours ago [-]
How? It’s a contract. Should I be allowed to pay a landscaper less if I agree to not do part of the work myself?
I don’t understand the desire to tell totally anonymous third parties the terms under which they are allowed to deal with each other.
InsideOutSanta 23 hours ago [-]
Trade-ins are voluntary.
efficax 23 hours ago [-]
so are buybacks. you choose to sign the contract. there's no way they didn't have an escape clause, although likely it meant not using the cloud provider anymore
InsideOutSanta 22 hours ago [-]
> so are buybacks. you choose to sign the contract
Right, so they're not voluntary.
anomaly_ 22 hours ago [-]
Literally what? Do you understand what is being proposed?
InsideOutSanta 22 hours ago [-]
Yes. "You choose to sign the contract" is not a reasonable argument, for two reasons:
1. There is one supplier, so you have no choice.
2. Even if you had a choice to sign the contract, this still means that it's not the same as a trade-in, because trade-ins are always voluntary, but once you have signed the contract, a right of first refusal is not.
In general, the "you chose to sign the contract" argument is a poor justification for bad contracts. If the contract is bad, it is bad regardless of whether you chose to sign it.
fsuts 14 hours ago [-]
B2b contracts have less protection in law than consumer contracts
As businesses are expected to be more informed and equal in the negotiations
efficax 21 hours ago [-]
there are lots of cloud suppliers besides whichever big name cloud this is, you can even find other cloud providers in seattle
23 hours ago [-]
2III7 1 days ago [-]
[flagged]
HPsquared 24 hours ago [-]
Good old "win-win-lose"
cyanydeez 23 hours ago [-]
The Grift Economy places all legalities on the marks and their inability to form legal fights.
echelon 1 days ago [-]
Wouldn't that be crazy - a hobbyist market for H100s?
Maybe someone could start a business buying up and rehousing these.
rtkwe 1 days ago [-]
They're pretty specific to the datacenter use case with no outputs and they need to be cooled externally, principally through the very loud high speed fans used in data centers. I suppose you could strap a fan to one and put it in a normal case or maybe make a dedicated after market cooler (like the water blocks made for water cooling cases).
jtolmar 24 hours ago [-]
I think there's a market for a home AI server that can run an LLM or video gen model behind a web frontend. Not literally a raspberry pi strapped to an H100, but something with lopsided enough specs that people joke it is.
And precisely because it's such a huge headache to do yourself, I think a small company could make a nice business wrapping up used datacenter cards in that sort of server.
everforward 19 hours ago [-]
I'm doubtful they can, based on current inference prices. An H100 draws about 50W while idling, which is ~$8/month at average US electricity prices. They also draw ~400W while active.
The electricity prices are relevant because if you paid $0 for your H100 and didn't use it a single time, you could buy millions of tokens in inference just on the electricity it draws while idling. If you can't keep that thing saturated through the night, you're probably underwater overnight. Likewise, it's too small to run even the frontier open source models so you need to be able to live with worse models.
Max power matters because you aren't going to run many of those H100s before you blow breakers in most houses. Newer houses in the US are 15A service to non-kitchen breakers, so 1650W (that might be peak rather than continuous, not sure). If you're plugging that into an existing run, you could maybe run 2 before you start blowing breakers? You can't just plug 4 H100s into the wall in a normal house.
Maybe I'm wrong, though. I'd be curious, it'd be neat to run my own inference for something more than what'll run on a 3080.
smalltorch 19 hours ago [-]
Most houses have a 200amp panel and extra space for larger breakers.
A electrician can plop in a electric car charger for instance, that is a 40-50 amp circuit.
rtkwe 5 hours ago [-]
Newer houses yes probably have the overhead. We put in a charger for my wife (because she drives 100 miles a day we needed a pretty beefy level 2) and had to replace our whole panel because 1) we had 125A so failed the load calc and 2) we had the infamous Zinsco panels and the electricians refused to squeeze one in even if we did pass the load calculations. Luckily we had a basement so they only had to extend the circuits a little instead of rerunning them all over the house.
HWR_14 14 hours ago [-]
I'd imagine that you can have the power to the board shut off when not in use.
But I never assumed it was to be cost competitive at current token rates. I assumed it was for the same reasons people might use open source hardware. Freedom to tinker, etc.
rtkwe 5 hours ago [-]
Basically sounds like a hobo version of PCI-e hotswapping which is a thing on enterprise motherboards but not many consumer boards. Also depends on if the board can actually power off the slots which I'm not sure many are setup to do.
trollbridge 15 hours ago [-]
A typical oven plug can offer 9.6kW (so your 24-node H100 can keep your kitchen nice and warm). My furnace was 16.5kW, which once failed in an odd way and was stuck on for about a week. The people living in the house simply opened the windows until I got around to fixing it. I estimate it wasted $400 on the power bill.
Bombthecat 18 hours ago [-]
Meh
Models are too big now
timmmmmmay 24 hours ago [-]
This is also true of the older P100 and hobbyists do this stuff now, today
chasd00 22 hours ago [-]
> Very loud high speed fans
if you haven't heard a 5u server intended for a datacenter rack come to life it's quite the experience. Sounds like a plane taking off.
gerdesj 20 hours ago [-]
"Sounds like a plane taking off."
Helicopter. I live in Yeovil, Somerset, UK - there's a helicopter factory just down the road. I had a IBM "AS/400" or whatever they are called now in our computer room rack for a customer and it made nearly as much noise as everything else put together. It was clearly tuned for start up noise to impress because they would fire up in sequence, rise to a crescendo and then slow down in sequence to just a din instead of painfully loud.
A switch or PC server on boot will normally run up cooling fans instantly to max as a default protection mechanism until the "OS" has started and sensors read and then the fans will slow down to deal with the actual thermal load.
If you switch off your air conn, it gets noisy, quickly. Recently in the UK we are seeing routine temperatures around 30C and we broke 200 odd year records for temperatures a few weeks back. I know its even worse elsewhere but our infrastructure is not designed for this. Here we are at the same latitude as Calgary AB!
rbanffy 21 hours ago [-]
My workstations are usually tower servers, which are the same design as their 4 and 5u rack counterparts.
Until the thermal management kicks in, they sound like jet planes. When the thermal management starts and assesses the required cooling, it’ll throttle down the fans to reasonable levels.
That is, until the moment you push the machine to its limits. When then happens, you might get back to the same levels of the boot time, but it’ll require you to push everything to the max - CPU, memory, storage (all 24 bays) and so on. For a normal user, there is a lot of room and it’s virtually impossible, even with a dozen of Teams windows open.
baby_souffle 19 hours ago [-]
Yes, but the line between "reasonably quiet" and 100% isn't a step function.
If these do end up in Home Labs, it's going to be a server rack in the basement and not the server rack at the other end of the room.
rbanffy 10 hours ago [-]
> If these do end up in Home Labs, it's going to be a server rack in the basement and not the server rack at the other end of the room.
If they are too loud, yes. In my part of the world basements are not common, but my shed would be a sensible place for one.
CamperBob2 17 hours ago [-]
Closest thing I can think of is the old LucasFilm "THX" trailer. Only there's no release segment, not until the BMC comes up a minute or so later.
gessha 19 hours ago [-]
If you check eBay for V100 SXM variant, you will often see them sold in combination with a cold plate, together with a PCIe carrier board. There’s definitely a market for them.
lelanthran 6 hours ago [-]
Meh.
You don't need much of a fan if you are prepared to pump a lot of water.
I mean, sure, if you're trying to cool the thing with 100ml of water, then yeah, you need a fan.
OTOH I have an unused car radiator in my garage, a quiet and cheap pump, and space to hang that radiator outside the window. It's barely an afternoon's worth of work if the H100s already have liquid-cooling intakes/exhausts on them already (I dunno, I have never seen one in the flesh).
I'm pretty certain 10+ litres of constantly circulating water (antifreeze, in my case) through a car radiator would be sufficient to cool down 2x H100s. Hell, add in 10 RPM used car cooling fan and I can probably cool 20x of those running at full-bore.
If anyone wants to ship me a bunch of H100s, I'll happily build it all out, take pics and videos and post it back here. Just sayin' :-)
bayindirh 1 days ago [-]
Or you can secure yourself a DLC one and try to feed it with the correct regime (liquid composition, temperature and flow rate). I'd say good luck.
These things get hot and are fussy about their requirements.
wmf 23 hours ago [-]
H100s (and the PCIe converter card you need) are available on eBay.
trollbridge 15 hours ago [-]
Yeah, but an RTX 6000 is better for almost any application other than training, and the latter only really is better if you’ve got multiple H100s.
mschuster91 1 days ago [-]
The GPU alone has a TDP of 700W, together with everything else (CPU, RAM, storage, fans) you're looking at 1500W+. Depending on the country, that may be enough to saturate your home's electricity uplink...
throw0101d 6 hours ago [-]
> […] you're looking at 1500W+. Depending on the country, that may be enough to saturate your home's electricity uplink...
1500W is the power of the typical American microwave or (tea) kettle at 120V. (Convert over to a NEMA 6 plug and dual-pole breaker and you can get 3000W.)
3000W is nothing special in the rest of the world with >200V wall plugs.
InsideOutSanta 23 hours ago [-]
That's actually less bad than I assumed. High-end gaming GPUs are already almost at 600W with peak usage above that. I assumed it was much worse than that; that seems absolutely feasible for running at home.
elictronic 22 hours ago [-]
Never heard electrical uplink before but to those wondering Italy, India, and Japan all have requirements around 3kw. I was slightly surprised by this, but in the end if your running this in a tiny space with that low of power your already probably not buying used H100s.
holoduke 1 days ago [-]
Hell no. 1.5kw is half a socket capacity. A heat pump or many kitchen appliances are using much more. Our car charger uses more as well.
jermaustin1 1 days ago [-]
That's why they said "depending on country"
In North America we are on 120V, making a standard 15A outlet only 1500W max, and something like 1200W sustained. To use higher wattage appliances, we have to upgrade our outlets to 20A (2000/1600W) or up our voltage to 240V, but that carries a different set of plugs and outlets as well.
datadrivenangel 1 days ago [-]
Usually the home service is 200+ amps these days which is 24 KW total across all circuits, though you're right that most individual circuits are only 10-15 amp.
Kirby64 14 hours ago [-]
48kW across all circuits. The 200 amp rating is at 240V for US electrics.
You’re limited to 24kW per leg if you ran everything only on 120V, but any really large server is going to be running off 240V anyways.
cyberax 23 hours ago [-]
> In North America we are on 120V, making a standard 15A outlet only 1500W max, and something like 1200W sustained.
It's 1800W for short periods and 1500W sustained.
trollbridge 15 hours ago [-]
1,800 watt for periods less then 3 hours and 1,440 watts for continuous loads over 3 hours.
Realistically, a 15A breaker won’t trip on a 2,300 watt load for at least a few minutes, and often not until a few hours. A listed breaker is expected to trip in several seconds to 3 minutes for a 3,600 watt continuous load.
Take all those numbers and go up by 33% for a house or apartment with 20A circuits (quite common). Go up by 667% for an oven plug.
As anyone who takes care of rentals during winter knows, a typical circuit can tolerate two 1,500 W space heaters without tripping very much (although this is unsafe and a bad idea).
jermaustin1 5 hours ago [-]
Sorry, I was doing conservative napkin math, better safe than burnt to a crisp, as grandpa NEVER said... I'm currently having to rewire a lot of the house that he built and I bought so he could retire. There have been a lot of electrical issues that the inspection never found.
He hated outlets, and hardwired almost every appliance directly into whatever circuit was closest that fit the amperage, including literally cutting the plug off a cord, to nut the wires directly into the Romex.
I still haven't found what circuit my range hood is on. I turned off all 120V and it remained blowing, and knowing grandpa, he powered it off a single leg of a 240V. And I know he didn't use junction boxes, so the splices are likely all inside the walls.
tyfon 1 days ago [-]
In the US they have a 120V system so it might actually saturate a normal socket over there. My PC room has a 16A fuse and 230V though, so should be plenty :)
AshleyGrant 16 hours ago [-]
The US has a 240V system for residential power, not a 120V system.
Most homes in the US have 200 amp service. That's 48 kW of power available to the house, though code states that you can only pull 80% of that continuously, so 160 amp/38.4 kW.
The 120V misconception comes from the fact that it is delivered as split phase on two 120V legs. Any competent electrician can run a 240V, 50-amp (40 amps/9.6 kW continuous) circuit to any room in a house. Plenty of homes have these circuits for electric ranges, EV chargers, or RV power outlets, and there's nothing stopping anyone from having one of these circuits installed in whatever room they'd like it installed in.
Heck, my old landlord and I installed one ourselves to provide an outlet for an EV charger in the garage of a house I lived in about a decade ago.
400 amp and even higher service is available, though it is fairly uncommon. Plenty of large houses will have a 400-amp service, though, which means you can double the power numbers available that I mentioned above.
chollida1 15 hours ago [-]
Most North American homes would have 240v and 200Amp service. A 16Amp fuse seems down right tiny:)
We have 50 amp service running a hot tub and another 50 amp service line to charge an electric car.
tyfon 13 hours ago [-]
The 16A is for the PC room, the house has a 40A "400V" 3 phase TN connetion :)
The quotes are in since the difference between the phases is 230V, but in practice it is almost the same as a 40Ax400V single phase connection.
chollida1 3 hours ago [-]
gotcha. I've never heard of 16amp service, in Canada we have 15 amp.
But its very common to have many 15amp and 240V lines run in a house in Canada.
My kitchen alone has 4 different lines. i don't think your scenario is all that uncommon given that almost all new builds will have the same lines run in their home.
httpz 24 hours ago [-]
A typical hair dryer uses about 1500W. I guess GPUs are mechanically not much different from a hair dryer.
Crunchified 23 hours ago [-]
Install it in your bathroom, use it like a restroom hand dryer.
rbanffy 21 hours ago [-]
> use it like a restroom hand dryer
One that runs continuously
23 hours ago [-]
sidewndr46 22 hours ago [-]
isn't Ferrari the brand that requires any purchaser to be an existing owner? I could see NVIDIA going for something like that.
metadat 21 hours ago [-]
Invite-only is only for special edition hypercars and halo models. Standard models can be purchased by anyone with the funds and desire.
driverdan 19 hours ago [-]
Porsche is doing that for RS cars.
olyjohn 16 hours ago [-]
Yeah and now they have lost 98% of their profits in the last year. How's that catering to the ultra-wealthy working out for them? They used to build attainable cars that were nearly as cheap as a Corvette. Now they're double the price, and most certainly aren't double the performance.
rgmerk 16 hours ago [-]
From what I’ve read the only part of their business that is holding up is the 911 and particularly the high-end variants. It’s everything else that’s struggling.
fancyfredbot 1 days ago [-]
The relatively slow depreciation of GPU value is an artifact of supply constraints. If you run fp4 inference and could choose freely between Hopper and a Rubin, the performance per watt would make the Hopper unattractive even if you paid zero for the hardware and only for the power.
You can't get the Rubin, or even the Blackwell, so you will pay for the H100 but this won't last if fabs ramp up capacity.
tracker1 24 hours ago [-]
Not to mention, physical limits to lithography are slowing down significantly... so tech will continue to evolve more slowly... it'll never be the jump from 1080-1990 again, for example, even though 1990-2000 was pretty close, 2000-2010 much slower and since 2010 slower still.
What's as or more weird is how much hardware is backordered, and how much live hardware is allocated, but waiting on facilities for operation. And how many facilities are years behind at this point already... all on various credit and dept swaps between all the involved companies... it's not just a balloon, it's a house of cards balanced on a balloon.
chuckadams 22 hours ago [-]
The process node size in William the Conqueror's time was really off the charts.
rbanffy 21 hours ago [-]
Plotting that backwards shows a transistor would be the size of a continent, at least.
lelanthran 6 hours ago [-]
> Plotting that backwards shows a transistor would be the size of a continent, at least.
Means that the computer would be the size of earth... maybe the mice really were onto something :-)
mickael-kerjean 19 hours ago [-]
The infamous attention is all you need paper has this: "The Transformer ... reach a new state of the art in translation quality after being trained for as little as twelve hours on eight P100 GPUs.", the p100 is now quite a bit under 100$ and that's from a time where nvidia was valued 100x less than it is today. Lets wait and see
jbarberu 11 hours ago [-]
Beating the GPU performance bump from 1080-1990 would be hard indeed...
rbanffy 21 hours ago [-]
> since 2010 slower still
The increase in PFLOPS/dollar has continued accelerating, a lot from process, but also a lot by simplifying the architecture- if you had placed an H100 worth of transistors on a CPU-like architecture, you wouldn’t reach the same peak performances.
cyanydeez 23 hours ago [-]
supposedly, the next step is into fiber & optics.
but I generally agree, people put a lot of faith in the exponential leaps vs the exponential space.
You tell them we're not living on mars any time soon and they'll bring up christopher columbus.
wyre 22 hours ago [-]
How would fiber and optics help if the bottleneck is computation efficiency and not data transfer?
cyanydeez 18 hours ago [-]
the optics are doing the gates. literally, shine a light down a path and uts doing calcs.
theptip 16 hours ago [-]
One does not simply “ramp up fab capacity”.
This is all downstream of ASML who is putting out tens of EUV machines per year.
pianopatrick 22 hours ago [-]
Are fabs going to ramp up capacity? Wouldn't it make more business sense for them to just not ramp up capacity and enjoy the higher prices?
fancyfredbot 20 hours ago [-]
Yes they are going to ramp up capacity. It would only make business sense to do nothing if all your competitors were also doing nothing, which would probably require some level of illegal collision.
If everyone ramps up then in the best case everyone has the same sized slice of a bigger pie. So in theory it makes business sense. But the more realistic possibility is that you end up with oversupply, crash the market and everyone loses. This is what normally seems to happen with DRAM.
14 hours ago [-]
hinkley 24 hours ago [-]
I wonder what shovels were worth after the Gold Rush faltered. Blacksmiths probably had all the scrap iron they could ever care for.
NohatCoder 23 hours ago [-]
Nah, "selling shovels" is mostly a metaphor, the amount of iron that went into mining equipment was insignificant, beyond some local demand peaks. The majority of the business was consumables.
A key thing to understand about the gold rush is that it was not a major economic event, or at least nowhere near as big as the participants thought it would be, hence the tradegy.
The AI gold rush is different in that there actually is a mountain of "shovels" large enough to flood the global market quite severely.
oersted 13 hours ago [-]
Railway tycoons were the real beneficiaries, and the Gold Rush led to broader organic economic development in the West, so railways remained relevant.
Not sure what that means for the metaphor, but there you go.
gnfargbl 23 hours ago [-]
> The job of an operations team is to keep all of this in steady state. They know which racks run hot in summer, which cooling loops have been flaky since the last firmware update, which jobs to re-route when a node degrades but has not failed yet. None of that knowledge is written down. It lives in the team.
Hmmn. All of this information should live in the monitoring system, in which case any frontier model will be able to get to grips with it in short order. It feels like the author doesn't really fully understand the changes brought about by the systems they are writing about.
Planktonne 22 hours ago [-]
It's generated prose; the author's understanding doesn't factor in, because they didn't actually write it.
NohatCoder 11 hours ago [-]
I suspect you are right, but it is getting really difficult to tell the difference between AI and someone who simply don't know what they are writing about.
hellohackers127 21 hours ago [-]
[flagged]
AYBABTME 18 hours ago [-]
I don't get how this is unique to GPU clusters. As a general rule, underwriters are not qualified to operate and maintain the assets they underwrite loans for. That's why houses, cars, equipment, ... go at auction at a fraction of their value. And why lenders have insurance.
roadbuster 17 hours ago [-]
> I don't get how this is unique to GPU clusters
It isn't, and you're right. This is just a long-winded article by someone who thinks they've come across a deep, crucial insight.
I witnessed a bank foreclosure stemming from large, unpaid loans to a lumber mill. The bank absolutely didn't want to take possession of the operation, but had no choice when the founder decided to call it a day.
The bank had no idea what to do with finished lumber sitting in the drying ovens, let alone the entirety of mill infrastructure itself. After struggling to find a buyer, they hired the founder as a consultant to handle liquidation. The same would happen with a repossessed datacentre.
danielmarkbruce 3 hours ago [-]
In many cases they could literally do it but just don't have the scale - a small team of folks can finance a lot of real estate development for example. They can't do all the work of the development itself even if capable. It's just too much work.
venussnatch 18 hours ago [-]
Or... auctions are a terrible way to buy big expensive risky purchases.
Auction prices account for that risk, you don't get to do all of the verification you do for normal purchases.
marticode 17 hours ago [-]
[dead]
narrator 23 hours ago [-]
One data point: 512GB Mac Studios are selling on ebay for double what they were selling for new at the beginning of the year.
oersted 13 hours ago [-]
I don't understand why, that is not HBM nor a real GPU. Fine the model might fit in RAM, kinda. And it a bit like a server that you can keep on all the time, which must feel nice if you only ever had a MacBook.
But it can still not run truly capable coding agents at reasonable speeds right?
That doesn't make sense for the cost of quite a nice new car. If you want to spend that much, build yourself a proper workstation at least. Clearly a computer with the form-factor of an Apple TV is not going to be ideal for getting the highest performance for your buck.
Is it just wannabe vibentrepreneurs with too much spare money hyping each other up?
mtmail 22 hours ago [-]
https://news.ycombinator.com/item?id=49005798 "I have an M3 ultra mac studio with 512 GB of memory. I want to sell it, [...] Given the high cost of the Mac studio ($20,000 or more)"
girvo 20 hours ago [-]
My DGX Spark-alike is currently selling in my country for nearly double what I paid for it too. We live in a silly, silly world.
wmf 23 hours ago [-]
Anything with RAM in it is an appreciating asset.
trklausss 12 hours ago [-]
The article frames a GPU cluster as if it was rocket science. If you swap a GPU cluster with, say, a power plant, the question doesn't change a bit. Also high complexity, multiple variables, you don't know the current output power, scheduled maintenance, skilled labor needed etc.
In the end, they will either figure it out, or contract a company that can figure it out. Truth is, there will be many companies wanting to do the rental of those clusters for AI...
cmiles8 23 hours ago [-]
All indications are there will be a lot of repossessed GPUs appearing on the market before too long. Likely to be messy for a while but will open up a lot of possibilities when it’s easy to get your hands on some secondhand GPUs.
tomaskafka 23 hours ago [-]
This sounds like weapons market after the crash of soviet republic. Never has been a better time to get a functioning tank or parts of nukes.
17 hours ago [-]
pinkmuffinere 17 hours ago [-]
What are people going to do with all the GPUs? I guess we'll all have really good gaming rigs, lol.
Sincerely, perhaps there will be a new flood of crypto farms, which I suppose would drive the prices of crypto down. Other ideas for what changes are enabled by massive supply of cheap GPUs?
baalimago 5 hours ago [-]
Run AI locally instead of in cloud
acd 1 days ago [-]
Used GPU hosting for pension funds.
You go to ebay search for a used GPU. You get a price.
Neither is used servers a new thing or used routers. There are established used server companies.
I like Meg a lot a human, but Meg is all doom and gloom. Every single post she makes is about how GPUs fail [0] and now she's onto how financing is a big thing just waiting to crash and "nobody knows what a used GPU cluster is worth"...
Actually, we do, people offer them to me all the time. A used box of MI300x is $257k. "There is no GPU futures market"... actually there are a few of them that people have pitched to me.
This article is a lot of words from someone who isn't actually buying or deploying compute. My point is... take it all with a grain of salt.
Somehow at any given moment every possible topic to discuss on HN belongs to either the set of "it's amazing and nobody can say anything bad about it" or the set of "this thing sucks and nobody is allowed to say anything positive about it" and I never have ANY idea which one any given topic will be in on any given day.
fragmede 1 days ago [-]
latchkey is being restrained. He runs a data center filled with AMD GPUs. He's got a lot more insight to the business of it than the post does.
cgyvbunji 24 hours ago [-]
That's great, their comment should get more attention then.
I was complaining that it's obviously incorrect that nobody knows what used GPUs are worth, not about latchkey.
I have NO idea why anyone is upvoting a post titled "nobody knows what a used GPU cluster is worth", that is a WILD claim.
latchkey 23 hours ago [-]
she makes a lot of wild claims for clicks and gets to the front page with them... sigh.
owebmaster 22 hours ago [-]
In 2008, was the opinion of a banker more "insightful" than they opinion of journalists, bloggers and normal people talking about the imminent subprime crash?
CookieCrisp 21 hours ago [-]
Cherry picked your example there a bit
fragmede 24 hours ago [-]
latchkey is being restrained. He runs a data center filled with AMD GPUs. He's got a lot more insight to the real, lived experience of such a business than the post seems to have.
Der_Einzige 18 hours ago [-]
It’s always the people who know the least about GPU economics who want to doom post. Until I can reliably get on demand A100s in large quantities for less than 2.00 an hour there’s no AI bubble.
Most of the people who shit talk GPUs don’t even understand why it relatively speaking doesn’t matter that much if I’m on ampere or Vera Rubin for full precision models. These same people are wildly misinforming investors.
To quote the ADATA ceo, “don’t talk about an AI bubble until 2040”
tharmas 23 hours ago [-]
Aren't the AI Datacenters for the Digital Control Grid, so govt contracts?
gadders 6 hours ago [-]
A lot of big infra projects end up going bust, and then someone else buying up the company and its assets because they no longer have the original build costs to bear - EG EuroTunnel, Iridium etc.
Someone should stand up a vulture fund to buy these data centres off companies when they go bust.
somat 23 hours ago [-]
Isn't it worth exactly what you can sell it for. a few ways to do this.
slow and awkward, best market match: The auction. sell to highest bidder.
faster and more customer friendly but poor market match until a lot of units sold: The store. guess price, adjust up or down to reach sell frequency desired.
fast and good market match but takes a knowledgeable customer base: The reverse auction. Start with price too high lower it over time until it sells.
4 hours ago [-]
charcircuit 20 hours ago [-]
>adjust up or down to reach sell frequency desired.
>Start with price too high lower it over time until it sells.
These are the same strategy.
PeterStuer 8 hours ago [-]
Learning from crypto, a GPU's life foes not end because it fails, althougg ofc that happens, but because an alternative comes along that makes running the old equipment no longer power efficient under the new market conditions.
thhfgujgik5 8 hours ago [-]
Except, there's no real secondary market for SXM cards.
ponkpanda 23 hours ago [-]
Very difficult to sense check that substack post without access to the TLB credit agreement.
There are likely management service agreements from xAI proper -> SPV to cover precisely what the author talks about. Clearly, xAI could play games but without seeing the docs (which are not public), it's very difficult.
This article's basic point is right though. On the other hand, the LTV of this deal was approx 50% debt-financed (not too high; very much depends on the "V"). At 12.5%, it's not as if its being priced as a high quality asset.
Overall, substack post was too bearish. The wider point is that there's a lot of froth tied to what has now become systemically opaque - namely the circular deal flow that every hyperscaler, nvidia, neoclouds and friends are now engaged in. When the proverbial hits the fan, that stuff will be difficult to price and find few willing buyers with the competence to underwrite.
The systemic issues are the bigger concern than one specific deal imo.
guess_who_is 8 hours ago [-]
American AI corps will soon find out the worth once Huawei brings the cheap accelerators in the market. 6-9 months left.
Zigurd 24 hours ago [-]
Time to dig up the spreadsheet models of what dark fiber is worth.
effnorwood 4 hours ago [-]
I do. As much as they will pay.
jmyeet 21 hours ago [-]
This is a good article. I've been curious about how this is going to play out. A couple of data points:
1. An enthusiast had a project to get a V100 working on his PC [1]. This was a ~$10k GPU 10 years ago. It's now sold for scrap;
2. The A100 came out in 2020 and cannot run a large model like DeepSeek v4 Pro. It can run Flash. You need a 16xH100 cluster to run Pro and that's a ~4 year old GPU and AFAICT 8xB100 or 4xB200;
3. We're about to roll out R100/R200s.
I'm surprised that NVidia is moving to a 1 year product cycle (per this article) because the big question I've had is what's that going to do to existing investments in GPUs. Why? Because if 4xR100 can do the work of 32xH100 then that's a massive advantage in performance-per-Watt, which I think is going to be the only metric that ends up mattering.
In addition to raw power, new capabilities are developed and come online. For example, certain smaller, more efficient quantization methods just didn't exist on older hardware.
Oh, another thought from this: a 9% annual failure rate just goes to show you how ridiculous the idea of orbital data centers really is. Orbital DCs were always just a pump-and-dump scheme for SpaceX's IPO.
Currently it gets expensive to run models larger than ~31B locally. You start to need some pretty expensive hardware. That's going to change. I don't expect we'll be running 1T+ models on a Macbook Pro within 5 years (at reasonable inference rates) but I think people today will be shocked at what's being run locally in 5 years and that'll easily be 100-200B+ models.
Maybe this article has something useful to say but the painfully LLM-generated prose is too distracting to make it evident.
mrhottakes 1 days ago [-]
Did an AI make all the text and kerning so ugly? Why do web sites look like this now?
aqfamnzc 23 hours ago [-]
Isn't this just Substack?
ijidak 1 days ago [-]
How do you know? I feel that it's just poorly edited.
In my experience, AI is easier to read than this was.
Game_Ender 1 days ago [-]
Some sentences feel pretty AI like:
> These are not catastrophic events. They are the steady state.
> There is no GPU futures market, no standardized residual value curve, and no way to lock in a forward rental rate. The premium is is the price of underwriting in the dark.
The headings are also AI like, a lot of essays before usually did not have titled sections but now they do and they all feel like these.
In addition the diagrams themselves look pretty AI generated.
owebmaster 22 hours ago [-]
Soon we will have people asking if these comments are AI-generated.
chasil 19 hours ago [-]
Should I be worried that my plant is owned by Apollo?
No, because I am retiring next month!
andix 19 hours ago [-]
I'm wondering if all those GPUs ever end up on the second hand market for us to buy. Or if they will get refurbished multiple times and get used in second tier data centers until the chips die.
trollbridge 15 hours ago [-]
They’re available but rarely a good idea for the average person who wands small scale AI - for the same money you’re far better off getting an RTX 6000, 5090, or R9700.
>... when I left Paperspace in mid-2024, our M4000 GPUs, nine-year-old GPUs, were still consistently utilized at near-total capacity. That’s not a typo. Nine. Years. Old. Still booked, still working, still generating revenue.
As such you'd assume these cloud providers to want faster depreciation of their GPU assets rather than slower? I suppose in this case they do have an incentive to show bigger revenue numbers, but there seems to be a trade-off here that is not being discussed?
thelastgallon 12 hours ago [-]
Why would anyone buy these used when the new one, Vera Rubin can generate 10× more tokens per megawatt? Electricity costs are the biggest cost. More importantly, DCs are of a fixed power capacity. A power efficient GPU will mean it can fit in legacy DCs, don't need mega scale newfangled DCs. Old GPUs in the recent years should be worthless.
imtringued 9 hours ago [-]
The price of the used stuff will drop until both options look equally good.
The thing that might be difficult to understand is that the new operators might not run the GPUs at full capacity and therefore their energy bills are much lower than you're assuming. Hence the price doesn't drop to nothing, it stays at some non zero value.
1-6 24 hours ago [-]
The GPU cluster's RAM is probably worth more than the flops these days.
trollbridge 15 hours ago [-]
I’ve wondered when people will start finding a way to disconnect the HBM DRAM from e.g. old V100s.
Zarathustra30 22 hours ago [-]
Is the 30-50% of face value realistic for "Liquidation Value"? If one organization has to liquidate, sure. But if the bubble pops and many groups have to liquidate at once? Owners will be lucky if they can dodge the recycling fees.
1 days ago [-]
jgalt212 6 hours ago [-]
9% failure rate seems very high. What's the failure rate of a server CPU? Is the GPU failure rate high because it's 1000s of CUDA cores instead of 128 (Xeon)? If so, can't the H100 just work around the busted CUDA core(s)?
grim_io 1 days ago [-]
Nothing, because those "GPU's" are special proprietary hardware and are not what most people are capable of plugging in into anything.
cgyvbunji 1 days ago [-]
People are selling adapter boards to plug some data center GPUs into a regular pcie slot.
baalimago 5 hours ago [-]
Is this the CDOs of the AI-bubble?
AtlasBarfed 17 hours ago [-]
Well it certainly doesn't appreciate.
And I feel that AI hardware is going to rapidly depreciate faster than even typical PC hardware
jujube3 4 hours ago [-]
A lot of hardware literally has appreciated in the last year or two.
tim-tday 23 hours ago [-]
I’ll take “people who sell used hardware” for 500 Alex.
nikanj 22 hours ago [-]
Essentially nothing, fractions of a penny on a dollar, because players who could afford paying real money won't risk the crusty old hardware - so you're limited to buyers who still need a massive cluster but don't have AI infinite money glitch enabled
andix 19 hours ago [-]
Are failing GPUs really a big issue? I don't know how they behave when they die, but if it can be detected quickly, the affected nodes can just be removed from the pool.
claytonjy 17 hours ago [-]
I think there’s a bit more nuance and complexity here.
I built and operated an application that used 1-2000 L4 GPUs in production for a couple of years. Long-term GCP reservations running in a GKE cluster. At that scale we had a few GPU failures per week, and once saw 3 in a day. Nearly every failure required an engineer to manually intervene to get rid of the bad node, and then file a ticket with GCP support as they requested.
NVIDIA GPUs throw an “XID” code when they fail, which can be seen from serial port logs for GCP compute nodes (not through k8s!). If you’re lucky, it fires right as your application starts to fail, but there’s often a delay of several minutes. Even when you get an XID, by default GKE only responds to one or a few of them. You can expand the list via configuration, but the reconciliation loop is so slow that might still take 10-15 minutes during which a pod is puking errors and someone might be getting paged.
They’re working on it, and we never saw a single XID after we migrated to H100s, so the situation is improving. I imagine other clouds are even worse, though iirc Azure was leading some effort to improve k8s node problem detector to include accelerator problems so maybe they have a better story.
Training workloads are rather more sensitive to a node failure given that many modern training runs (including SFT etc) need multiple nodes where topology matters, and they might not have another e.g. 8xH100 box in the right place when one fails.
BenFranklin100 1 days ago [-]
Would it make economic sense to strip it down and sell for parts? That’s how it’s done now for older data centers, where the obsolete equipment is sent off to China, stripped down for parts, and sold on the secondary market. I’ve picked up several older but still useful RAID hardware cards off eBay this way.
I’m mainly interested in getting some DDR4/5 and RTX5090s on the cheap :).
petesergeant 8 hours ago [-]
“Nobody knows” = “the method sophisticated lenders are using to determine this is a commercial secret”
If you are the CFO of a credible AI company you can pick up the phone and get some advice exceptionally accurate numbers on what the market things the value is very quickly.
aslkalska 1 days ago [-]
so you're telling me you can't use 1 or 2% percent of 5 billions dollars to rebuild a team that runs GPU clusters for like a couple of years ?! and the guy that borrowed billions from these banks would want to mess up that relationship for what? I mean he is a stupid narcissist but not to that degree. This whole article makes no sense to me.
dfedbeef 14 hours ago [-]
\__
Joel_Mckay 24 hours ago [-]
Given the price for a rack of bc-250 after the crypto hype cycle, the expected value of the hardware will be around 5% to 10% of the original retail price.
Without other market influences, that is a >90% expected discount when the over-provisioned market must inevitably self-correct.
If the Market follows what Samsung/SK Hynix did to the South Korean exchange this week, than the "AI" bubble will hit harder than the dot com crash.
I like the Shrek Movie correlation theory, as they always happen just before Debt-backed investors get hit hard... And the new film is due out in 2027. =3
totetsu 12 hours ago [-]
I am hearing this point come up a lot. That in previous bubbles, there at least was useful infrastructure left in the aftermath. In the case of an AI bubble there is a lot of shortlived chips, and high maintenance infrastructure being invested into.
owebmaster 22 hours ago [-]
> If the Market follows what Samsung/SK Hynix did to the South Korean exchange this week, than the "AI" bubble will hit harder than the dot com crash
Can you tell us more about this? Or some link
HAL3000 19 hours ago [-]
Samsung and SK Hynix together account for around 60% of the Kospi's (SK stock exchange) market capitalization.
Over the past few weeks Kospi index has tumbled 25% since its June peak, resulting in a $1 trillion wipeout and its chipmaker duo have both lost at least 30% of their value. There have been days of near 10% plunges followed by sharp rebounds driven entirely by shifting confidence in whether AI spending is sustainable.
Joel_Mckay 19 hours ago [-]
Don't worry about it... I am more focused on the bizarre Shrek film timing phenomena, and pondering whether the pattern will hold again. =3
Patrick Boyle gives a summary of the situation, but not the underlying Shrek issue:
> Silent data corruption (SDC) is the most expensive, where a faulty GPU produces wrong answers without crashing anything, which means a multi-day training run can complete normally and the resulting model weights are quietly poisoned.
Anyone know how these get caught ultimately?
wmf 23 hours ago [-]
Usually silent data corruption isn't caught. If a file fails to open you might realize it's been corrupted.
timmmmmmay 24 hours ago [-]
[flagged]
mlyle 24 hours ago [-]
That still doesn't tell you what GPUs will be worth in 5 years, because it depends upon what inference demand is like and what the alternatives are. What will an hour of NVL72 be worth in 2029?
(And other things, that we know partially but not fully-- like what failure rate for current generation parts will be under this loading).
So we have big uncertainties about the revenue, moderate uncertainty about the proportion of the asset that will survive, and some uncertainty about what operating costs will be. It's difficult to turn this into a residual value.
Finally, the whole "operating the big facility" thing is not likely to be plug-and-play for a new technical team following a default. How much outage/disruption ensues?
timmmmmmay 18 hours ago [-]
perhaps! but you can look at what a five year old GPU (Nvidia A100 80gb) sells for today. it's not cheap! I wish it was!
mlyle 18 hours ago [-]
Yes, because we're currently in a GPU and RAM supply crunch, things have held value well compared to historical averages.
We are increasing production of GPUs and RAM a lot.
munchler 24 hours ago [-]
Nobody knows what anything is going to be worth in 5 years.
mlyle 22 hours ago [-]
Strictly true, but finance being able to predict this pretty closely is how the entire modern economy works.
throwaway85825 23 hours ago [-]
If that was true the insurance and loan industries wouldn't exist.
DiabloD3 24 hours ago [-]
GPU clusters have, largely, no actual value.
If anything, you might have to pay to have them disposed of, they don't really have any meaningful used eBay market outside of the randos that want to do high end extreme local inference in their basement.
Also, as for RAMmageddon, the inference SBCs that all of the AI bros bought don't have DIMMs, they're not even the right chip: its all GDDR and LPDDR. The only DDR DIMMs being consumed are for regular non-inference machines that help run the business and service infrastructure behind the scenes.
pier25 1 days ago [-]
so when will xAI default on its debt?
wmf 23 hours ago [-]
Never? They're currently renting out those GPUs at a profit.
owebmaster 22 hours ago [-]
Their debt expects a much (much!) bigger margin tho
cgyvbunji 1 days ago [-]
If nobody knows what a used GPU cluster is worth it means nobody is doing anything important with GPUs - time to short everything. Do you believe it?
That's for good NVidia H100 units.[1] There's a shortage of those. That seems to be the price after removal, cleaning, testing and refurbishing. Raw units removed from a shutdown will not be as valuable.
H100 units are available on eBay, but multiple sellers are using the same picture of a new unit in its original packaging, a bad sign.[2] Some even have pictures with the logos of a competitor.
[1] https://introl.com/blog/secondary-gpu-markets-buying-selling...
[2] https://www.ebay.com/shop/nvidia-h100-gpu?_nkw=nvidia+h100+g...
If you search ebay for server equipment like a Dell R840 with 768GB RAM, the same sort of dealers who are selling that and have thousands of feedback (at 98.5% of greater rating) are the ones I would consider much less risk.
From an annualized number on the llama 3 training report. would be interesting to see if we have a better idea given that we're already on rubin.
Also consider mixing in "was dropped during shipping" or "was stored improperly".
Why would hardware failure raise exponential with linear scale? Doesn't make sense
That's what you get from a constant failure rate per unit of time.
It's saying that hardware fails in the same way that Carbon-14 dating works.
Most if not all hardware failure on my side is def not memoryless besides the one lightning damage thingy but thats not hardwares fault.
When I tested "ECO" mode on my AMD CPU, performance was ~97% of regular mode, and temperatures dropped 5-10C. I think ECO mode drops TDP from something like 100W->65W. Cheaper and cooler to run for basically no cost.
If you are bitcoin mining or selling GPU capacity, you are incredibly conscious of your energy bill, so it only makes sense to optimize the power draw.
The difference is that much of the hardware crypto was running on was dirt cheap in comparison. Add to that a lot of crypto was being run on discount gear in cobbled together setups. Some people were even stealing electricity which often involved installing large clusters of relatively cheap devices in places without proper cooling.
Your are measuring two very different things.
And you are only considering electricity bills for the GPU. Higher TDP also means more heat which means more cooling.
It's not really easy to make an off hand comment about costs. A 30-40% decreaae in electricity usage and heat output for a fleet of GPUs definitely makes a massive difference to your opex bottom line.
Performance per Watt is the key metric for data centers full of GPUs.
https://en.wikipedia.org/wiki/Avalanche_effect
It was clean, cheap, and is still going strong for daily gaming.
In its working life it was undervolted and probably cooled better than in my rig.
It's wild to think that the system now is worth at least as much as I paid for it then if not much more than that. I saw a similar one going for $2500.
besides what intel did with Raptor Lake. This series had massive issues with oxidiation. Was the first time ever i became aware of this issue on scale.
Nontheless there are papers out there that silicon can degenerate and does.
I hope not though, perhaps I can pick up a H100 in a few years if they get sold on the open market.
why would anyone sign such a contract?
If you want to switch back to on prem, there's probably a way to structure acquiring hardware so it doesn't break the contract. Maybe you lease it, maybe the purchase happens through a related company, maybe there was no way for the contracted cloud to find out...
Plausible - the hardware might be leased, and so you have no right to modify it. You have to pay them to modify it.
Same as if you leased a car, you cannot do the services yourself, you have to pay them (and an approved agent of theirs) to do it.
Where folks (often) get lazy is the resulting math over what the real bean-counters care about (but are too lazy to check often).
In a past life, I worked on costing models for a Cable/Fiber contract house, to help the company decide 'what was profitable to keep in house' versus 'what do we subcontract' (sometimes that could even mean we just 'rented' a machine and had a qualified operator using it, based on that employee's hourly rate and expected L2R for taxes... so many spreadsheets...)
And from from my 'I don't know all the factors for this but I've seen how people screw up the big ones' view (and frankly, I'm guessing a lot of us have seen and dealt with the same category of 'bad math' around outsourcing IT work...)
An on-prem data center means:
- You need to account for electricity costs - i.e. CA vs midwest electric rates.
- cooling and power backup capability - Smaller factor but real
- personnel cost - e.x. there's probably cases where a smaller org could be better off with 'on-site' server admins that have other roles based on local wages. Kinda case specfic but it's a case.
- whatever the 'space' holding the stuff costs
Still, IMO a lot of orgs don't do the right math around these decisions, or just buy into the 'Well trends can change' as though they can decide as an org they need to suddenly triple capacity in a month and it would be able to organically happen in the first place.Frankly, the orgs that 'might' need that either have their arch set up where they are in cloud, or they are onprem but can scale to cloud if needed in interim.
It will be perfect for stuff like GPU-accelerated query engines, "classical ML" and every other CPU-based workload that could conceivably be offloaded to GPU
You can get used 16GB P100s on AliExpress for ~$100 if you want obsolete GPUs. Allegedly new AMD BC 250s are only slightly more.
I've looked at this some but I already have a GTX1070 which is only supported upto CUDA 11.9.
That's precludes some interesting modern optimizations out of the box. I've spend a lot of LLM tokens backporting some things, but I'm really not sure the hassle is worth it.
New hardware is just better. I think in maybe 5 years when supply and demand are back in equilibrium we are going to have some killer technology for decent prices, and 15yo H100s won't look attractive.
In a recent Gamer Nexus video with Level 1 Tech, they mention V100s are also quite useful for many applications that use FP64: stuff four in a workstation, and many PhD candidates would be quite happy with the throughput they can get for certain scenarios.
This is assuming that there is no big, big disruption to the semiconductor industry (e.g. TSMC getting attacked), in which case... well, I am gonna treat each stick of RAM I currently have like it's a faberge egg.
Is the idea that previously maintaining GPU programs was expensive whereas now AI makes it cheap? If so, I could buy that line of reasoning.
Maybe relatedly, I expect (hope) the hardware manufacturers will ramp up supply in the meanwhile which would also put downward pressure on GPUs. Right now though this hardware crunch is making me sad, not even because of GPUs but also because of general memory / disk.
TL;DR.
> I, for one, welcome the coming age of the post-LLM-datacenter-overinvestment-bust-fueled backyard GPU supercomputer revolution.
> The Big Question is…
> Who is cultivating the option to snap up and repurpose vapourised datacenter investments at fire sale prices, soon as the "datacenter debt" cometh calling?
For example, when I worked at a VFX software company, we were exclusively with one hardware partner. This unlocked something like a 50% discount across all our infra needs (they were big enough to provide switches, racks, servers, storage)
another company it unlocked a 75% discount.
By doing that, you know upfront what the value of your used hardware will be at the time you decommission it. It removes a lot of the risk for buyers in a volatile market.
a) It's not fraudulent, and
b) Nvidia has already signed on to buyback any unused capacity from the DCs it is selling to.
I don't think they need some special protection against this kind of contract.
I don’t understand the desire to tell totally anonymous third parties the terms under which they are allowed to deal with each other.
Right, so they're not voluntary.
1. There is one supplier, so you have no choice. 2. Even if you had a choice to sign the contract, this still means that it's not the same as a trade-in, because trade-ins are always voluntary, but once you have signed the contract, a right of first refusal is not.
In general, the "you chose to sign the contract" argument is a poor justification for bad contracts. If the contract is bad, it is bad regardless of whether you chose to sign it.
As businesses are expected to be more informed and equal in the negotiations
Maybe someone could start a business buying up and rehousing these.
And precisely because it's such a huge headache to do yourself, I think a small company could make a nice business wrapping up used datacenter cards in that sort of server.
The electricity prices are relevant because if you paid $0 for your H100 and didn't use it a single time, you could buy millions of tokens in inference just on the electricity it draws while idling. If you can't keep that thing saturated through the night, you're probably underwater overnight. Likewise, it's too small to run even the frontier open source models so you need to be able to live with worse models.
Max power matters because you aren't going to run many of those H100s before you blow breakers in most houses. Newer houses in the US are 15A service to non-kitchen breakers, so 1650W (that might be peak rather than continuous, not sure). If you're plugging that into an existing run, you could maybe run 2 before you start blowing breakers? You can't just plug 4 H100s into the wall in a normal house.
Maybe I'm wrong, though. I'd be curious, it'd be neat to run my own inference for something more than what'll run on a 3080.
A electrician can plop in a electric car charger for instance, that is a 40-50 amp circuit.
But I never assumed it was to be cost competitive at current token rates. I assumed it was for the same reasons people might use open source hardware. Freedom to tinker, etc.
Models are too big now
if you haven't heard a 5u server intended for a datacenter rack come to life it's quite the experience. Sounds like a plane taking off.
Helicopter. I live in Yeovil, Somerset, UK - there's a helicopter factory just down the road. I had a IBM "AS/400" or whatever they are called now in our computer room rack for a customer and it made nearly as much noise as everything else put together. It was clearly tuned for start up noise to impress because they would fire up in sequence, rise to a crescendo and then slow down in sequence to just a din instead of painfully loud.
A switch or PC server on boot will normally run up cooling fans instantly to max as a default protection mechanism until the "OS" has started and sensors read and then the fans will slow down to deal with the actual thermal load.
If you switch off your air conn, it gets noisy, quickly. Recently in the UK we are seeing routine temperatures around 30C and we broke 200 odd year records for temperatures a few weeks back. I know its even worse elsewhere but our infrastructure is not designed for this. Here we are at the same latitude as Calgary AB!
Until the thermal management kicks in, they sound like jet planes. When the thermal management starts and assesses the required cooling, it’ll throttle down the fans to reasonable levels.
That is, until the moment you push the machine to its limits. When then happens, you might get back to the same levels of the boot time, but it’ll require you to push everything to the max - CPU, memory, storage (all 24 bays) and so on. For a normal user, there is a lot of room and it’s virtually impossible, even with a dozen of Teams windows open.
If these do end up in Home Labs, it's going to be a server rack in the basement and not the server rack at the other end of the room.
If they are too loud, yes. In my part of the world basements are not common, but my shed would be a sensible place for one.
You don't need much of a fan if you are prepared to pump a lot of water.
I mean, sure, if you're trying to cool the thing with 100ml of water, then yeah, you need a fan.
OTOH I have an unused car radiator in my garage, a quiet and cheap pump, and space to hang that radiator outside the window. It's barely an afternoon's worth of work if the H100s already have liquid-cooling intakes/exhausts on them already (I dunno, I have never seen one in the flesh).
I'm pretty certain 10+ litres of constantly circulating water (antifreeze, in my case) through a car radiator would be sufficient to cool down 2x H100s. Hell, add in 10 RPM used car cooling fan and I can probably cool 20x of those running at full-bore.
If anyone wants to ship me a bunch of H100s, I'll happily build it all out, take pics and videos and post it back here. Just sayin' :-)
These things get hot and are fussy about their requirements.
1500W is the power of the typical American microwave or (tea) kettle at 120V. (Convert over to a NEMA 6 plug and dual-pole breaker and you can get 3000W.)
3000W is nothing special in the rest of the world with >200V wall plugs.
In North America we are on 120V, making a standard 15A outlet only 1500W max, and something like 1200W sustained. To use higher wattage appliances, we have to upgrade our outlets to 20A (2000/1600W) or up our voltage to 240V, but that carries a different set of plugs and outlets as well.
You’re limited to 24kW per leg if you ran everything only on 120V, but any really large server is going to be running off 240V anyways.
It's 1800W for short periods and 1500W sustained.
Realistically, a 15A breaker won’t trip on a 2,300 watt load for at least a few minutes, and often not until a few hours. A listed breaker is expected to trip in several seconds to 3 minutes for a 3,600 watt continuous load.
Take all those numbers and go up by 33% for a house or apartment with 20A circuits (quite common). Go up by 667% for an oven plug.
As anyone who takes care of rentals during winter knows, a typical circuit can tolerate two 1,500 W space heaters without tripping very much (although this is unsafe and a bad idea).
He hated outlets, and hardwired almost every appliance directly into whatever circuit was closest that fit the amperage, including literally cutting the plug off a cord, to nut the wires directly into the Romex.
I still haven't found what circuit my range hood is on. I turned off all 120V and it remained blowing, and knowing grandpa, he powered it off a single leg of a 240V. And I know he didn't use junction boxes, so the splices are likely all inside the walls.
Most homes in the US have 200 amp service. That's 48 kW of power available to the house, though code states that you can only pull 80% of that continuously, so 160 amp/38.4 kW.
The 120V misconception comes from the fact that it is delivered as split phase on two 120V legs. Any competent electrician can run a 240V, 50-amp (40 amps/9.6 kW continuous) circuit to any room in a house. Plenty of homes have these circuits for electric ranges, EV chargers, or RV power outlets, and there's nothing stopping anyone from having one of these circuits installed in whatever room they'd like it installed in.
Heck, my old landlord and I installed one ourselves to provide an outlet for an EV charger in the garage of a house I lived in about a decade ago.
400 amp and even higher service is available, though it is fairly uncommon. Plenty of large houses will have a 400-amp service, though, which means you can double the power numbers available that I mentioned above.
We have 50 amp service running a hot tub and another 50 amp service line to charge an electric car.
The quotes are in since the difference between the phases is 230V, but in practice it is almost the same as a 40Ax400V single phase connection.
But its very common to have many 15amp and 240V lines run in a house in Canada.
My kitchen alone has 4 different lines. i don't think your scenario is all that uncommon given that almost all new builds will have the same lines run in their home.
One that runs continuously
You can't get the Rubin, or even the Blackwell, so you will pay for the H100 but this won't last if fabs ramp up capacity.
What's as or more weird is how much hardware is backordered, and how much live hardware is allocated, but waiting on facilities for operation. And how many facilities are years behind at this point already... all on various credit and dept swaps between all the involved companies... it's not just a balloon, it's a house of cards balanced on a balloon.
Means that the computer would be the size of earth... maybe the mice really were onto something :-)
The increase in PFLOPS/dollar has continued accelerating, a lot from process, but also a lot by simplifying the architecture- if you had placed an H100 worth of transistors on a CPU-like architecture, you wouldn’t reach the same peak performances.
but I generally agree, people put a lot of faith in the exponential leaps vs the exponential space.
You tell them we're not living on mars any time soon and they'll bring up christopher columbus.
This is all downstream of ASML who is putting out tens of EUV machines per year.
If everyone ramps up then in the best case everyone has the same sized slice of a bigger pie. So in theory it makes business sense. But the more realistic possibility is that you end up with oversupply, crash the market and everyone loses. This is what normally seems to happen with DRAM.
A key thing to understand about the gold rush is that it was not a major economic event, or at least nowhere near as big as the participants thought it would be, hence the tradegy.
The AI gold rush is different in that there actually is a mountain of "shovels" large enough to flood the global market quite severely.
Not sure what that means for the metaphor, but there you go.
Hmmn. All of this information should live in the monitoring system, in which case any frontier model will be able to get to grips with it in short order. It feels like the author doesn't really fully understand the changes brought about by the systems they are writing about.
It isn't, and you're right. This is just a long-winded article by someone who thinks they've come across a deep, crucial insight.
I witnessed a bank foreclosure stemming from large, unpaid loans to a lumber mill. The bank absolutely didn't want to take possession of the operation, but had no choice when the founder decided to call it a day.
The bank had no idea what to do with finished lumber sitting in the drying ovens, let alone the entirety of mill infrastructure itself. After struggling to find a buyer, they hired the founder as a consultant to handle liquidation. The same would happen with a repossessed datacentre.
Auction prices account for that risk, you don't get to do all of the verification you do for normal purchases.
But it can still not run truly capable coding agents at reasonable speeds right?
That doesn't make sense for the cost of quite a nice new car. If you want to spend that much, build yourself a proper workstation at least. Clearly a computer with the form-factor of an Apple TV is not going to be ideal for getting the highest performance for your buck.
Is it just wannabe vibentrepreneurs with too much spare money hyping each other up?
In the end, they will either figure it out, or contract a company that can figure it out. Truth is, there will be many companies wanting to do the rental of those clusters for AI...
Sincerely, perhaps there will be a new flood of crypto farms, which I suppose would drive the prices of crypto down. Other ideas for what changes are enabled by massive supply of cheap GPUs?
You go to ebay search for a used GPU. You get a price.
Neither is used servers a new thing or used routers. There are established used server companies.
https://en.wikipedia.org/wiki/The_Emperor%27s_New_Clothes
Actually, we do, people offer them to me all the time. A used box of MI300x is $257k. "There is no GPU futures market"... actually there are a few of them that people have pitched to me.
This article is a lot of words from someone who isn't actually buying or deploying compute. My point is... take it all with a grain of salt.
[0] https://x.com/meggmcnulty/status/2040851080066859386
I was complaining that it's obviously incorrect that nobody knows what used GPUs are worth, not about latchkey.
I have NO idea why anyone is upvoting a post titled "nobody knows what a used GPU cluster is worth", that is a WILD claim.
Most of the people who shit talk GPUs don’t even understand why it relatively speaking doesn’t matter that much if I’m on ampere or Vera Rubin for full precision models. These same people are wildly misinforming investors.
To quote the ADATA ceo, “don’t talk about an AI bubble until 2040”
Someone should stand up a vulture fund to buy these data centres off companies when they go bust.
slow and awkward, best market match: The auction. sell to highest bidder.
faster and more customer friendly but poor market match until a lot of units sold: The store. guess price, adjust up or down to reach sell frequency desired.
fast and good market match but takes a knowledgeable customer base: The reverse auction. Start with price too high lower it over time until it sells.
>Start with price too high lower it over time until it sells.
These are the same strategy.
There are likely management service agreements from xAI proper -> SPV to cover precisely what the author talks about. Clearly, xAI could play games but without seeing the docs (which are not public), it's very difficult.
This article's basic point is right though. On the other hand, the LTV of this deal was approx 50% debt-financed (not too high; very much depends on the "V"). At 12.5%, it's not as if its being priced as a high quality asset.
Overall, substack post was too bearish. The wider point is that there's a lot of froth tied to what has now become systemically opaque - namely the circular deal flow that every hyperscaler, nvidia, neoclouds and friends are now engaged in. When the proverbial hits the fan, that stuff will be difficult to price and find few willing buyers with the competence to underwrite.
The systemic issues are the bigger concern than one specific deal imo.
1. An enthusiast had a project to get a V100 working on his PC [1]. This was a ~$10k GPU 10 years ago. It's now sold for scrap;
2. The A100 came out in 2020 and cannot run a large model like DeepSeek v4 Pro. It can run Flash. You need a 16xH100 cluster to run Pro and that's a ~4 year old GPU and AFAICT 8xB100 or 4xB200;
3. We're about to roll out R100/R200s.
I'm surprised that NVidia is moving to a 1 year product cycle (per this article) because the big question I've had is what's that going to do to existing investments in GPUs. Why? Because if 4xR100 can do the work of 32xH100 then that's a massive advantage in performance-per-Watt, which I think is going to be the only metric that ends up mattering.
In addition to raw power, new capabilities are developed and come online. For example, certain smaller, more efficient quantization methods just didn't exist on older hardware.
Oh, another thought from this: a 9% annual failure rate just goes to show you how ridiculous the idea of orbital data centers really is. Orbital DCs were always just a pump-and-dump scheme for SpaceX's IPO.
Currently it gets expensive to run models larger than ~31B locally. You start to need some pretty expensive hardware. That's going to change. I don't expect we'll be running 1T+ models on a Macbook Pro within 5 years (at reasonable inference rates) but I think people today will be shocked at what's being run locally in 5 years and that'll easily be 100-200B+ models.
[1]: https://www.hackster.io/news/hacking-a-server-grade-nvidia-g...
In my experience, AI is easier to read than this was.
> These are not catastrophic events. They are the steady state.
> There is no GPU futures market, no standardized residual value curve, and no way to lock in a forward rental rate. The premium is is the price of underwriting in the dark.
The headings are also AI like, a lot of essays before usually did not have titled sections but now they do and they all feel like these.
In addition the diagrams themselves look pretty AI generated.
No, because I am retiring next month!
Which has this anecdotal data point:
>... when I left Paperspace in mid-2024, our M4000 GPUs, nine-year-old GPUs, were still consistently utilized at near-total capacity. That’s not a typo. Nine. Years. Old. Still booked, still working, still generating revenue.
Also, I won't claim to understand accounting, but in general it seems it is advantageous to accelerate depreciation schedules for high CapEx industries because they lower taxes: https://leyton.com/us/insights/articles/what-is-accelerated-...
As such you'd assume these cloud providers to want faster depreciation of their GPU assets rather than slower? I suppose in this case they do have an incentive to show bigger revenue numbers, but there seems to be a trade-off here that is not being discussed?
The thing that might be difficult to understand is that the new operators might not run the GPUs at full capacity and therefore their energy bills are much lower than you're assuming. Hence the price doesn't drop to nothing, it stays at some non zero value.
And I feel that AI hardware is going to rapidly depreciate faster than even typical PC hardware
I built and operated an application that used 1-2000 L4 GPUs in production for a couple of years. Long-term GCP reservations running in a GKE cluster. At that scale we had a few GPU failures per week, and once saw 3 in a day. Nearly every failure required an engineer to manually intervene to get rid of the bad node, and then file a ticket with GCP support as they requested.
NVIDIA GPUs throw an “XID” code when they fail, which can be seen from serial port logs for GCP compute nodes (not through k8s!). If you’re lucky, it fires right as your application starts to fail, but there’s often a delay of several minutes. Even when you get an XID, by default GKE only responds to one or a few of them. You can expand the list via configuration, but the reconciliation loop is so slow that might still take 10-15 minutes during which a pod is puking errors and someone might be getting paged.
They’re working on it, and we never saw a single XID after we migrated to H100s, so the situation is improving. I imagine other clouds are even worse, though iirc Azure was leading some effort to improve k8s node problem detector to include accelerator problems so maybe they have a better story.
Training workloads are rather more sensitive to a node failure given that many modern training runs (including SFT etc) need multiple nodes where topology matters, and they might not have another e.g. 8xH100 box in the right place when one fails.
I’m mainly interested in getting some DDR4/5 and RTX5090s on the cheap :).
If you are the CFO of a credible AI company you can pick up the phone and get some advice exceptionally accurate numbers on what the market things the value is very quickly.
Without other market influences, that is a >90% expected discount when the over-provisioned market must inevitably self-correct.
If the Market follows what Samsung/SK Hynix did to the South Korean exchange this week, than the "AI" bubble will hit harder than the dot com crash.
I like the Shrek Movie correlation theory, as they always happen just before Debt-backed investors get hit hard... And the new film is due out in 2027. =3
Can you tell us more about this? Or some link
Over the past few weeks Kospi index has tumbled 25% since its June peak, resulting in a $1 trillion wipeout and its chipmaker duo have both lost at least 30% of their value. There have been days of near 10% plunges followed by sharp rebounds driven entirely by shifting confidence in whether AI spending is sustainable.
Patrick Boyle gives a summary of the situation, but not the underlying Shrek issue:
https://www.youtube.com/watch?v=nJtL9MBVj48
Anyone know how these get caught ultimately?
(And other things, that we know partially but not fully-- like what failure rate for current generation parts will be under this loading).
So we have big uncertainties about the revenue, moderate uncertainty about the proportion of the asset that will survive, and some uncertainty about what operating costs will be. It's difficult to turn this into a residual value.
Finally, the whole "operating the big facility" thing is not likely to be plug-and-play for a new technical team following a default. How much outage/disruption ensues?
We are increasing production of GPUs and RAM a lot.
If anything, you might have to pay to have them disposed of, they don't really have any meaningful used eBay market outside of the randos that want to do high end extreme local inference in their basement.
Also, as for RAMmageddon, the inference SBCs that all of the AI bros bought don't have DIMMs, they're not even the right chip: its all GDDR and LPDDR. The only DDR DIMMs being consumed are for regular non-inference machines that help run the business and service infrastructure behind the scenes.