Rendered at 11:54:24 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
mikewarot 14 hours ago [-]
It's my long standing opinion[1] that we shouldn't accept any RAM that can be subject to random bit flips. Most of the mitigations are just security through obscurity.
My computer with a lot of ECC DDR5 sometimes catches several bitflips in a day. I know this because they are reported to dmesg and sometimes I look at dmesg.
Intel decided long ago that you would need to pay more to not be subject to random bitflips. So it's an AMD system.
userbinator 6 hours ago [-]
My computer with a lot of ECC DDR5 sometimes catches several bitflips in a day.
Replace your RAM or find what's causing this disturbance in the environment. Sooner or later you'll get a pattern the ECC won't be able to correct.
Ironically if you didn't have ECC you would instantly RMA such RAM as defective because even a simple memtest would quickly discover the problem, and "several bitflips in a day" would cause very obvious crashes all the time.
e2le 14 hours ago [-]
I was shocked to learn that around 10% of Firefox crashes are caused by bit flips.
Most software is like this. Factorio had a weird one where some routers dropped UDP packets with zero checksum.
userbinator 6 hours ago [-]
100% agree. This is a pure case of defective product. Memory not behaving like memory.
It's very telling that when Rowhammer first appeared, authors of popular memory testing utilities added tests for it, and then quickly hid and disabled them by default because "too much RAM would test defective". I have no idea how they were convinced to do so.
SideQuark 5 hours ago [-]
Single event upsets are forced on you by physics. There is no possible way to stop them all.
reliabilityguy 8 hours ago [-]
> It's my long standing opinion[1] that we shouldn't accept any RAM that can be subject to random bit flips.
Every system fails eventually.
adrian_b 1 hours ago [-]
That is why every subsystem must have means to detect the failures, to allow the replacement of the bad components.
ECC in DIMMs achieves this.
For myself, this has been the greatest advantage of always using ECC, that I have been warned early about the great increase in the frequency of errors in some DIMMs, after many years (over 5 years) of continuous usage, which has allowed me to replace the offending DIMMs and continue to use those systems for some more years.
anonymous_user9 6 hours ago [-]
Yes things fail, but currently random bit flips are not considered a failure.
adrian_b 1 hours ago [-]
The big companies consider them a failure in their own computers, so they use appropriate memory, but they consider that "consumers" either must accept them as normal and not a failure, or pay a great overprice in comparison with what a big company pays, in order to have memory devices with error detection.
Retr0id 14 hours ago [-]
Intel MEE addressed this (https://eprint.iacr.org/2016/204.pdf), by treating DRAM as entirely untrusted (assuming an attacker with arbitrary physical read/write capabilities), but then they discontinued it.
Regarding your linked comment:
> If we had properly tested and validated RAM, RowHammer wouldn't work, ever.
While it's exacerbated by physical defects and tight timings, it's really a fundamental problem with how DRAM works. It's frankly a miracle it works in the first place.
inigyou 14 hours ago [-]
No. If it was designed really well, it would trigger extra refreshes as needed. It would know how much accessing a row disturbs nearby rows and how much disturbance a row can suffer before it needs a refresh to prevent data corruption. It would know how long a row can be held open for, too. There is no fundamental theorem that rowhammer must work. It was an engineering tradeoff.
I don't think we can blame DRAM designers for flying a teensy bit too close to the sun here, since this is no problem in normal operation and only appears under adversarial scenarios. We can blame them for not fixing it once it was discovered. And we can definitely blame Intel for making ECC RAM a market segmentation feature.
Retr0id 14 hours ago [-]
What you describe is the Target Row Refresh (TRR) mitigation, and it can be bypassed via "half-double" attack.
inigyou 13 hours ago [-]
Where did I imply that they should use an incorrect underestimation of row disturbance? If they actually measured these things, they'd know how far away a disturbance can be created before it becomes so weak the normal refresh cycle fixes it.
Retr0id 13 hours ago [-]
It is not feasible for DRAM (or the memory controller) to maintain a fully accurate simulation of DRAM. If it was easy they'd have done it.
gpm 13 hours ago [-]
It doesn't need to be fully accurate, merely always conservative. It's a question of cost (how much performance does the conservative estimate waste) not feasibility.
adrian_b 56 minutes ago [-]
This problem has never existed with bigger DRAM cells.
At some point in time, a few generations of DRAM ago, they have reduced the dimensions so much and without discovering adequate mitigations for the problems introduced by this, that the DRAM reliability has become inadequate.
The reason why they did this was to reduce the fabrication costs. It is likely that the pressure to reduce the fabrication costs has been caused more by the desire to increase the profit margins than by the intention to enable any price reductions, because even before the recent price increases there have been around 15 years with only negligible reductions in memory prices.
By increasing the fabrication costs, it would be easy to eliminate the RowHammer problem, while still having memory prices several times lower than the current prices.
However, the vendors do not want this. They want to find some kind of mitigation that would not cause any measurable increase in the fabrication costs. Until now they have failed to do this, but it is not clear how hard they have tried.
It is very likely that their failure to find anything that works has been caused in a good part by the secrecy that is typical for nowadays.
In the earlier times of the semiconductor industry, every manufacturing problem was described in public research papers, with complete details, and usually the right solution was found by someone else and then it spread quickly in all the industry, with much less concerns about "IP" than today.
Only this openness has allowed the creation of the successful semiconductor industry and of the "Silicon Valley".
inigyou 13 hours ago [-]
Okay then refresh the whole thing and not just a single row whenever TRR is triggered.
Or refresh the whole thing after a certain number of row activations.
Or interleave activations with refreshes. Say, after every 5 activations, refresh the next row in the refresh cycle.
I'm not a DRAM expert. They can figure it out. I promise if you refresh the whole chip after every row activation you won't have rowhammer. It'll be too slow though. Somewhere in between is the fastest point where there isn't rowhammer.
reliabilityguy 8 hours ago [-]
> Okay then refresh the whole thing and not just a single row whenever TRR is triggered.
How would you refresh the whole DRAM? Refresh is simply reading the row and writing it back, it is sequential in nature.
> Or interleave activations with refreshes. Say, after every 5 activations, refresh the next row in the refresh cycle.
It makes zero sense. Refresh is disturbance in itself. You are entering a recursion here where the mere act of refreshing increases the counter values of victim rows making refresh even more frequent. At the end you have DRAM that you cannot read from or write to at all because it’s always refreshes itself.
> I'm not a DRAM expert
Yes
simoncion 7 hours ago [-]
> At the end you have DRAM that you cannot read from or write to at all because it’s always refreshes itself.
Be so kind as to run the math on this for me? There must be something that I'm missing, because it looks like you're describing a system in which a refresh causes so much disturbance that another refresh is immediately required. Such a system seems incapable of reliably storing a byte and reading it back. The colloquial term for such a system is "unreliable garbage".
13 hours ago [-]
simoncion 7 hours ago [-]
> If it was easy they'd have done it.
Why? Something like 70->90% of the world's DRAM is produced by three companies that are very friendly to each other. It makes no business sense to fix problems that you know your peer companies will not fix... that's spending money that you absolutely do not have to.
If the business/economic theory doesn't sway you, look way back to what happened to ISP speeds and pricing when Google Fiber so much as credibly threatened to start providing service in an area served by a mono/duo/triopoly. Or -more recently- how SpaceX demonstrated that the defense-contractor-owned space launch companies had spend decades wasting enormous amounts of taxpayer money by refusing to do any significant amount of research into bringing the cost to launch down substantially. ISPs and the space launch companies had no peers that would spend the resources required to provide a better and/or cheaper service to their customers, so it made absolutely no sense for any one of them to spend resources to break the truce and make them all far less money in the long run.
inigyou 3 hours ago [-]
Isn't this a memory controller concern, anyway?
simoncion 1 hours ago [-]
I don't know. But I do strongly suspect that if rowhammer could be eliminated by a change in memory controller behavior, it would have been eliminated long ago.
bell-cot 13 hours ago [-]
It doesn't need to be a fully accurate simulation, any more than your car needs a fully accurate ICE simulation to decide "Change the engine oil".
Just count accesses to each row, and refresh the adjacent row(s) when some limit is reached.
inigyou 12 hours ago [-]
Half-Double is an attack where you hammer the rows 2 spaces away so that the automatic refresh on the rows 1 space away is what actually hammers the target row.
bell-cot 12 hours ago [-]
Yes? When a given row's counter hits the limit, you refresh every other row which your actually-competent testing has shown might maybe possibly be compromised by that activity.
Same as a good car's computer doesn't make a bunch of rosy assumptions about whether the oil needs changing or not - the automotive engineers actually do their jobs, test the crap out of their engine designs, and base the oil-change criteria on the real-world test results.
(JIC: "Adjacent" has a range of meanings in English. Those go from "the singular closest neighbor, within further constraints on type, orientation, etc." to "relatively close to by some metric". But I am not an EE, let alone a chip architect, to know the "real insider" lingo here.)
nomel 6 hours ago [-]
The later I get into my career the more I'm convinced that the biggest engineering challenge is to convince others there's a problem to begin with. This is also why I think innovation generally only happens in smaller companies: less convincing.
inigyou 3 hours ago [-]
Which is the whole chip, because if you only refresh any limited number of rows, the next row outside that range becomes a viable indirect target as shown by Half-Double.
Alternatively you would also count the refresh as a hammer on its adjacent rows.
- The entire chip is already refreshed every 32ms to 64ms, because the capacitors which implement DRAM lose their charges over time.
- The time required to induce an exploitable bit flip is (in one system tested) was ~22ms
So: Even if it was the whole chip - vs., say, the 6 nearest rows to the highly-accessed row - the more-frequent refreshes would not be a big deal.
> Alternatively you would also count the refresh ...
Sure. Or once any row access counter triggers a refresh, extend that to every row with a row access counter within (say) 25 of its limit. And if that ends up refreshing more than (say) 25% of the chip, then just refresh the entire chip.
anon_rh_guy 12 hours ago [-]
Little known fun fact: the MEE's checksum updates... were one of the most successful Rowhammer patterns... :-p
Retr0id 11 hours ago [-]
Interesting - but presumably used to attack rows outside of SGX?
[1] https://news.ycombinator.com/item?id=33860477
Intel decided long ago that you would need to pay more to not be subject to random bitflips. So it's an AMD system.
Replace your RAM or find what's causing this disturbance in the environment. Sooner or later you'll get a pattern the ECC won't be able to correct.
Ironically if you didn't have ECC you would instantly RMA such RAM as defective because even a simple memtest would quickly discover the problem, and "several bitflips in a day" would cause very obvious crashes all the time.
https://news.ycombinator.com/item?id=47252971
It's very telling that when Rowhammer first appeared, authors of popular memory testing utilities added tests for it, and then quickly hid and disabled them by default because "too much RAM would test defective". I have no idea how they were convinced to do so.
Every system fails eventually.
ECC in DIMMs achieves this.
For myself, this has been the greatest advantage of always using ECC, that I have been warned early about the great increase in the frequency of errors in some DIMMs, after many years (over 5 years) of continuous usage, which has allowed me to replace the offending DIMMs and continue to use those systems for some more years.
Regarding your linked comment:
> If we had properly tested and validated RAM, RowHammer wouldn't work, ever.
While it's exacerbated by physical defects and tight timings, it's really a fundamental problem with how DRAM works. It's frankly a miracle it works in the first place.
I don't think we can blame DRAM designers for flying a teensy bit too close to the sun here, since this is no problem in normal operation and only appears under adversarial scenarios. We can blame them for not fixing it once it was discovered. And we can definitely blame Intel for making ECC RAM a market segmentation feature.
At some point in time, a few generations of DRAM ago, they have reduced the dimensions so much and without discovering adequate mitigations for the problems introduced by this, that the DRAM reliability has become inadequate.
The reason why they did this was to reduce the fabrication costs. It is likely that the pressure to reduce the fabrication costs has been caused more by the desire to increase the profit margins than by the intention to enable any price reductions, because even before the recent price increases there have been around 15 years with only negligible reductions in memory prices.
By increasing the fabrication costs, it would be easy to eliminate the RowHammer problem, while still having memory prices several times lower than the current prices.
However, the vendors do not want this. They want to find some kind of mitigation that would not cause any measurable increase in the fabrication costs. Until now they have failed to do this, but it is not clear how hard they have tried.
It is very likely that their failure to find anything that works has been caused in a good part by the secrecy that is typical for nowadays.
In the earlier times of the semiconductor industry, every manufacturing problem was described in public research papers, with complete details, and usually the right solution was found by someone else and then it spread quickly in all the industry, with much less concerns about "IP" than today.
Only this openness has allowed the creation of the successful semiconductor industry and of the "Silicon Valley".
Or refresh the whole thing after a certain number of row activations.
Or interleave activations with refreshes. Say, after every 5 activations, refresh the next row in the refresh cycle.
I'm not a DRAM expert. They can figure it out. I promise if you refresh the whole chip after every row activation you won't have rowhammer. It'll be too slow though. Somewhere in between is the fastest point where there isn't rowhammer.
How would you refresh the whole DRAM? Refresh is simply reading the row and writing it back, it is sequential in nature.
> Or interleave activations with refreshes. Say, after every 5 activations, refresh the next row in the refresh cycle.
It makes zero sense. Refresh is disturbance in itself. You are entering a recursion here where the mere act of refreshing increases the counter values of victim rows making refresh even more frequent. At the end you have DRAM that you cannot read from or write to at all because it’s always refreshes itself.
> I'm not a DRAM expert
Yes
Be so kind as to run the math on this for me? There must be something that I'm missing, because it looks like you're describing a system in which a refresh causes so much disturbance that another refresh is immediately required. Such a system seems incapable of reliably storing a byte and reading it back. The colloquial term for such a system is "unreliable garbage".
Why? Something like 70->90% of the world's DRAM is produced by three companies that are very friendly to each other. It makes no business sense to fix problems that you know your peer companies will not fix... that's spending money that you absolutely do not have to.
If the business/economic theory doesn't sway you, look way back to what happened to ISP speeds and pricing when Google Fiber so much as credibly threatened to start providing service in an area served by a mono/duo/triopoly. Or -more recently- how SpaceX demonstrated that the defense-contractor-owned space launch companies had spend decades wasting enormous amounts of taxpayer money by refusing to do any significant amount of research into bringing the cost to launch down substantially. ISPs and the space launch companies had no peers that would spend the resources required to provide a better and/or cheaper service to their customers, so it made absolutely no sense for any one of them to spend resources to break the truce and make them all far less money in the long run.
Just count accesses to each row, and refresh the adjacent row(s) when some limit is reached.
Same as a good car's computer doesn't make a bunch of rosy assumptions about whether the oil needs changing or not - the automotive engineers actually do their jobs, test the crap out of their engine designs, and base the oil-change criteria on the real-world test results.
(JIC: "Adjacent" has a range of meanings in English. Those go from "the singular closest neighbor, within further constraints on type, orientation, etc." to "relatively close to by some metric". But I am not an EE, let alone a chip architect, to know the "real insider" lingo here.)
Alternatively you would also count the refresh as a hammer on its adjacent rows.
Here's the actual Half-Double paper:
https://www.usenix.org/system/files/sec22-kogler-half-double...
Note two things:
- The entire chip is already refreshed every 32ms to 64ms, because the capacitors which implement DRAM lose their charges over time.
- The time required to induce an exploitable bit flip is (in one system tested) was ~22ms
So: Even if it was the whole chip - vs., say, the 6 nearest rows to the highly-accessed row - the more-frequent refreshes would not be a big deal.
> Alternatively you would also count the refresh ...
Sure. Or once any row access counter triggers a refresh, extend that to every row with a row access counter within (say) 25 of its limit. And if that ends up refreshing more than (say) 25% of the chip, then just refresh the entire chip.