Assembly Hall of Shame (github.com)
282 points by piotrgrabowski 10 hours ago
56767865678 a few seconds ago
Sahil
Retr0id 9 hours ago
Related, and linked in the readme: https://github.com/xoreaxeaxeax/smiiiiiiiiiiiiiiii (using the slow instructions to break SMI)
jonathrg 7 hours ago
I wish they would just explain it in normal terms instead of this nasty LLM "engaging blog post" style
twothreeone 4 hours ago
Chris usually takes an educational angle, I don't think this is LLM-generated content at all it's just his style. I highly encourage watching some of his DefCon or BlackHat talks, they're fun!
Brian_K_White 3 hours ago
I looked at both links and don't see anything weird or annoying, and I hate overblown styles myself.
kazinator 2 hours ago
Bus cycles can be arbitrarily long on any processor that has memory cycles with a hand shake requiring an ack, with no timeout.
E.g. we can build a board around a MC68000 where we make it lock up forever in a bus cycle, waiting for a DTACK that doesn't arrive.
Some early microprocessors had clocked bus cycles without handshaking. They would put out an address on some address lines and signal some line together with a read/write indication, and then expect the transfer to be completed within some clock cycles. If nothing is attached to the address, they would read whatever values are on the bus, like maybe all 1's if it is an open drain system that requires the transmitting device to pull to ground to indicate zero.
I'd say that kind of thing belongs to a hall of shame; it requires software hacks to interface with anything that can't keep up with the prescribed bus cycle.
monocasa 7 hours ago
It says in the rules
> Trapped/emulated/virtualized instructions may only time the trap, not the handler.
But I feel like that 12ms write to an ACPI IO port at current leaderboard position 8 is probably trapping to SMM and being handled there.
layer8 9 hours ago
Nop should be #1, because it is infinitely slow for what it does. ;)
jooops1 9 hours ago
It increments rip by one.
EvanAnderson 5 hours ago
I thought I remembered reading somewhere re: the 8086 microcode disassembly that NOP, which is encoded as XCHG AX,AX actually does run the XCHG microcode and uses an internal scratchpad register to do the exchange.
JoeAltmaier 5 hours ago
fluoridation 6 hours ago
No, that's done by the decoder. It actually does nothing.
loeg 4 hours ago
dlcarrier 3 hours ago
russdill 6 hours ago
hyperhello 2 hours ago
It's a little faster than yep.
mito88 8 hours ago
Strategy: nop does nothing. It opens the leaderboard accordingly.
Score: 1 cycles Time: 0 nanoseconds
layer8 8 hours ago
It opens the leaderboard as #27, so in the last place.
TomatoCo 10 hours ago
This author also has other things like: A compiler that emits only `mov` instructions and another compiler that deliberately messes with the control flow so that, if disassembled, common debuggers will draw symbols like skulls or threats. https://github.com/xoreaxeaxeax/repsych
inigyou 9 hours ago
He also bruteforced the entire opcode space to find undocumented instructions (sandsifter).
codeshaunted 10 hours ago
what im seeing from this chart is that we should be using the nop instruction for everything
bee_rider 9 hours ago
Well the best code is no code. Nop could be second best though.
inigyou 9 hours ago
Instructions unclear. Set the NX bit to ensure no code, and got a general protection fault.
markus_zhang 8 hours ago
Does that mean Chris Domas is ready for his next adventure?
simonebrunozzi 6 hours ago
Related, somehow: Core War [0].
michalsustr 9 hours ago
Very cool! Also, huh interesting. I’ve used rdtsc to measure cycle diffs but had no idea its execution takes that long. Is that common across architectures?
pbsd 7 hours ago
The cycle count for RDTSC is ~25 cycles on Skylake-era microarchitectures. The 49 number shown in the OP seems off.
inigyou 9 hours ago
AFAIK it acts as some kind of execution barrier, to give meaningful timing.
rrampage 9 hours ago
Isn't that rdtscp ( https://www.felixcloutier.com/x86/rdtscp )?
metadat 10 hours ago
It’s crazy how computers still seem to get perceivably slow every few years, given how many instructions can be executed in 1ms. Shameful, even..
What’s that law called about programmers wasting all the compute on abstraction?
AceJohnny2 5 hours ago
There are two aspects to compute performance: latency and throughput.
There has been a ton of work to improve throughput, but that's often come at the cost of latency, because a great technique to improve throughput is batching, to avoid the per-task overhead cost.
We've also added many layers of abstraction.
A seminal example is Dan Luu's "computer latency" table (2017) https://danluu.com/input-lag/, which measures the time between a keypress and visible changes on the screen, and wherein an Apple IIe has latency 5x lower than a Lenovo X1 Carbon on Windows.
From some perspective, you could say that the Lenovo on Windows example is an impressive feat of engineering, considering all the subsystems involved (USB, interrupt dispatcher, input layer, windowing system, double-buffering...)
This throughput-over-latency tradeoff can be seen across the entire computer landscape design.
loeg 4 hours ago
Idk man, my computers don't seem to get any slower over time -- no upgrades to any components, either.
metadat 4 hours ago
Does responsiveness get better or worse with updates, in general?
loeg 2 hours ago
mwigdahl 9 hours ago
Wirth's Law I believe.
inigyou 8 hours ago
And remember to call him by name, not by value!
HappyPanacea 10 hours ago
The new windows notepad is a disgrace
adamrezich 8 hours ago
The new mspaint fucked, then unfucked, then refucked my decades-old muscle memory of Win+R mspaint Enter Ctrl+E 1 Tab 1 Enter Ctrl+V to open Paint, resize canvas to minimum, then paste from clipboard. When you press Ctrl+E now, the Units control is selected by default, for some completely asinine reason!!!
fluoridation 6 hours ago
EvanAnderson 4 hours ago
sitzkrieg 6 hours ago
inigyou 9 hours ago
Andy and Bill's Law
summarybot 9 hours ago
The OS should do less not more
RiverCrochet 6 hours ago
vi should be a kernel-level system call, tunable with sysctls. For agentic management.
LoganDark 10 hours ago
Huh? A millisecond is an eternity!
m463 9 hours ago
I remember reading once somewhere:
If some app responds in 10ms or less, it is INTERACTIVE.
makes you think.
Xirdus 8 hours ago
faresahmed 6 hours ago
vardump 10 hours ago
A great resource for any performance deoptimization.
baddash 6 hours ago
just curious, how much do these actually discover useful practices or pitfalls, on top of just being for fun?
spoocecow 9 hours ago
Oh wow, glad to see Chris Domas active online again!
achierius 10 hours ago
It'd be really interesting to see whether the winning (losing?) instructions/strategies would be different on other architectures. At least right now the top spot (`fxrstor64` on MMIO, starve PCIe) seems relatively architecture-independent, but maybe something about MMIO ordering rules on e.g. POWER would be different enough to change that -- or perhaps open up new avenues?
I wonder what the actual limit on this `fxrstor64` is right now. If you can stall the PCIe bus for that long, then why not indefinitely? Certainly there's no forward progress guarantee here.
IshKebab 8 hours ago
Using MMIO is cheating and makes the results very boring.
It would be much more interesting to know the results if you're only allowed to use main memory.
arn3n 10 hours ago
There’s definitely strategies here; A lot of the floating point operations use subnormals, and a lot of the worst instructions are slowed down by really, really fucking with MMIO.
eek2121 6 hours ago
This is neat!
darksim905 2 hours ago
Seems like spam from this creator since there are two things on the front page?
john_strinlai 2 hours ago
submitted by two different people, both with year+ old accounts and decent karma. i dont think either is the author. the other submitter probably read this one, looked at the github, saw something else cool and posted it. (i almost did the same, but bookmarked it instead)