There are literally thousands of compiler engineers who pore over the assembly a compiler generates, and then finds ways to make it better. I get paid to figure out where the compiler can do better, and often a step in that is to hand-code my own replacement. I then teach the compiler to do that.
But even beyond that, the compiler can't make certain assumptions that an assembly writer can. Such as whether a callee-saved register really does need to be saved in some particular routine. Or even pushing an extra parameter in unusual cases.
So it is entirely possible to beat the compiler, it's doable under certain circumstances, even today.
But you also have the danger of your loving hand-crafted assembly beating the compiler today. But next year the compiler is even smarter, the hardware may have changed in subtle ways, and the compiler will know and improve the code it generates.
Your hand-written code won't change unless you revisit it.
With hand-written assembly, you can (with effort and care) ensure that the code runs in constant time and doesn’t leak any information through side channels (such as which memory or cache addresses it accesses). That’s important for most encryption code. It’s difficult to ensure constant time execution in pure Rust (or C, or most high level languages in general) because you can’t tell if some future compiler optimization will break your attempts at constant-time code.
An optimizer is trying to balance between compile time, runtime speed, and code size, and most optimizations will win you one axis at the cost of one or both of the other axes. (The rare optimizations that win on all 3 are already all implemented in the compiler.) Compiler developer time is also a scarce resource; I know of so many more optimizations I could implement, but without demonstrable code that would actually benefit, it's not a good use of my time to implement them. Compilers tune this balance by making lots of heuristic decisions, and these heuristics are tuned by large benchmarks, which often times involve a lot of flat code profiles (i.e., no code is worth spending a lot of time really nailing down the best code layout).
One of the advantages of hand-written assembly is that you get to opt out of the compiler heuristics and commit to being able to spend the time to optimize the one bit of code that you know is really important for runtime as perfectly as you want, instead of relying on the compiler to get it close enough to perfect before it exhausts its budget of caring about optimizing it.
You can usually hint the compiler, but the best idea is to provide it with more data. PGO with real-world data and LTO for whole-program analysis enable it to do much better decisions.
1. Better memory locality by knowing what you load and when, exactly.
2. The ability to "cheat" on calling conventions.
3. The ability for techniques like threaded code, and in general, better cache-awareness.
4. Less mov's.
5. Guaranteeing no spilling in important loops.
6. Compilers don't do well with flags registers and you can't read/write them in high level languages. You're hoping your `if (result < a) { carry = 1; }` becomes a direct flag test. Especially important in bignum, you can't really utilise adcx/adox directly from high-level code.
7. Hot/cold layout without PGO. Yes PGO is good but sometimes you know better and PGO isn't very suitable for "configurable" code.
9. Exploiting uninitialised memory for classic party tricks like not initialising a buffer fully (let's say you have a library function with a return buffer. You don't want dynamic allocations for some reason. You can simulate this with a pointer return into a let's say a static 4KB buffer and a count return, you only initialise it until the count. Caller has the responsibility not to overread.)
"We’d like to initialise our Vec in parallel, otherwise we’d have to wait for the main thread to fill the entire Vec with a placeholder value only to then have our threads overwrite those placeholder values."
Talk about overengineering :P Multithreaded vector initialisation instead of just...skipping it?
I'm so far downstream from the source language that I know of none specific to Rust. And note someone else's specific answer to your question above that has nothing Rust specific either.
> Such as whether a callee-saved register really does need to be saved in some particular routine.
Ironically this is your preconceived abstraction of how a compiler has to operate. An ideal compiler could allocate registers differently for each called function: F1()->F2()->F3(), F1 uses r0-5, F2 uses r6-10, F3 uses r11-15, no register saving required in the whole chain. There's no need for a fixed ABI. Such a compiler would look very different from today's ones.
Modifying the ABI of a function requires being able to track down all of the call-sites of the function, which is less trivial than you might assume. ABI concerns also tend to baked in relatively early in the optimization pipeline because you just simply can't get the ABI wrong, and I can think of several instances where the ABI decision causes missed optimizations.
There is also the other issue that a good algorithm for optimizing a problem like register allocation tends to be super-linear (e.g., quadratic), and if you shift the model from "allocate on a per-function basis" to "allocate all functions", the N in the O(N²) goes from "size of function" to "size of program," which is now suddenly a lot more compiler time spent for very modest gains. If register spilling across a function call is a noticeable component of runtime, then you're probably better off inlining that function in the first place!
One of the hardest parts of optimizations is not being penny wise and pound foolish. This sort of thing is a great example of that.
Additionally, a hard part is that all of this can change over time with new hardware! Some patterns that were crucial before everything gained branch predictors are irrelevant now, etc.
Yeah, that's what MSVC and GCC did on x86, called "custom calling conventions" on MSVC and the regparm attribute on GCC.
All of these were dropped on x64, on x64 (and ARM) you get standard calling conventions for just about everything with proper unwind tables for functions.
It doesn't really "cheat" on the registers unless it inlines a function entirely. LLVM has support for custom calling conventions and pragmas to specify them, this is used by GHC on Haskell and other things, but it's practically unheard of in "normal" C/C++ code.
If it were practical and if there were significant performance benefits (on modern heavyweight CPUs) to tailoring the calling convention on a per-function basis, I'd expect to see it done in optimising JIT engines like Java HotSpot. As far as I know, they don't bother.
I'm no expert but I suspect jcranmer's comment has it right that you end up doing cross-function register-allocation while foregoing the other benefits of just inlining. I also suspect the payoff would be minimal on modern heavyweight hardware. I can see it making more of a difference on a very minimal embedded processor, or if optimising for the smallest binary possible.
As a compiler engineer, it's not hard to find opportunities for the compiler to produce better code. What's difficult is turning those into generalizable patterns without introducing bugs or regressing performance elsewhere. You're right that as time goes on hand-written assembly won't change, but I see that as increasingly less of an issue now that we have LLMs to pour over and re-analyze what's going on as new releases come out.
> `feature(explicit_tail_calls)` is currently incomplete and may not work properly.
Works on my machine (tm). At least with toy examples. Including in full optimizationless debug mode, turning `call`s into `jmp`s ensuring `factorial(usize::MAX)` won't stack overflow probably maybe.
Compiler isn't really difficult in the way that many think it is, because it really isn't that difficult to write a basic C compiler for example, and writing a parser is pretty mechanical that there are a multitude of parser generators. The difficult part is that writing an optimized compiler is much more difficult than writing a correct compiler, and certain language features, like generics, closure, and async, propagates and touch every part of the language that the complexity grows exponentially, and it's really hard to balance performance, compile speed, and correctness against miscompiles.
You can kinda see that in the many "compile TypeScript to native via LLVM" projects that showed up a lot recently. From my testing, none of them could beat V8/Node JIT in the majority of cases, and most of them are generally 10x-40x slower.
But, we did have a great number of innovations in language design over the last decades that really closes the gap on how optimized a compiler can be over writing assembly directly: Rust's exhaustive match default null-less error handling and language level MIR, immutable data structures from functional languages to mainstream ones, TypeScript's compile time constraints as core part of the language, and Zig's `comptime` turning compile time metaprogramming to an integrated part of the language instead of C++ template metaprogramming.
Obviously, it's not really possible to beat hand optimized C/C++ or directly authored assembly, but I think a well-designed compiler/language can potentially beat idiomatic C/C++ in performance.
And this is speaking as someone who learned compiler design solely from having every one of his vibe-coded projects turn into either a compiler or a kernel for some reason.
Complexity does not grow exponentially and not even polynomially because you lower things. At the backend, generics, closures, async and many other features are completely non-existent (except for debuginfo. Debuginfo is complicated).
That does not mean writing an optimizing compiler is easy, of course. The frontend for many modern languages (especially Rust) if also far from trivial.
> In modern times, everyone knows that writing assembly is a fool's errand
ffmpeg is like 10% assembly. I think something similar is true of all video encoders. OpenSSL and libsodium write some of their core math routines in assembly (e.g,. NTT).
Ideally, a specific well optimized code for a problem domain could always outperform a generic optimized code for the same problem domain.
This is because the specific solution can make assumptions that generic cannot.
This usually holds true everywhere, not just for compilers.
This is not an excuse for avoiding generic solutions. But where performance matters absolutely and where the problem space is sufficiently constrained, specific solutions become the valid path.
Compilers are pretty good. Really good, even. But languages, even C, are abstractions, which necessarily constrain the level below. This threaded-code jump thing is just not possible to express in C. Even the best abstraction can often be beaten by something the abstraction can't express. Self-modifying code is one example. So is this threaded interpreter with its non-structured control flow.
But it takes longer. It's more difficult. That's why the abstraction exists and is still very useful despite its limitations.
Self-modifying code is a meme unless you're writing an obfuscator... (but even then you can hook the execution so it's a pretty low-tier anti-reversing effort ngl)
For perf reasons it's the equivalent of shooting your leg off to lose weight. You're flushing the instruction cache and breaking prefetch, leading to a huge stall. Then you do it again. And again. It hasn't been in vogue since the 80s...
Dynamic languages, like Common Lisp, where things can be redefined at run time likely require modification of running code to achieve high efficiency (the alternative is to just leave a general mechanism in place and accept the runtime overhead and loss of optimization opportunities from that.)
That depends how far in advance of running the code you're modifying it. You could see a JIT compiler as an extreme type of SMC. Or a Monero miner - its hashing algorithm relies on running randomly generated programs.
JIT and SMC are two different things because JIT is write-once-then-execute, SMC is write-many-then-execute-many. Even with reoptimisation like the JVM, you're not modifying the existing code but writing it onto a new page. That is the key difference, you're not modifying existing written-out instructions, you're creating new ones.
Consider a kernel that boots on several different machines with different capabilities. Patching the call instructions at init time, to point to either one version or another of a frequently used function, saves time.
There are literally thousands of compiler engineers who pore over the assembly a compiler generates, and then finds ways to make it better. I get paid to figure out where the compiler can do better, and often a step in that is to hand-code my own replacement. I then teach the compiler to do that.
But even beyond that, the compiler can't make certain assumptions that an assembly writer can. Such as whether a callee-saved register really does need to be saved in some particular routine. Or even pushing an extra parameter in unusual cases.
So it is entirely possible to beat the compiler, it's doable under certain circumstances, even today.
But you also have the danger of your loving hand-crafted assembly beating the compiler today. But next year the compiler is even smarter, the hardware may have changed in subtle ways, and the compiler will know and improve the code it generates.
Your hand-written code won't change unless you revisit it.
What do you think the most impactful savings from hand-written assembly are over optimized Rust today?
With hand-written assembly, you can (with effort and care) ensure that the code runs in constant time and doesn’t leak any information through side channels (such as which memory or cache addresses it accesses). That’s important for most encryption code. It’s difficult to ensure constant time execution in pure Rust (or C, or most high level languages in general) because you can’t tell if some future compiler optimization will break your attempts at constant-time code.
An optimizer is trying to balance between compile time, runtime speed, and code size, and most optimizations will win you one axis at the cost of one or both of the other axes. (The rare optimizations that win on all 3 are already all implemented in the compiler.) Compiler developer time is also a scarce resource; I know of so many more optimizations I could implement, but without demonstrable code that would actually benefit, it's not a good use of my time to implement them. Compilers tune this balance by making lots of heuristic decisions, and these heuristics are tuned by large benchmarks, which often times involve a lot of flat code profiles (i.e., no code is worth spending a lot of time really nailing down the best code layout).
One of the advantages of hand-written assembly is that you get to opt out of the compiler heuristics and commit to being able to spend the time to optimize the one bit of code that you know is really important for runtime as perfectly as you want, instead of relying on the compiler to get it close enough to perfect before it exhausts its budget of caring about optimizing it.
You can usually hint the compiler, but the best idea is to provide it with more data. PGO with real-world data and LTO for whole-program analysis enable it to do much better decisions.
1. Better memory locality by knowing what you load and when, exactly.
2. The ability to "cheat" on calling conventions.
3. The ability for techniques like threaded code, and in general, better cache-awareness.
4. Less mov's.
5. Guaranteeing no spilling in important loops.
6. Compilers don't do well with flags registers and you can't read/write them in high level languages. You're hoping your `if (result < a) { carry = 1; }` becomes a direct flag test. Especially important in bignum, you can't really utilise adcx/adox directly from high-level code.
7. Hot/cold layout without PGO. Yes PGO is good but sometimes you know better and PGO isn't very suitable for "configurable" code.
8. Computed goto. See https://github.com/python/cpython/issues/128563 , who doesn't like 10% free performance?
9. Exploiting uninitialised memory for classic party tricks like not initialising a buffer fully (let's say you have a library function with a return buffer. You don't want dynamic allocations for some reason. You can simulate this with a pointer return into a let's say a static 4KB buffer and a count return, you only initialise it until the count. Caller has the responsibility not to overread.)
Hand-optimized assembly tends to be ephemeral. My codebase constantly changes, which means hand-written assembly had to be redone.
It's rarely worth the effort.
There a lot of juice in improving the data structures that a compiler cannot do.
usually with Rust impactful savings come from writing the Rust differently, not from doing hand-written assembly instead.
See stuff like https://davidlattimore.github.io/posts/2025/09/02/rustforge-...
This doesn't mean Rust is near perfect, it's just that your first move should be "how do I make the Rust better" and not "I need to drop into asm."
"We’d like to initialise our Vec in parallel, otherwise we’d have to wait for the main thread to fill the entire Vec with a placeholder value only to then have our threads overwrite those placeholder values."
Talk about overengineering :P Multithreaded vector initialisation instead of just...skipping it?
That sentence is specifically talking about not initializing with a placeholder value, and instead letting it write over uninitialized values.
Entirely situation and application specific. Just like with every other language.
So there's no patterns at all?
I'm so far downstream from the source language that I know of none specific to Rust. And note someone else's specific answer to your question above that has nothing Rust specific either.
As one example, LLVM still routinely does dumb stuff like this: https://github.com/llvm/llvm-project/issues/53348
Regarding ABI, calee-saved registers also often result in useless data shuffling and prevent the compiler from using them for argument/result passing.
With whole program analysis (LTO, usually, though not only), LLVM can create custom ABIs.
Any “Multimedia Extensions” or streaming or whatever beyond what’s available past a 486.
Rust has autovectorization, but a developer knows their algorithms best.
Also AES-NI vs software is no contest.
When people talk about out coding ‘to the metal’ you have to consider what ‘the metal’ provides
> Such as whether a callee-saved register really does need to be saved in some particular routine.
Ironically this is your preconceived abstraction of how a compiler has to operate. An ideal compiler could allocate registers differently for each called function: F1()->F2()->F3(), F1 uses r0-5, F2 uses r6-10, F3 uses r11-15, no register saving required in the whole chain. There's no need for a fixed ABI. Such a compiler would look very different from today's ones.
Modifying the ABI of a function requires being able to track down all of the call-sites of the function, which is less trivial than you might assume. ABI concerns also tend to baked in relatively early in the optimization pipeline because you just simply can't get the ABI wrong, and I can think of several instances where the ABI decision causes missed optimizations.
There is also the other issue that a good algorithm for optimizing a problem like register allocation tends to be super-linear (e.g., quadratic), and if you shift the model from "allocate on a per-function basis" to "allocate all functions", the N in the O(N²) goes from "size of function" to "size of program," which is now suddenly a lot more compiler time spent for very modest gains. If register spilling across a function call is a noticeable component of runtime, then you're probably better off inlining that function in the first place!
One of the hardest parts of optimizations is not being penny wise and pound foolish. This sort of thing is a great example of that.
Additionally, a hard part is that all of this can change over time with new hardware! Some patterns that were crucial before everything gained branch predictors are irrelevant now, etc.
> penny wise and pound foolish
I love this saying. The general problem, optimizing the wrong metric, shows up all over the place.
Yeah, that's what MSVC and GCC did on x86, called "custom calling conventions" on MSVC and the regparm attribute on GCC.
All of these were dropped on x64, on x64 (and ARM) you get standard calling conventions for just about everything with proper unwind tables for functions.
It doesn't really "cheat" on the registers unless it inlines a function entirely. LLVM has support for custom calling conventions and pragmas to specify them, this is used by GHC on Haskell and other things, but it's practically unheard of in "normal" C/C++ code.
If it were practical and if there were significant performance benefits (on modern heavyweight CPUs) to tailoring the calling convention on a per-function basis, I'd expect to see it done in optimising JIT engines like Java HotSpot. As far as I know, they don't bother.
I'm no expert but I suspect jcranmer's comment has it right that you end up doing cross-function register-allocation while foregoing the other benefits of just inlining. I also suspect the payoff would be minimal on modern heavyweight hardware. I can see it making more of a difference on a very minimal embedded processor, or if optimising for the smallest binary possible.
This is already done in the form of deciding whether or not to inline a function.
As a compiler engineer, it's not hard to find opportunities for the compiler to produce better code. What's difficult is turning those into generalizable patterns without introducing bugs or regressing performance elsewhere. You're right that as time goes on hand-written assembly won't change, but I see that as increasingly less of an issue now that we have LLMs to pour over and re-analyze what's going on as new releases come out.
Now I'm wondering if Claude can beat the compiler if I ever need to optimize something a bit more.
> At this time there is no portable way to produce computed gotos or tail call optimization in compiled machine code from Rust.
Nightly rust now has the `become` keyword:
https://doc.rust-lang.org/std/keyword.become.html
> `feature(explicit_tail_calls)` is currently incomplete and may not work properly.
Works on my machine (tm). At least with toy examples. Including in full optimizationless debug mode, turning `call`s into `jmp`s ensuring `factorial(usize::MAX)` won't stack overflow probably maybe.
https://rust.godbolt.org/z/7xf836E8K
> checked_sub? wrapping_mul? MaulingMonkey, what's wrong with you?
Eliminating debug-mode panic boilerplate.
Compiler isn't really difficult in the way that many think it is, because it really isn't that difficult to write a basic C compiler for example, and writing a parser is pretty mechanical that there are a multitude of parser generators. The difficult part is that writing an optimized compiler is much more difficult than writing a correct compiler, and certain language features, like generics, closure, and async, propagates and touch every part of the language that the complexity grows exponentially, and it's really hard to balance performance, compile speed, and correctness against miscompiles.
You can kinda see that in the many "compile TypeScript to native via LLVM" projects that showed up a lot recently. From my testing, none of them could beat V8/Node JIT in the majority of cases, and most of them are generally 10x-40x slower.
But, we did have a great number of innovations in language design over the last decades that really closes the gap on how optimized a compiler can be over writing assembly directly: Rust's exhaustive match default null-less error handling and language level MIR, immutable data structures from functional languages to mainstream ones, TypeScript's compile time constraints as core part of the language, and Zig's `comptime` turning compile time metaprogramming to an integrated part of the language instead of C++ template metaprogramming.
Obviously, it's not really possible to beat hand optimized C/C++ or directly authored assembly, but I think a well-designed compiler/language can potentially beat idiomatic C/C++ in performance.
And this is speaking as someone who learned compiler design solely from having every one of his vibe-coded projects turn into either a compiler or a kernel for some reason.
Complexity does not grow exponentially and not even polynomially because you lower things. At the backend, generics, closures, async and many other features are completely non-existent (except for debuginfo. Debuginfo is complicated).
That does not mean writing an optimizing compiler is easy, of course. The frontend for many modern languages (especially Rust) if also far from trivial.
Previous discussion from 2024 (linked in the "Post-publication notes" section):
https://news.ycombinator.com/item?id=40948353
Thanks! Macroexpanded:
Beating the Compiler - https://news.ycombinator.com/item?id=40948353 - July 2024 (71 comments)
> In modern times, everyone knows that writing assembly is a fool's errand
ffmpeg is like 10% assembly. I think something similar is true of all video encoders. OpenSSL and libsodium write some of their core math routines in assembly (e.g,. NTT).
So maybe this myth should die?
I was playing a video using an old codec and VLC printed a warning “No hardware support on your Intel GPU”
But the software fallback was >5% CPU on a mobile Ice Lake in low power mode. ffmpeg is Good Stuff
For a second I thought this was a post about Marathon. I’m tired.
There are dozens of us who both play Marathon and care about the Rust compiler. Dozens!
Nice.
Ideally, a specific well optimized code for a problem domain could always outperform a generic optimized code for the same problem domain.
This is because the specific solution can make assumptions that generic cannot.
This usually holds true everywhere, not just for compilers.
This is not an excuse for avoiding generic solutions. But where performance matters absolutely and where the problem space is sufficiently constrained, specific solutions become the valid path.
Compilers are pretty good. Really good, even. But languages, even C, are abstractions, which necessarily constrain the level below. This threaded-code jump thing is just not possible to express in C. Even the best abstraction can often be beaten by something the abstraction can't express. Self-modifying code is one example. So is this threaded interpreter with its non-structured control flow.
But it takes longer. It's more difficult. That's why the abstraction exists and is still very useful despite its limitations.
Self-modifying code is a meme unless you're writing an obfuscator... (but even then you can hook the execution so it's a pretty low-tier anti-reversing effort ngl)
For perf reasons it's the equivalent of shooting your leg off to lose weight. You're flushing the instruction cache and breaking prefetch, leading to a huge stall. Then you do it again. And again. It hasn't been in vogue since the 80s...
JIT compilers are essentially producing self-modifying programs (or, indeed, dynamically loaded libraries).
Dynamic languages, like Common Lisp, where things can be redefined at run time likely require modification of running code to achieve high efficiency (the alternative is to just leave a general mechanism in place and accept the runtime overhead and loss of optimization opportunities from that.)
That depends how far in advance of running the code you're modifying it. You could see a JIT compiler as an extreme type of SMC. Or a Monero miner - its hashing algorithm relies on running randomly generated programs.
JIT and SMC are two different things because JIT is write-once-then-execute, SMC is write-many-then-execute-many. Even with reoptimisation like the JVM, you're not modifying the existing code but writing it onto a new page. That is the key difference, you're not modifying existing written-out instructions, you're creating new ones.
Consider a kernel that boots on several different machines with different capabilities. Patching the call instructions at init time, to point to either one version or another of a frequently used function, saves time.