The linux distribution I co-maintain uses mold as our bootstrap linker to bootstrap rust itself, and it saved us -hours- on long version-by-version build chains. Mold being in c meant we could build it very early and use it as the default linker distro wide and enjoy build speedups everywhere.
Now sadly we will have to fork and maintain the c version as mold2 forever.
Rust is not actually the right tool for all problems.
I have been thinking about Rust bootstrapping recently. Couldn't a Rust compiler without borrow checker be put together relatively easily?
Assuming the source code contains no issues (which can be checked later once the Rust compiler is built), one could leave that piece behind, and take the shortest path from .rs to executed code (C transpilation, or even an interpreter).
That's your choice and your responsibility. You don't get to bind upstream, at least not without paying money. You can't say Rust is a bad solution, just because you developed a bootstrap chain that happens to find useful the fact that mold is written in C.
Edit: sorry, didn't mean to throw any shade at stage0 or their right to complain, just saying they don't really get to factor into mold's decisions.
I do understand what you are trying to say, but I think that this fails to be kind to Irvick and the other maintainers of the stage0 project who are actually being quite underfunded (someone like OAI should fund them in the name of security!)
So asking them to pay money is well.. not quite the solution.
What can happen as it often happens, is this, Mold creator created the project, Stage0 found it useful, Mold ports itself to rust, Stage0 now founds it not useful, Stage0 can comment on the new update and be slightly disappointed and comment how they would've preferred to not rust in this particular case for them.
Yes this doesn't prevent mold from changing to rust or anything as it was shown but I guess we can allow the ability for Irvick to drop a comment I guess without saying that's your responsibility while they are just being maintainers of the open source and an underfunded/mostly volunteer one of work at that.
Also, @Irvick, stage0 is extremely cool, I hope that a lot more companies sponsor the work that stage0 contributors are doing. I think that it can help the supply-chain issues that the industry is facing at, and is honestly just a really cool idea and I love reading your comments and thank you!
If the need is well justified, maybe there is great chance to ask adding #[no_mangle] and extern C support? Since release is very fresh. If that is causing the problem. Or is some dependency the issue?
I'm a bit confused why a decision to decrease the maintenance burden of Mold by switching languages makes Rust the wrong tool here? Fearless concurrency sounds like a huge benefit for what they're doing, given the resources they have.
One of the explicit goals for Mold 3.0 is to promote its use as the default linker in Linux distributions, and bootstrapping complexity is certainly worthy of consideration in that realm.
> We will then conduct extensive compatibility testing and work closely with Linux distribution developers to make it practical for them to adopt mold as /usr/bin/ld. Making this happen is one of our highest priorities for mold 3.x.
It's simply an ordering problem. When building the entire world from scratch, usually the C and C++ toolchains are built near the very first, and Rust toolchains built somewhat later. Anything written in Rust must come after the Rust toolchain is built. You need a linker as part of your C++ toolchain, so it must be written in a language ready to go at that point. If it is written in C, you are done. If it is written in Rust you have to wait. So a Rust-based linker can no longer be used from that very, very early C bootstrap.
It's not a big deal for normal users, where you have Rust ready to go. Kind of a bummer in this case, but this is a specialized one.
You could. But now you are trusting an entirely new toolchain that can build the Rust-based linker, even if it is statically linked. This toolchain is much larger than a comparatively small and much more easily understood C bootstrap toolchain.
This effectively triples or quadruples (maybe even more) the amount of code you need to trust for the cold bootstrap.
>Rust is not actually the right tool for all problems.
"my problems (that are not mold's) are not solved by mold. How dare mold make those decisions?"
Maybe you should rewrite more of your linux distribution in rust so it's available earlier in the build process and get back to it being the default linker.
No, they were never one of mold's goals. You don't get to assign solutions to people _and then blame them_. Hyrum's Law being a load bearing part of your infrastructure isn't Mold's problem.
I can't really agree with that. If you ship your linker with a disclaimer saying 'don't use this for early stage stuff, one day we might decide to rewrite the whole codebase in a language that won't be available early enough' then you'd have a point. But if you didn't that is precisely the kind of use case where mold shines, and so inevitably it will be adopted. Your downstream should be precious. At a minimum you should communicate such intentions, do so timely and hear what others have to say, even if afterwards you decide to push on.
Pretty much every piece of open source software ships with a legal document saying that the software is provided as-is without any warranty or guarantee of functionality (i.e. https://github.com/rui314/mold/blob/main/LICENSE ). Unless I had a support agreement with the maintainers that superseded that document, I would not expect any special treatment from upstream. Personally, whenever I add a new dependency, I do it with the knowledge that I may have to either replace it or take on maintenance myself in the future.
I actually made a Linux distribution that relies on this being your latest comment, since you didn't say this wouldn't always be your latest comment. If you post any newer ones, it will break.
Yeah, for stuff we wanna rely on, this should ideally be the default. I'd go one step further and say no breaking changes past 1.0.0 at all. Instead people should favor creating entirely new projects (forks or not) and jump over to those, leaving the old one behind, if they want to do massive changes to something.
Of course, no one would be forced to do this, but it feels like if more did this, long-term supporting stuff that depends on those things would be a lot easier, if things could just be instead of changing under our feet all the time. Thank god for Nix and NixOS, even with their warts.
If not, downstream may maintain mold2 themselves forever because their _extremely specific_ use case is neither a promise nor a valuable thing for mold to maintain. Debian understood this a while ago and isn't whining when they have to maintain their own fork. Maintainers maintain.
If enough downstreamers are unhappy about choices, they are also welcome to fork, until the base project is abandoned. Or they realize their usecases are extremely narrow.
(In addition, it's an extremely hypocritical and purist demand, because I am pretty certain that their distribution has, at some point, an arbitrary executable to make a compiler from. So they're probably okay with blobs, just not that one in particular.)
Was the conversion assisted? I get the sense that Mold's author is an extremely competent guy and it would be quite feasible for someone like him to do it on his own.
I'm curious about motivation - bounds checks for corrupted inputs seems like it would be one, but it also seems that fixing corrupted input handling in a C/C++ code base would not be too hard, and probably less of an effort? So why did you choose the rewrite?
And second, did you use any AI tools for the rewrite?
> fixing corrupted input handling in a C/C++ code base would not be too hard
The best programmers on the planet have tried and failed with this task for 50 years now, so I don't think this is true.
The main disadvantage of Rust right now is not supporting some more obscure platforms, but because mold wouldn't support them anyway I don't see that as a problem.
> The main disadvantage of Rust right now is not supporting some more obscure platforms
Last time I checked, I got impressed by the wide platform support, once you go down the tier list (https://doc.rust-lang.org/nightly/rustc/platform-support.htm...). What "obscure platform" specifically are you thinking about, that is currently missing from those lists?
Just to be clear, I think this is a very small disadvantage. GCC and therefore C/C++ supports some old stuff like SuperH, Intel Itanium, PA-RISC and a bunch of microcontroller archs that LLVM does not.
However, there is now a Rust codegen plugin for GCC, so even this disadvantage is now basically moot.
I think the claim is that GCC supports a few obscure platforms that LLVM doesn't. Don't ask me which, that is just the claim that I heard multiple times. So it's not Rust that limits platforms, but LLVM. And they say the GCC backend efforts of Rust are meant to deal with that.
There are many. Two big ones are Alpha and PA-RISC. NetBSD and Linux continue to support both. Linux distro choices are pretty much limited to Gentoo though.
1. You've got the causality backwards. LLVM backends don't come for free, you have to write code to enable support for them.
2. LLVM does not support as many backends as GCC does, so even if you did get 100% of the LLVM supported backends up and running, you'd still be missing some.
My very naive understanding is that a part of what makes mold fast is concurrency, which I'd expect to be a lot more error-prone in C/C++. Not having to worry about data races might give more confidence with trying out more complex techniques for how to split up work in a way that ends up making things faster
It's possible to write safe code in C or C++ but it's extremely difficult to read existing code and prove it's safe, without using as much effort as it takes to write it in the first place. This includes the code you wrote last month whose surrounding code has changed. And you have to be right every time while the attacker only needs you to be wrong once. The problem is not writing the code, it is continually verifying it.
I suspect fearless concurrency is another motivating factor. Better perf can be achieved by squeezing more parallelism, but without borrow checking it's difficult to do fine-grained parallelism correctly.
Bound checks, which are part of what you get with Rust, also exist in ReleaseSafe Zig, that is true. But if that is what you wanted to say, your statement is very confusing and also does not answer the GP. If you meant to say ReleaseSafe and Rust are equivalent, then this is just false. ReleaseSafe still has many UBs (most importantly use after free and double free), not to mention that Rust solves many problems at compile time and Zig only at runtime.
Zig makes you choose that at build time. You can use safe builds for development to catch bugs, and fast builds for release. You can even use safe builds in of the release, and optimize hot paths with fast builds.
In practice, projects written in Zig very much can choose both.
What warrants an off the cuff comment like this?
I can take any piece of software and just say a vague statement like that. All of software... What does it even bring to the conversation? Just vague "they"s again we keep on propping up? Who is unethical? The mold team? The AI? The rust foundation? The ASCII character set? The cloud system it has been compiled on? Oh no obviously the backdoors in there?! Oh let us guess!!
2. Most of those are obviously LLM-induced, unless there's a new breed of human trained on just the failures of LLMs and not the successes or other relevant information
3. I therefore assume that the rest of the project is an unverified LLM psychosis without meaningful human review
That's not the worst thing in the world for all software maybe. I recently found out my apartment complex in their latest AI rewrite generates SMS OTP based purely on the timestamp and ignores passwords, so I can log in as anyone else by knowing their email and having a valid email to grab the current OTP. That's a major security failing, but how bad is it really? I can grab the last-4 of their credit card numbers, pay their rent, see how much other tenants are being screwed, and so on, and only if I know their email addresses (solvable with a tiny bit of social engineering, but let's assume that's also moderately hard). How bad is that really? On the one hand, it's terrifying, since I presume they have the same level of attention to detail with respect to payment methods and PII despite my having opted out of having them stored, but (a) that's all already been exposed via dozens of breaches and is being handled behind the scenes by my bank anyway, and (b) if we examine the immediate blast radius of the known bugs then there's approximately fuck-all an attacker can do with that information.
For a multi-threaded linker? Come the fuck on. I don't care if it's written in Rust. If you ignore all of the memory ordering intrinsics and `unsafe` then maybe it's more okay, but not having easily available UAF and other memory bugs is very different from actually implementing the correct behaviour, and for a linker where you're explicitly joining together multiple independent binary blobs and choosing what and how to execute on a machine instruction level, Rust's guarantees 100% don't save you from broken, unvalidated "business logic."
(a) That's mostly irrelevant to my point. This project has major issues with an LLM signature in a safety-critical context. If we assume that the bun rust rewrite went off without a hitch and that doing so is repeatable then it's worth trying to understand what the difference was.
(b) I'm not sure those assumptions hold. Bug reports spiked 2-3x when the rust version was introduced. HMR broke. There were utf-16 parsing panics. The result was slower out of the box and needed a lot of hand-tuning. There were FFI regressions. The rust rewrite had its own concurrency bugs despite rust's "fearless concurrency." In a JS runtime where the rest of the ecosystem is already usually broken, especially since the bugs seemingly weren't safety-related, maybe that's fine, but it wasn't exactly a "no major issues" rewrite.
The days when AI always means slop are gone. It still often does, but the latest models can be used to get very high quality code in many situations. I haven't checked the code here but I would be surprised if it's bad.
"Recent advances in AI-assisted coding have made large-scale rewrites considerably more practical, but they do not eliminate those risks, and we do not take them lightly."
Last time I checked it didn't make any difference or whatsoever when compared to lld so I would not go that far by saying "mold is a reference in the world of linkers" - made a test just few days ago on 10M LoC C++. So, rewriting the codebase in another language just because doesn't seem like an investment worth doing when there's much bigger fish to fry.
If it was transpiled there would be what GNU calls a "preferred form" of the linker source in which it wasn't written in Rust.
This is true for - as an example - the WUFFS GIF decoder. You can get C which decodes GIFs and was transpiled from WUFFS, but that's awful code and nobody wants to modify that code, whereas the WUFFS source code for the decoder is fine.
When we look at rui314's changes to mold today after 3.0 release, they just modify the mold source code in Rust, as you'd expect if this was in fact now written in Rust.
Now sadly we will have to fork and maintain the c version as mold2 forever.
Rust is not actually the right tool for all problems.
Would that meant that building rust is slower, but everything else is the same, and you don't have to maintain a fork of another complex project?
Assuming the source code contains no issues (which can be checked later once the Rust compiler is built), one could leave that piece behind, and take the shortest path from .rs to executed code (C transpilation, or even an interpreter).
Edit: sorry, didn't mean to throw any shade at stage0 or their right to complain, just saying they don't really get to factor into mold's decisions.
So asking them to pay money is well.. not quite the solution.
What can happen as it often happens, is this, Mold creator created the project, Stage0 found it useful, Mold ports itself to rust, Stage0 now founds it not useful, Stage0 can comment on the new update and be slightly disappointed and comment how they would've preferred to not rust in this particular case for them.
Yes this doesn't prevent mold from changing to rust or anything as it was shown but I guess we can allow the ability for Irvick to drop a comment I guess without saying that's your responsibility while they are just being maintainers of the open source and an underfunded/mostly volunteer one of work at that.
Also, @Irvick, stage0 is extremely cool, I hope that a lot more companies sponsor the work that stage0 contributors are doing. I think that it can help the supply-chain issues that the industry is facing at, and is honestly just a really cool idea and I love reading your comments and thank you!
Then you can just use the latest built version to build mold and the next version of rust.
I know this is common but it seems like either an aesthetic decision, or glibc cruft.
> We will then conduct extensive compatibility testing and work closely with Linux distribution developers to make it practical for them to adopt mold as /usr/bin/ld. Making this happen is one of our highest priorities for mold 3.x.
It's not a big deal for normal users, where you have Rust ready to go. Kind of a bummer in this case, but this is a specialized one.
This effectively triples or quadruples (maybe even more) the amount of code you need to trust for the cold bootstrap.
"my problems (that are not mold's) are not solved by mold. How dare mold make those decisions?"
Maybe you should rewrite more of your linux distribution in rust so it's available earlier in the build process and get back to it being the default linker.
Yeah, for stuff we wanna rely on, this should ideally be the default. I'd go one step further and say no breaking changes past 1.0.0 at all. Instead people should favor creating entirely new projects (forks or not) and jump over to those, leaving the old one behind, if they want to do massive changes to something.
Of course, no one would be forced to do this, but it feels like if more did this, long-term supporting stuff that depends on those things would be a lot easier, if things could just be instead of changing under our feet all the time. Thank god for Nix and NixOS, even with their warts.
Is my downstream paying me for the maintenance ?
If not, downstream may maintain mold2 themselves forever because their _extremely specific_ use case is neither a promise nor a valuable thing for mold to maintain. Debian understood this a while ago and isn't whining when they have to maintain their own fork. Maintainers maintain.
If enough downstreamers are unhappy about choices, they are also welcome to fork, until the base project is abandoned. Or they realize their usecases are extremely narrow.
(In addition, it's an extremely hypocritical and purist demand, because I am pretty certain that their distribution has, at some point, an arbitrary executable to make a compiler from. So they're probably okay with blobs, just not that one in particular.)
Edit: this seems to have been cooking for a while when the first commit dropped: https://github.com/rui314/mold/commit/f41bfcd5c72ca30cce6498...
Yes. Comment by mold's author:
https://www.reddit.com/r/rust/comments/1w45j6n/comment/p7ac5...
(Wild is another fast linker, that only supports Linux)
And second, did you use any AI tools for the rewrite?
The best programmers on the planet have tried and failed with this task for 50 years now, so I don't think this is true.
The main disadvantage of Rust right now is not supporting some more obscure platforms, but because mold wouldn't support them anyway I don't see that as a problem.
Last time I checked, I got impressed by the wide platform support, once you go down the tier list (https://doc.rust-lang.org/nightly/rustc/platform-support.htm...). What "obscure platform" specifically are you thinking about, that is currently missing from those lists?
However, there is now a Rust codegen plugin for GCC, so even this disadvantage is now basically moot.
Rewrite all the things!
There are many. Two big ones are Alpha and PA-RISC. NetBSD and Linux continue to support both. Linux distro choices are pretty much limited to Gentoo though.
2. LLVM does not support as many backends as GCC does, so even if you did get 100% of the LLVM supported backends up and running, you'd still be missing some.
IMO, given the recent commits: almost certainly.
It's like "in a Zodiac boat / aircraft carrier navy". This customary putting C and C++ into the same bucket is as amusing as it is unproductive.
In practice, projects written in Zig very much can choose both.
Zig isn’t even in the same league as Rust regarding these things. Zig may still be around and active 10 years from now, Rust is guaranteed to be.
LLM are excellent translation tools. Nothing is learned or stolen from them out of this exercise if this is a copyright issue you are getting at.
Then Rust? Why picking on it, the language is morally corrupt? In what way? If it were a rewrite in Object Pascal it would be better?
I will not explain why.
1. There are major obvious flaws
2. Most of those are obviously LLM-induced, unless there's a new breed of human trained on just the failures of LLMs and not the successes or other relevant information
3. I therefore assume that the rest of the project is an unverified LLM psychosis without meaningful human review
That's not the worst thing in the world for all software maybe. I recently found out my apartment complex in their latest AI rewrite generates SMS OTP based purely on the timestamp and ignores passwords, so I can log in as anyone else by knowing their email and having a valid email to grab the current OTP. That's a major security failing, but how bad is it really? I can grab the last-4 of their credit card numbers, pay their rent, see how much other tenants are being screwed, and so on, and only if I know their email addresses (solvable with a tiny bit of social engineering, but let's assume that's also moderately hard). How bad is that really? On the one hand, it's terrifying, since I presume they have the same level of attention to detail with respect to payment methods and PII despite my having opted out of having them stored, but (a) that's all already been exposed via dozens of breaches and is being handled behind the scenes by my bank anyway, and (b) if we examine the immediate blast radius of the known bugs then there's approximately fuck-all an attacker can do with that information.
For a multi-threaded linker? Come the fuck on. I don't care if it's written in Rust. If you ignore all of the memory ordering intrinsics and `unsafe` then maybe it's more okay, but not having easily available UAF and other memory bugs is very different from actually implementing the correct behaviour, and for a linker where you're explicitly joining together multiple independent binary blobs and choosing what and how to execute on a machine instruction level, Rust's guarantees 100% don't save you from broken, unvalidated "business logic."
(b) I'm not sure those assumptions hold. Bug reports spiked 2-3x when the rust version was introduced. HMR broke. There were utf-16 parsing panics. The result was slower out of the box and needed a lot of hand-tuning. There were FFI regressions. The rust rewrite had its own concurrency bugs despite rust's "fearless concurrency." In a JS runtime where the rest of the ecosystem is already usually broken, especially since the bugs seemingly weren't safety-related, maybe that's fine, but it wasn't exactly a "no major issues" rewrite.
"Recent advances in AI-assisted coding have made large-scale rewrites considerably more practical, but they do not eliminate those risks, and we do not take them lightly."
This is true for - as an example - the WUFFS GIF decoder. You can get C which decodes GIFs and was transpiled from WUFFS, but that's awful code and nobody wants to modify that code, whereas the WUFFS source code for the decoder is fine.
When we look at rui314's changes to mold today after 3.0 release, they just modify the mold source code in Rust, as you'd expect if this was in fact now written in Rust.