Smaller argparsh binaries
Making argparsh smaller and 3x faster by optimizing binary size!
After my last post, I realized that a bottleneck might be the binary size itself.
Each invocation of argparsh must load the binary into memory.
While each invocation is likely to load the binary from cache, it’s still overhead to map in the code pages and execute them.
By default, the release cargo build optimizes for performance, not code size.
Performance optimizations likely don’t matter much since argparsh doesn’t have a lot of loops or branching anyway.
This is easy to change by setting opt-level in Cargo.toml to either “s” or “z”.
“z” is the most aggressive option, disabling some optimizations that “s” includes.
The nostd build from the last post, was already using “s”, so I tried changing it to “z”, results below.
In the std case, the binary went from 1.6M to 885K with “s” and 805K with “z”.
However, In the nostd case, the binary went from 70K to 69K, suggesting that “z” might not be optimal here.
# previous-post: Baseline
Benchmark 1: env PARSER=argparsh bash bench.sh a -i 100 foo qux
Time (mean ± σ): 64.4 ms ± 12.4 ms [User: 37.8 ms, System: 27.0 ms]
Range (min … max): 48.3 ms … 105.7 ms 27 runs
# previous-post: best build with std
Benchmark 1: env PARSER=./target/x86_64-unknown-linux-musl/release/argparsh bash bench.sh a -i 100 foo qux
Time (mean ± σ): 31.7 ms ± 1.4 ms [User: 8.1 ms, System: 25.0 ms]
Range (min … max): 29.0 ms … 35.6 ms 83 runs
# previous-post: best build with nostd
Benchmark 1: env PARSER=./nostd-demo/target/x86_64-unknown-linux-musl/release/argparsh PARSE_PARSER=./target/x86_64-unknown-linux-musl/release/argparsh bash bench.sh a -i 100 foo qux
Time (mean ± σ): 18.6 ms ± 4.9 ms [User: 5.8 ms, System: 14.0 ms]
Range (min … max): 2.1 ms … 23.3 ms 118 runs
# small binary (w/ std)
Benchmark 1: env PARSER=./target/x86_64-unknown-linux-musl/release/argparsh bash bench.sh a -i 100 foo qux
Time (mean ± σ): 7.0 ms ± 2.6 ms [User: 3.9 ms, System: 3.4 ms]
Range (min … max): 6.0 ms … 33.9 ms 353 runs
# small binary + nostd (with std binary for final "parse" invocation)
Benchmark 1: env PARSER=./nostd-demo/target/x86_64-unknown-linux-musl/release/argparsh PARSE_PARSER=./target/x86_64-unknown-linux-musl/release/argparsh bash bench.sh a -i 100 foo qux
Time (mean ± σ): 4.1 ms ± 0.3 ms [User: 2.6 ms, System: 1.8 ms]
Range (min … max): 3.6 ms … 7.0 ms 528 runs
Just changing the optimization to “z” yields a 4.5x speed up in both the std and nostd case! The speedup in the nostd case is coming from the final invocation to the std binary instead of from the intermediate steps. Starting from the baseline, this is almost a 16x speedup!
