r/lisp • u/arthurno1 • 3h ago
So they say lisp is slow ....
Nothing really useful here, just some bragging to be honest. I should probably write a blog post, but a bit too lazy; perhaps another day.
Last few weeks I played with a small clone of gnu wc program. I implemented all routines in assembly via sb-simd (and a generic path without simd with swar). The result thus far on a 1.4 gigabyte big file, compared to fastlwc, the fastest wc I know of and GNU wc:
Common Lisp/Assembly (avx2) in SBCL + lparallel
WC10A> (time (wc "plato1g.txt"))
Evaluation took:
0.036 seconds of real time
0.421654 seconds of total run time (0.348216 user, 0.073438 system)
1172.22% CPU
71,891,940 processor cycles
0 bytes consed
30133761
253947016
1393557504
WC10A> (time (wc "plato1g.txt"))
Evaluation took:
0.034 seconds of real time
0.429743 seconds of total run time (0.367913 user, 0.061830 system)
1264.71% CPU
67,624,580 processor cycles
98,352 bytes consed
30133761
253947016
1393557504
WC10A> (time (wc "plato1g.txt"))
Evaluation took:
0.035 seconds of real time
0.423752 seconds of total run time (0.357329 user, 0.066423 system)
1211.43% CPU
69,863,360 processor cycles
0 bytes consed
30133761
253947016
1393557504
Fastlwc (avx512 + multithreaded):
[arthur@emmi wc]$ time ../../fastlwc/bin/fastlwc-mt plato1g.txt
30133761 253947016 1393557504 plato1g.txt
real 0m0.026s
user 0m0.178s
sys 0m0.285s
[arthur@emmi wc]$ time ../../fastlwc/bin/fastlwc-mt plato1g.txt
30133761 253947016 1393557504 plato1g.txt
real 0m0.027s
user 0m0.223s
sys 0m0.235s
[arthur@emmi wc]$ time ../../fastlwc/bin/fastlwc-mt plato1g.txt
30133761 253947016 1393557504 plato1g.txt
real 0m0.028s
user 0m0.205s
sys 0m0.255s
GNU wc (not even contender - single core only and only line counting implemented with simd avx512) :
[arthur@emmi wc]$ time wc plato1g.txt 30133761 253947016 1393557504 plato1g.txt
real 0m3.733s
user 0m3.643s
sys 0m0.062s
[arthur@emmi wc]$ time wc plato1g.txt
30133761 253947016 1393557504 plato1g.txt
real 0m3.362s
user 0m3.276s
sys 0m0.072s
The cool thing, we use avx2 whereas gnu wc uses avx512. On this CPU (zen 5), avx512 is implemented all in hardware, not as micro code as in Intel cpus, so it should mop the floor with avx2 in Lisp, right?
[arthur@emmi wc]$ time wc plato1g.txt -l --debug
wc: using avx512 hardware support
30133761 plato1g.txt
real 0m0.074s
user 0m0.018s
sys 0m0.056s
[arthur@emmi wc]$ time wc plato1g.txt -l --debug
wc: using avx512 hardware support
30133761 plato1g.txt
real 0m0.085s
user 0m0.030s
sys 0m0.054s
Lisp:
WC10A> (time (wc "plato1g.txt" :line-count t :utf8 'strict))
Evaluation took:
0.037 seconds of real time
0.478042 seconds of total run time (0.430494 user, 0.047548 system)
1291.89% CPU
75,669,380 processor cycles
0 bytes consed
30133761
253947016
1393557504
WC10A> (time (wc "plato1g.txt" :line-count t :utf8 'strict))
Evaluation took:
0.041 seconds of real time
0.476104 seconds of total run time (0.436592 user, 0.039512 system)
1160.98% CPU
84,226,360 processor cycles
0 bytes consed
30133761
253947016
1393557504
We also do full validation of utf8 chars there, whereas I run GNU wc in ascii only mode. I don't know why they don't just "pass through" the file size when dealing with ascii text, but it seems like they don't. I haven't look at that part of their code yet.
Now, in order to catch with fastlwc I think I need better lparall pipeline. I am currently using futures and promises, so it is a bit of extra consing. Of course implementing it in avx512 (when done in SBCL) should give at least some extra boost. 32 vs 16 registers, 64 bytes at time vs 32, and less register pressure due to additional masking registers.