pperl 0.6.21 - addendum

This blog entry has no release attached to it. 0.6.21 is an internal version, not for release. In the four days since 0.6.20 a few things happened, but this post is not about new things. It is about old things the 0.6.20 post forgot, about how a project moving at this pace should announce anything at all, and about a question that deserves an answer in public: who pperl is for, at a time when you can ask an AI to port your Perl script to the language of your choice and be done.

What the last post left out

"The Player of Games" covered the Benchmarks Game scoreboard, Memoize, the GPU binding and four bullet points. The two months it reported on contained more than that. The following landed between 0.6.8 and 0.6.20, is in the binary you can download, and was not mentioned:

  • The JIT rebuild is complete. The restart announced in Shall we play a game? landed, and the previous engine was deleted, about 19,700 lines of it. There is one JIT in the tree.
  • Encode went from 21 encodings to all 124 of the upstream distribution: the single-byte tables generated from the Unicode .ucm data, UTF-16/32 and UCS-2, the 21 multi-byte CJK encodings, ISO-2022, UTF-7 and the MIME encoded-word codecs.
  • Storable writes and reads perl5's format, byte for byte, including hooks, code references, regexps and the pre-0.7 streams. A file stored by perl5 retrieves on pperl and the other way round.
  • POSIX, Fcntl and Socket are complete against upstream's export lists: locking, sigaction, the termios constants, the socket option packers, sockaddr_un.
  • DBD::SQLite and DBD::MariaDB joined the drivers of the 0.6.8 post, with DBD::mysql now the compatibility layer over the latter. DBD::Pg gained arrays, COPY, savepoints and asynchronous queries. Peta::FFI gained closures, which is what SQLite's user-defined functions run on.
  • Coro 6.57 runs, the whole stack: EV 4.37 with libev 4.33 and Event 1.28 underneath, AnyEvent and Guard beside it.
  • HTTP::Parser::XS and Time::Piece are native.
  • The command line: the program is read from stdin the three ways perl5 reads it, the -d family reaches the debugger, -s, -T, -t, -W and -X behave as perl5's do, an assignment to $0 rewrites the process title, and --cpan installs modules.
  • The distributions' own test suites run inside our harness: 28 modules, 440 files, each compared against perl 5.44 running the same file.

The four days since 0.6.20 added 147 commits, among them Net::SSLeay with IO::Socket::SSL, the DBM family and the parallelizer on the new JIT (mandelbrot 10.85s -> 4.2s wall time). They belong to the next release post, if it remembers them.

The compression problem

The list above is in a different form from our usual posts, on purpose.

0.6.20 was 1,374 commits and about 454,000 lines of code change. The post announcing it was 12 kB, which is under nine bytes per commit. No prose compresses that far. What the post did instead was select: one topic in depth, two more in a few paragraphs, and the rest dropped. The selection followed what we had been working on in the last weeks before the release, and whatever was finished in August was no longer in view in late September. That is how a complete Encode goes missing.

There are three ways out, as far as we can see.

Less prose, more changelog. The nine items above take about 2 kB, a sixth of the last post. It is the efficient form, and it is substantially drier. A reader learns that Storable is byte-compatible and does not learn why it was not before, what the bump arena under binary-trees has to prove, or why asynchronous compilation lost. Those are the parts we would want to read ourselves.

More posts, smaller ones. A post per topic, when the topic is finished, as Doing a REJIT was. Nothing gets forgotten that way, because nothing waits for a release. The cost is on the reader's side: at the current pace this would be a post every few days, on a blog shared with other people.

Both, in different places. The complete record goes into a changelog that ships with each release and lives on the website, written when the work lands and not reconstructed afterwards. The blog post no longer tries to be that record. It links to it, and spends its 12 kB on the two or three things that are worth an explanation.

We lean towards the third. It keeps the prose where prose earns its space, and it removes the selection problem, since leaving something out of the post no longer means leaving it out. This post is the experiment: the section above is the changelog form, this section is the other one. If you read these posts and prefer one to the other, the comment field is open.

Who is this for, when AI can port your script?

A language model can transcribe code: small programs, medium ones, and with enough patience large ones. Whoever needs a Perl program to exist in another language can have that today, and nothing in this project argues against it.

What a language model cannot do is miracles. Give the best one available the task of writing a Perl entry for the Benchmarks Game that is valid under the Game's rules and beats the fastest node.js entry, and on perl5 it will not deliver, today or in the foreseeable future. On n-body the Game lists the fastest Perl entry at about seven minutes and node.js under nine seconds, and no arrangement of Perl source closes a gap of that size, whoever or whatever writes it. A top Formula 1 driver does not get a Lada Niva to 300 km/h either, and nobody concludes that the driver is the problem.

The scoreboard of the last post is the same statement from the other side. Those ten programs are Perl, they run on perl5 as well, and on pperl nine of them are ahead of the fastest scripted entry and the tenth is level with it. The source did not become cleverer in between; the car changed.

That is what pperl is: an option for Perl. Where the requirement reads "do X with Perl" and X is out of perl5's reach, X may be within pperl's. X can be speed, as above. It can be ease of use: a database script that connects without a driver build, a C compiler or an XS toolchain on the machine. It can be functionality: a GPU from Perl, or a hash whose integer-key access compiles to a probe in a register. Where perl5 reaches X, perl5 remains a perfectly good answer, and where no Perl is required at all, a port may be the shorter way. pperl widens the set of things for which Perl can be the answer, and a language model, however good, can only work inside that set.

  • Richard C. Jelinek, PetaMem s.r.o.

Leave a comment

About PetaMem

user-pic All things Perl.