Peta::NN - Surface Detail
Peta::NN is a neural network library written in Perl: models are defined, trained, measured and shipped from a Perl script, with no dependency beyond perl 5.36. It is not a binding to somebody else's framework and it does not try to be PyTorch. It trains micro models - a few thousand to a few hundred thousand weights, each answering one narrow question about a string - and composes them: in a pipeline, or fused into one model that runs a whole batch on the graphics card without coming back to Perl in between. It runs on perl5 and on pperl, on three backends, and under pperl one of them is the GPU. Two models of 18,761 and 10,002 weights form the plural of 98.3% of German nouns they were never shown, and of every single noun of the core vocabulary.
Why this exists
PetaMem has done natural language processing in Perl for twenty-five years, and for most of them the craft was rule sets: inflection, language identification, diacritics, word classes. The Lingua distributions on CPAN are the tiny public part of that. Rules have a known weakness. They answer for what their author thought of, and the language has more words than that.
A small network trained on what the rules and the lexicon say closes that gap: it also answers for the words nobody listed. That is the whole use case, and it is a modest one. The title is meant literally: these models read the surface form of a word, often just its last six characters, and nothing else.
Perl had no tool for this. On CPAN, AI::MXNet binds a framework
that has been retired, AI::TensorFlow::Libtensorflow runs models
trained elsewhere, the AI::NeuralNet::* family is old and small,
and PDL is the numeric base without a network layer on top. None of
them trains a model to a stated quality from Perl data, and none
composes models. The usual advice is to do the training in Python
and carry the result over. We did not take it.
What it is not
Three kinds of layer, gradients written by hand and checked against numeric differentiation, no automatic differentiation, no tensor library, no images, no sound, no large models. Whoever needs those has PyTorch and Keras - for now. ;-) The library is about 7,000 lines of Perl, documentation included, and means to stay small: three concepts - data, a model, a chain - and one way to do a thing.
A model is declared, and trained to a goal
use Peta::NN::Data;
use Peta::NN::Model;
use Peta::NN::Chain qw(chain);
my $nouns = Peta::NN::Data->read('nouns.tsv',
fields => [qw(singular gender plural)])->hold_out(0.2);
my $plural = Peta::NN::Model->new(
kind => 'edit',
from => 'singular',
to => 'plural',
given => ['gender'],
reads => { end => 6 },
goal => { unseen => 0.93 },
);
$plural->train($nouns);
print scalar $plural->predict('zeitung', gender => 'feminine'); # zeitungen
That declaration is everything the user decides: which field the model reads, which it answers, what it is told beside the word, how much of the word it looks at, and what it has to reach. There is no layer list, no optimizer, no loss, no device, and no training loop. The kind of model settles the loss. Training settles the rest: it runs until the goal is met on nouns the model was not shown, widens the network if it stops getting closer, and then trains a second model from another random start to confirm that the first was not a lucky draw.
An edit model does not generate characters. Its answer is one of
the edits it has seen - "cut two characters, add en" - which makes
it a classifier, and its mistakes wrong words of the right shape
(Agronome for Agronomen) and never noise.
The data is held to the same standard. Which records are held out follows from a hash of a field's value, so the same noun is on the same side on every run, every machine and every perl. And before training, the library says what no model of that shape can learn:
7 training pairs the model cannot tell from a pair with a
different answer; a wider window would (mutter → mütter as
magnetmutter → magnetmuttern, ...)
Mutter and Magnetmutter agree in gender and in their last six characters and differ in the plural. Nothing that reads six characters gets both, so they are set aside and named, not silently counted as failures of the training.
Small, composable
The philosophy in one sentence: a task that is too much for one small model is usually several small tasks. The German dative plural is three. Change the vowel (Apfel, Äpfel). Rewrite the ending (Mann, Männer). Add the case (Äpfeln). Three models in a row are easier to get right, and to check one by one, than one larger model that has to learn all of it at once.
my $dative = chain(plural => $plural, case => $case);
print scalar $dative->predict('hund', gender => 'masculine',
case => 'dative'); # hunden
$dative->save('dative.chain');
A chain is itself a model. It is asked, scored, saved as one file,
loaded with one line, and can be a part of a further chain. Its
parts stay parts: one of them can be trained further, measured and
replaced alone, and the rest is untouched. Parameters travel by
name, so the chain above takes gender for its first part and
case for its second without anyone wiring them.
Chains are not only series. A model can judge a whole text by pooling what it says about each word, and a model's answer can choose which follower runs. The word-class example does both: one model decides whether a text is Czech, German or English, and hands each word to that language's tagger. Four models, one chain, one file.
What this buys, measured on the examples in the distribution:
| task | models | weights | not shown | core |
|---|---|---|---|---|
| German noun, singular to plural | umlaut, ending | 28,763 | 98.3% | 100% of 9,778 |
| and on to a plural case | + case | 29,419 | 98.4% | 100% |
| plural back to singular | singular | 23,783 | 96.2% | 99.8% |
| Czech word classes | class-ces | 130,460 | 93.1% | 100% of 4,306 |
| English word classes | class-eng | 173,423 | 80.8% | 100% of 3,070 |
"Core" is the part of the vocabulary a model has to get right without exception; a goal can say so, and training does not stop before it holds. The 0.2% missing on the way back are plurals with two singulars, of which only one can come back. The Czech grammar example trains 108 such models in 87 seconds; 91 of them reach 100% and the weakest 98.5%. The plural chain is a file of 118 kB.
The English row is there on purpose. 80.8% on unseen words is what a model that reads only the spelling of an English word can know about its class, and no amount of training will change that. The remedy is context, which is a further model and not a larger one.
Fusable
A pipeline has Perl between its models: the first model's answer comes back, Perl builds the new string, the second model gets it. That is fine on a CPU and fatal on a GPU, where the transfer costs more than the computation.
A fused chain is the same models, as they are, in one model, with the hand-over inside it. Nothing is retrained and no weight changes. On the CPU the fused chain agrees with the pipeline to the last bit, answers and confidences. On the card it runs in 32-bit floats, and over all 45,218 German nouns not one answer differs.
my $fused = Peta::NN::Chain->load('dative.chain')->on('gpu');
my @forms = $fused->predict_all(\@nouns, gender => 'neuter',
case => 'dative');
Milliseconds per string, singular to a plural case, three models, by how many strings go through at once:
| at once | pipeline | fused, cpu | fused, gpu |
|---|---|---|---|
| 1 | 0.197 | 0.176 | 1.52 |
| 16 | 0.124 | 0.111 | 0.061 |
| 256 | 0.199 | 0.185 | 0.011 |
| 32,768 | 0.170 | 0.136 | 0.016 |
One string alone is ten times slower on the card, and from a few hundred strings on the fused chain is ten times faster than the pipeline. Most of what remains is Perl turning strings into numbers and edits back into strings; the card's own share is under 0.002 ms per string.
There is a third form, fused and consolidated: one model trained on what the chain does. We tried it once - 21,865 weights for 28,828, 98.1% for 98.4% - and it is not the default, because it gives up what the parts were for.
That is the bet of the project: many small models, each trained in seconds to minutes and checked on its own, then composed, as an alternative to training one large model - keeping what a large model gives up, which is knowing what each part does.
perl5 and pperl, three backends
A backend owns the tensors and implements about a dozen operations
on whole batches. plain is flat Perl arrays, needs only perl, and
is bit-identical on every perl. pdl is the CPU workhorse. gpu is
one WGSL compute shader per operation on Peta::WebGPU, the binding
the 0.6.20 post, "The Player of Games",
introduced for the other kind of games; it needs a pperl built with
the webgpu feature. A saved model carries no backend and loads on
any.
Training the smallest model of the benchmark set (944 weights, batches of 16), samples per second, ThinkPad P53 with a Quadro RTX 5000:
| plain | pdl | gpu | |
|---|---|---|---|
| pperl | 31,400-32,000 | 48,400-51,400 | 22,300-24,200 |
pperl --no-jit | 5,730-5,840 | 49,400-51,700 | 22,400-24,400 |
| perl 5.44 | 5,230-5,430 | not installed | none |
| perl 5.42 | 4,420-4,570 | 38,600-40,100 | none |
Two readings. In plain Perl the JIT is worth a factor of six: the same nested loops over Perl arrays, unchanged source. And the card loses. A training step on the GPU costs about a millisecond whatever is in it, so for a model this small the CPU is the faster engine, and it stays that way until the model or the batch grows:
| weights | batch 16: pdl | gpu | batch 128: pdl | gpu | batch 512: pdl | gpu |
|---|---|---|---|---|---|---|
| 944 | 56,000 | 17,200 | 160,400 | 123,400 | 229,900 | 306,000 |
| 4,136 | 35,900 | 15,700 | 54,500 | 86,300 | 60,500 | 226,600 |
| 19,080 | 14,800 | 14,500 | 22,800 | 104,700 | 24,400 | 386,300 |
| 37,768 | 8,600 | 14,500 | 12,800 | 109,200 | 13,300 | 615,500 |
At batches of 16 the two meet at about 20,000 weights, at 128 at about 2,000, and at 512 the card is ahead for every model, 46 times at the bottom right. Further up, at 1.6 million weights, a training step takes 31-38 ms on the card and 1.5 s on PDL.
Nobody should have to carry that table in their head, so a model that is not told where to train finds out: when it is trained, the library times a few steps on every backend this perl and this machine have, and goes on where a step is fastest. It assumes nothing about the order. A graphics card may not be there, PDL may not be installed, perl5 has no JIT, and each of the three wins somewhere.
Keeping the card busy
The first GPU run of the word-class example took 42 minutes, with the card idle most of the time: a second of training, then six seconds of one CPU core checking the model. It takes 4 minutes now. What brought it there was no faster shader but fewer crossings: the check of an epoch runs on the card, through a fused model of that one model, and the samples go up once and stay there, a step picking its batch on the card.
One thing that looked right and was not: gathering the embedding gradient in one pass per table row is a sixteenth of the work and ran five times slower. A hundred long loops do not fill a card; sixteen hundred short ones do.
Try it
Peta::NN is on CPAN, with the German noun data and the three trained chains in the distribution, so the getting-started page runs as printed:
cpanm Peta::NN
Every block of code in the documentation is run by a test, and its output is what the page shows. The suite is 774 tests on pperl; on a perl without PDL or without a card it skips what needs them and passes on the plain backend alone.
The caveats, as always. It is early and there is no API promise yet. A model asked about something outside its world answers anyway: the three-language identifier shown French picks one of its three, sure of itself. And these are models of spelling; what needs a sentence to decide is not in them yet.
Do not take our word for any number in this post. The German noun example ships with its data and trains on your perl, on whatever backend you have, in a few minutes.
- Richard C. Jelinek, PetaMem s.r.o.
Leave a comment