Passion Projects in the Time of AI

Allow me to set the scene. Back in 2021, I was working as a quality control statistician for a major player in the Philippine fast-food industry. Actually, since the term was coming in vogue, my job title was gradually transitioning into in-house data scientist, as were my day-to-day responsibilities, which little by little started to include development tasks.

This was not a problem, as programming had always been something I’d enjoyed on the side. My passion wasn’t so much a two-track thing (whereby people had a main thing and a side thing) than it was a daisy-chain of all sorts of different hobbies and interests in time-varying and unpredictable levels of engagement: statistics, abstract mathematics, guitar (a recent addition), chess problems, and so on. Since my father taught me to use the Windows Command Prompt and programming using Borland Turbo C++ back in the sixth grade as his pilot student for a new class he was teaching, I’d prided in my ability to gain immediate fluency in any computer language that let me do anything cool.

So back to my work: one of the projects our team was pitching at the time was an on-demand analysis platform that could serve the entire department. The system at that time was that when anyone from the other teams (supply chain, restaurant systems, and what have you) needed any kind of analysis done, even if it were just the simplest t-test on the smallest set in the database (is this temperature significantly different from standard? or perhaps, is serving time consistently getting longer for this branch?), they would file a ticket on our team and we’d prepare the report for them. Sometimes there wouldn’t even need to be a report (a simple chart might have sufficed), but we’d make one anyway in keeping with the formality of the exchange.

I can’t remember now if it was me or my manager who came up with the TED talk pitch of imagining a website where anyone in the department could log in and, instead of filing a ticket, do the analysis themselves. It would be intuitive, and responsive! It would always be updated because the website would pull directly from the database! Regardless of who came up with the idea, the actual work, of course, fell upon me. Towards the end of that year, while cooped up in our family’s loft at the height of the COVID quarantine, I started developing the website that would – I suspect – land me a juicy talent retention bonus by the end of the following year.

Here’s how you build a live analysis platform. You build a simple CRUD app that pulls from the database and converts it into an all-around analysis format like R’s tidy data specification. Then you need a statistics engine living somewhere in your app’s model so that a t-test or a linear regression can be applied and spit back the results. If you’re building this in Python, such as via Django or FastAPI, then you’re in luck: you can practically just import your choice of out-of-the-box functionalities like statsmodels or scipy. But I didn’t know Python at the time. I did know C++, which meant I already knew a little bit of PHP and Javascript, and as it happens I’m the type of hammer-wielder who really insists that everything is a nail. There were some analysis packages for Javascript at the time, but they didn’t play well with each other and were incomplete for my purposes. I knew I had to build at least part of it from scratch.

I was now faced with an interesting conundrum: I knew I was going to really enjoy building this thing. Statistics, it shouldn’t have to be said, takes up a good amount of space in the aforementioned daisy-chain. Programming, even more. In fact, when I landed in Statistics due to my university admission test results, I wasn’t yet quite sold on the field. That was until I found out you really needed a lot of programming to do solid, contemporary statistics, especially with machine learning on its way in during my undergrad years. But at the same time I didn’t want to dedicate hours of labor on a project I would lose access to the moment I left the company. I wanted to own it. And in case anyone else wanted it (unlikely – data science on Javascript wasn’t then, and isn’t now a thing), I wanted to be able to show it to them and let them run with it.

The conundrum was solved thus. The CRUD application – the company owned. That was unquestionable. The statistics engine, on the other hand, was a dependency. It was an important piece of the thing, but it wasn’t exactly the thing, in the same way the Express server and the Vue framework were very important components. And like Express and Vue, I could simply import the statistics engine as another critical dependency from some other developer who also happened to be named Dominic Dayta.

I compartmentalized. While I was busy setting up the application and its interface, making sure it was well connected to the database, I let the other developer work on the statistics engine nightly. After clocking out of work, I’d put aside my work laptop and pull up my own personal laptop. I’d code the statistics engine as a separate entity. It had no reference to the company or its data. It was just a package that could manage dataframes like R’s tidy ecosystem, and perform basic hypothesis testing and modeling. It also borrowed heavily from Python and the C++ family’s object oriented philosophy in managing results objects.

The result was nodestat, which I published on the NPM repository and just had its version 1.2.0 come up this month, featuring new functionalities I never got to include back then, like Principal Components Analysis and random number generation. Nodestat became the precursor, the first of many so-called “passion projects” that I realized I never got to do anymore since coming under the pressure of college and building a career. It was a pleasure to see nodestat come alive and actually serve my company project quite well. It turned out that abstracting the results into a self-contained Javascript object made for transmitting data between server and client quite painless.

But it was even more of a pleasure to be working on nodestat. Statistics concepts can seem mundane to those who are only seeing its workings as printed out neatly by a statmodels or base R function call. But to actually write it in code, one grapples with the hidden complexities that are neatly abstracted away from the user. For instance, when building a regression model involving both categorical and quantitative variables, how do you manage building and organizing the dummy variables? For random sampling from a distribution, how do you convert probability values into pseudo-random numbers? For finding the probability in a normal distribution, how do you quickly and efficiently evaluate the gaussian integral?

And to be honest, the hardest problem of all for me in developing nodestat, leading to a feature that never shipped in my project, and one that the NPM package wouldn’t have until recently, was how to sort a dataframe object across multiple variables, in varying orders, without eating up the user’s memory? That was a tricky one, let me tell you. At least for a non-computer science major like myself.

Getting into the weeds of a project, and then seeing it come to life is one of the great joys of programming. By the time version 1 of nodestat was up, I must have stared at this line of code on my project’s README file so many times, just patting myself on the back that my crazy idea is now actually a reality.

const nstat = require('@dominicdayta/nodestat');
let stat = nstat.stat;
let titanic = stats.dataset("Titanic");
let freqDiedSexAge = titanic
.subset(col="Survived", function(x){ return(x == "No") })
.select(["Sex","Age","Freq"])
.aggregate(by=["Sex","Age"], stats.sum)
.data;
console.log(freqDiedSexAge);

Even the latest feature is a source of unabashed joy for me.

nstat.random.set_global_seed(2026);
const normal = nstat.random.normal(0, 1);
console.log(normal.pdf(0));
console.log(normal.cdf(1.96));
console.log(normal.sample(5));

Which finally brings me to the reason why this blog post is titled “Passion Projects in the Time of AI” and not “Check out my cool NPM package”. Something I’ve been grappling with in recent years, especially as I entered and thankfully eventually successfully left the job market, is whether passion projects hold much currency today. It used to be that if an aspiring developer wanted to prove their bona fides to prospective employers in the face of a lack of experience, they would set up a Github account with a portfolio of interesting projects like a custom chess engine, or a pong/pacman rewrite.

I could, for instance, simply sign up for a new Github account and quickly populate it with all sorts of projects written up in an hour or so each in Claude, Cursor, or Grok. I’m sure employers suspect the same. In fact I shudder to think of someone browsing through my repositories and shrug nodestat off as AI slop given its rarefied nature from among the rest of my code, which are usually just replication files in R or (more recently) Python accompanying my papers. That someone can just point to my sophomoric mistakes as standard AI misunderstanding human software design and not the fact that I began the project before I had any real knowledge not just in structuring a Node JS library, but also how Javascript even worked beyond the coverage of a 3 hour “Javascript Crash Course” video on Youtube. The only thing I could point to in my defense is the fact that nodestat was started before ChatGPT was even a twinkle in the vibecoder’s eye, and when Claude was just a name in an Information Science student’s waking nightmare.

So why not just vibe code away? Why get in the weeds of a programming project when Claude could probably do it much faster, and with complete documentation right from the get-go?

Passion projects, I think, are now starting to lose their currency among employers. That’s if they haven’t already completely done so. But they haven’t lost their touch to programmers who really enjoy the art and the downright messy science of it. I think back to my journey with nodestat, and the fact that it has 4 weekly downloads on npm purely thanks to (I suspect) bot activity. Nodestat never answered a burning industry need, and it never will. No one was ever going to say, “oh we Javascript developers were so bereft of real data science capability until you came along!” I knew from the beginning developers might see it as no more than a joke, like that programming language built out of the cat cheezburger meme. Impressive? To the right person sure. Employable? Ehhhhh. (Ok honestly I think any tech firm would hire the LOLCODE people.)

I think about these problems a lot because I happen to have been so unwise as to dedicate my life to two things that LLMs are gradually taking away from us bipedal meat bags, the other one being writing. I just don’t share the pessimism from my peers in either field. I’m not a complete Luddite as to completely disavow AI. I use it a lot, in fact. At work, when jumping into a new codebase, I often send it off to Cursor to give me a top-down view. It’s very good at flowcharting the whole thing and letting me know where I need to go to get started. The autocomplete feature is a nice touch too. No matter how much you enjoy coding, there’s no helping getting annoyed at the fifth module.exports or public static void main(String[] args){} you had to write in one session.

Don’t get me wrong, I think software development has changed, and will continue to change as these LLMs get better. But it doesn’t have to mean the end of human software engineering. And coming back to the topic of passion projects, I think it’s liberating to know that we can once again pursue projects that are as stupid, as useless, or as challenging as we pick them out to be, now that we can be liberated from the fear that someone at a future FAANG company we might be interested in down the line may be watching. If there’s one human talent I can be confident in that they can never train LLMs for, it’s the ability to come up with shit nobody asked for in the first place – like fucking LLM novelists for example.

By the way – did I mention I made a Jupyter Notebook companion to nodestat?

Leave a Reply