Evidence-Oriented Programming
transcript
Uh, okay. My name is uh Andrea Stefik. I'm a scholar that works for University of Nevada, Las Vegas. Um, I work in the computer science department. I'm an assistant professor. Uh, and I'm going to be talking to you about a paradigm change in computer programming languages that, uh, various scholars have called either evidenceoriented programming or evidence-based programming. Um, and this started, at least from my story, there's a number of scholars that have had slightly different go stories on this, but I'll tell you just mine because I'm me. Um, and I started working on this problem around 2006 or so when I was a PhD student investigating uh this idea of um different kinds of people that were learning computer programming. In my case, I knew a number of individuals that were working in industry that were actually blind. Right. So they uh I met a number of people on the internet that were professional blind programmers and uh one of the one example of one of the people that I know is uh named Cena Braum. He's what's a white house champion of change if you've heard of that program. He runs a company called Prime Access Consulting and he's an unbelievable programmer. I mean just in every respect but he happens to be blind, right? So I was really curious. I mean I was a PhD student. I was young and predominantly stupid. So I thought, well, that sounds interesting. So I wonder how hard it is. Like how hard is it to program blind? Or like how hard is it to debug? Like can you debug? Or like can you use all those cool editor features that we sort of take for granted like code completion and editor hints and refactoring tools and even those little block languages like scratchy type stuff, you know, can we do all that? So I just started by talking to people. I I started uh investigating by sitting down with a number of blind programmers typically over the internet. So not really sitting down and just asking them questions and just sort of becoming their friend and trying to think about these problems. And as is probably not that surprising, programming blind is really hard, right? Even for pros, right? This isn't like a uniquely novice problem. It's just that when you're trying to represent computer programming through audio, it's just damn difficult, right? Uh, so I thought, well, okay, I know a couple blind programmers that are professionals, but what what do kids do? Like, do they have programs for this? Like, there's somewhere in the ballpark of like 40k blind children in the US, something like that. Do they have any programs to like learn computer science? Because that's a job you could really do, right? Hypothetically. So, I started asking them and I just started looking around and saying, "Well, are there blind children that want to learn this stuff?" And it turned out that there was a lot uh and um we started having them try just doing various things like try C see see how it goes. How hard is it? Is it tough or is it easy? Um and the key observation is that it's not easy. It's not just hard. It's really that's three really hard, right? There's even italics on that one. So just be aware it's tough, right? You can imagine why. Um, by the way, at the bottom here, this is one of our programs that we have through my research lab called Epic program, which is was NSF NSF funded for many years and blah blah blah. Uh, this is nowadays about half teachers of the visually impaired and about half just regular K through2 schools that actually use our programming language technologies, uh, which is called Quorum, by the way. All right, but let's get back to this. So, programming was hard for a lot of reasons. And as I sort of just sat down talking to people and trying to figure this out, some of them jumped out. So C style syntax, believe it or not, it's pretty hard when it's in audio. Well, why is that? Because you have to actually physically say for left PN int i equals 0 semicolon i less than 10 semicolon i ++ right pern. Now, how many people in this room would think that's a little easy to miss one? Yeah, right. It is. Um, debuggers at the time were either practically or actually worthless. By pra by actually worthless, I mean it made no sound and thereby you can't use it blind. By practically worthless, what I really mean is that you could use it like it might make sound, but what it would say would be so unbelievably absurd that you can't really use it for debugging, right? So like you might it might tell you that there's a line of code and it'll give you an object code. Like it'll say you'll hit a line, it'll say 137 and you're like what? And that wouldn't be like a line number. that' be just be an arbitrary number for the graphic or something like that. So no real semantic meaning. And then besides that, a lot of the features that many of us take for granted, all the features in a normal development environment, they typically wouldn't translate to audio. So sometimes they'd have special graphics toolkits that render on top of things and they weren't hooked up to screen readers or or whatever. And then on top of that, there was effectively no educational programs in the United States uh to actually learn this stuff. which is hard. By the way, that's what they gave me my PhD for. So, oh wait, wait, I forgot my joke. I'm so sorry. I said being a naive little kitten, maybe I can fix some of these problems. Haha. Okay. Sorry, but so, okay, so understanding C and audio is hard. Well, that's cool. But I had this nagging question. Why does C have the design that it has? Like, what's the evidence for it? Like, do we have evidence for the word for? Like, why is it plus+? like what is the data on that and I was curious so I started to try to look it up so I just started reading a lot of papers in programming language design and there's lots of proofs I think most of us know that right we we do proofs on programming language design all the time and have for many many decades right and we gather lots of performance data right like we want to know how fast it is because that's really important right if you want to scale if you want it to actually work then you do proofs and performance data that's sort breadandbut programming language design computer science but I found something unexpected and that was that there didn't appear to be any human factors data on programming language design these changes were just sort of created so at this point in the talk I actually have to take a step backward in time just a little bit and the reason for this is I'm going to ask uh how many of you have actually studied experimental design in college like a handful how many of you know the history of behind experimental design. Yeah, like a small a couple people. So, usually I like to talk about this stuff a little bit to give you a sense of context, right? So, let's uh jump back in time and let's look at this paper by this fellow named Ted Captchuk. He's a a medical researcher that studies placebo and he wrote a paper a few years back that influenced me a lot on sort of the history of medicine. Why why are studies designed how they are? What's the evidence for how you gather evidence? Right? That's a little meta, I suppose. So um believe it or not there's about five phases according Capchuk um and he stops around the 1940s or 50s. Um the first uh phase in at least the medical history in some of those areas was really to detect fraud. Right? There was a lot of so to speak snake oil salesman in the late 18th century. By the middle of the 19th century it's there's this sort of weird quirk of history where the actual homeopathists were actually doing the best experiments for the time which is sort of ironic. We'll get to that. Excuse me. By the late 19th century, it turned out that psychology got involved and they started doing things like double blinds, had some randomization, stuff like that. uh but after that pharmarmacology got involved and then by around the 1935 or the 40s or 50s especially we started to see what has become sort of standard in effectively all other disciplines except computer science which is the randomized control trial um and even though captic actually stops around the 50s it's not as if other scientific disciplines said oh RCTs are the way that that must be the only thing we can do in fact other disciplines have gone substantially especially forward since the 50s. We now have uh been rethinking statistical practices to sort of increase replication in studies. We have uh standard replication packets. There's this concept of registered randomized control trials in medicine which became popular nowadays for corporate fraud actually um amongst other reasons. Anyway, you get the idea. So, let's actually talk about these phases individually to give you a sense of the kinds of things people were fighting against at the time. So like what kinds of science were they doing to try to do a little bit better in our knowledge? Well, it turns out that one of the first things that uh scientists fought against in the um late 1700s was actually spearheaded by Benjamin Franklin, like that actual guy on your all the $100 bills in our pockets, right? Or $1 bills. I'm a professor. Come on. Um I'm kidding. We get paid fine. uh and he was actually hired with a commission by Louis the 16th um to evaluate this thing called mesmeriism. How many people have heard of this? So there's like a million different versions of it that are all slippery and try to be as untestable as possible, but um it's basically a sham movement from the late 1700s that didn't work right. And Benjamin Franklin along with his crew um used this concept of sham treatments. Basically they tricked people into thinking the mess was healing them. they really weren't doing anything and lo and behold mesmeriism doesn't work. But interestingly, even though Franklin had basically discredited it by the end of the late 1700s, it basically persisted in medical science and textbooks and all that kind of good stuff for like a century or so, right? Because, you know, the experiments weren't good enough, right? Which they actually weren't, but that's irrelevant. So, so um by the mid 1830s or so, uh by the way, homeopathy came around also in the late 1700s. Uh however, it didn't really get tested for around 40 years, right? Because, you know, they were busy. Um and the first time that this was tested was done by this fellow named Arman Truso, who basically had did the first placebo test. At the time, they used something called bread pills. To my understanding, it just means bread in a pill. Uh but maybe I'm wrong but I think that's true. Um and by the late 1800s or by around that period we started to see more techniques like double blinding. Right now double blinding believe it or not at the time actually didn't work very well because blinding doesn't work very well without randomization which you'll see in a moment. Uh, anyway, homeopathy has been thoroughly debunked since literally 1834, yet is still funded by the UK today. Ouch. If you ever have time, economics. Say what? So is austerity economics. That I cannot disagree with that at all. Fair point, sir. Okay. Now, psychology started moving forward in part because they were investigating things like sensation. This is like if I had two weights and I wanted to see how much I could sense in the weight, that's like a sensation type test or perception tests. Um, but they were also investigating sort of fraud like psychics or talking to the dead or weird junk like that. Right now, psychics in this case, believe it or not, they actually introduced randomization approximately this time. Think of the idea like I have a deck of cards and I grab one at random and I try to psychically project it to you, right? Which doesn't work. Um, that you know that's randomization of a form. It's not very good randomization, but it's something again guard against quackery, fraud, and bias. Oh, you know what? I forgot to say that. By the way, if you don't know, quackery is like not like a making fun of term. Quackery means that you believe something that's thoroughly debunked, right? That's really all it means. Fraud means you're telling people something you know is a lie, right? All right. By phase four, we have this sort of icky story, but I'm going to tell you anyway. There's this guy named Charles Edward Brown Sequard, which I'm sure I'm butchering, but forgive me. And he actually took guinea pigs and pulled material out of their testicles and was injecting it into people, right? As you [Music] do. And interestingly at the time experimentation was considered like whoa what do you mean experiment that's BS and so what many people in the medical committees wanted was actually case histories. So think about what this means for a second. They say deploy the testicle material worldwide and then let's see what happens, right? Kind of like lambdas in C++, right? Just saying. I'm kidding. Actually, there is no studies of that, right? There actually isn't before it was deployed worldwide. Just saying. Obviously, it's not testicle material. But interestingly the best part about pharmarmacology is by this time they had seen experiments. So the community very quickly said what the hell and they started looking up and figuring out how to do a test. So within like two months of this guy having his claims they started doing experiments. Now their methodology was all wrong. They didn't really know how to do it. But they started experimenting using different kinds of placeos, actually having control groups and these sort of things that we sort of take for granted until we eventually hit phase 5, which is actually not modern. It's not modern randomized control trials, but it'll it'll give you a sense of the pathway. Specifically in 1935 this fellow named Fischer wrote this really influential book called the design of experiments which I imagine you can guess is about the design of experiments and in this he really formalized mathematically the concept of randomization. Now this was crucial to the design of experiments broadly. This is not just like a an immaterial mathematical theorem. Without this, we would not have randomized control trials. It's very important. Now, interestingly, at the time, many in the medical communities considered this to be esoteric, right? Like doctors are okay with blinding, but you know, people die if the treatment doesn't work. So, they would cheat a little bit or they would have these little tactics. It's it's not so much that they were trying to commit fraud. It's just they were worried about their patients. It's legitimate to a certain degree. But the problem is since you can't get good data, good information, you needed some way to guard against it. So eventually even though in 1955 the journal of the American Medical Association wrote uh treatments should be kept in the hands of the general practitioner and not on the basis of experiments even though that's true by the 50s eventually medicine formalized this into what we now know as phased control trials which are not perfect by any stretch. We still today have corporate fraud and all sorts of other problems. But nonetheless, it's an important step on the evolution in scientific experiment designing. So the first randomized control trial is actually done by this fellow Austin Bradford Hill. Oh, by the way, he didn't use the word randomization in his study because he was worried about scaring off his colleagues. It's interesting. He was controversial. All right, so here's the key observation. Scientists over time have increased their evidence standards to avoid bias, fraud, and other issues, settling on randomized controlled trials as the gold standard when you're evaluating things involving human beings, right? For good reason. I've given you plenty of examples about fraud and other issues. But let's get back to programming languages because this seems like it's unrelated. So, here I was. I was investigating programming tools for the blind. I was working with kids. I was talking to professionals and it seemed like things were hard. I was inventing a whole bunch of tools, auditory code hiders, little navigation tools, talking debuggers, all sorts of little stuff. Um, and they seem to help a little bit, but C itself seemed really difficult to hear. And I could got not get this nagging question out of my head. Why does it have the design it does? And where is the data? Where is the data? So I started looking and I recruited a team from all over the place. Germany now Finland, there's people in Chile, there's people all over the United States. And what we found was there was no evidence. In fact, systemically since the 70s, programming language designers have systematically refused to collect it. Think about that. They have ignored centuries of empirical study design and simply not collected data. Now, here's the key observation. Historically, anecdotes have not been helpful in evaluating things like medical science because it's too complicated. It's just the world is a tricky place. And anecdotal evidence is actually ills suited to a task as complex as understanding the programming language wars that all of us have to deal with all the time. even if we don't really think about it much. So we created a paradigm called evidenceoriented programming and I have to give credit where credit is due. There was another scholar in Germany named Stefan Hannenburgg who at the same time that I was working with blind kids was working on aspectoriented programming because he couldn't get this nagging question out of his head. What's the evidence that it's actually effective? And it turns out it's actually not but he didn't know that yet. And so he came up with the same idea at basically the same time. So here's the basic idea. Design features of a language must have corresponding evidence from appropriately constructed randomized control trials using modern standards that we know from other scientific disciplines like psychology, medicine and other places as best as we can. There's actually some systemic problems in computer science preventing us doing from some of the most modern especially registration of control trials. However, it's the best you can do given the culture. uh language authors must allow external community members to suggest changes to a programming language API or whatever using the actual scientific method, right? Not anecdotes, the actual scientific method as known in other disciplines. And finally, evidence from other techniques, software repository mining, educational data, whatever is fine to influence programming languages so long as it's rigorous. Now, that's an important thing because even though randomized control trials are useful for some things, they are generally smaller in scale than the things you can do by looking at like a a GitHub repository analysis because you get lots of data. So, it's okay to look at data across different paradigms and try to reverse engineer what will help humans. That's the idea. So then my wife and I, who is a compiler writer, um invented Quorum, which I keep hearing online was crowdfunded, but it was actually funded by the National Science Foundation. Um to my knowledge, it's the first and only programming language to use a scientific evidence standard for human factors. Of course, we run tests as much as you do too, to make it work, make sure it works, but that's a different issue. So let me give you sort of a short tutorial on the types of things we do when we evaluate things like syntax and semantics choices as opposed to what you're used to in programming languages. For example, how many can look at this phrase here just by show of hands and tell me approximately what you think that would do. Yeah. So can people with no training because we've tested it, right? Middle school kids, they know what that does, too. How efficient is that? It's just as efficient. Doesn't matter, right? It's just a phrase. Now, the way we do this is we usually start out with something like surveys. These are not RCTs, but they're useful because we can often figure out what kind of words and symbols make sense. So, let me sort of show you this. If you look at loops here, if you take non-programmers, people with no training. By the way, you want to do both because there's a bias effect and there's this equation where you can predict it. That's a long story. But noviceses usually tell us words like repeat make sense, which they kind of do. And ironically, if you've never heard this result before, the three worst choices in every replication of every study we've ever done have been four, while, and four each in that order. And I swear to God, I disbelieved that result so much I replicated at another university. Same result, which is funny. Um, so, so, but the thing is evidence-oriented programming doesn't mean that you like through magic have one language to rule them all, right? Like there are many contexts of use that people use programming in and that's reality, right? A NASA engineer just isn't the same thing as a child making a PowerPointish type slide thing in Scratch. It's just not the same. Uh and in addition, evidence from some of these st uh studies can actually be very confusing sometimes. It's not clear what the answer is. And a really good exh example of that is actually the exception system in Quorum, which looks kind of like this. Um, specifically whether exceptions are helpful in a language, I actually haven't found any evidence of yet, but um, whatever. Which words imply best the concepts is a little unclear, which I'll show you some data in a second. And the exact structure of the code, we have a little bit of evidence on, but I'm not very convinced by it. It's like a start. So, here's an example. If you look at some of these words, uh, by the way, we don't show people the words try, catch, and stuff like that. There's a scientific methodology behind how you do surveys and you can look that up. It's not like it's way too complicated to describe right now. But if you have non-programmers, you get words like check or test, things like that, which kind of sort of makes sense, but kind of not also. So maybe the non-programmers just didn't understand what we're asking. The programmers, they seem to sort of understand the concept. They use words like try, but they don't like them. So I don't know what does that mean? Well, it's hard to say exactly, right? But the thing is surveys can only tell you so much. So knowing the history, um the next thing we tried was sham treatments, right? Same thing, right? Kind of like placeos, except you have to come up with an idea for what a placebo would be in computer science. This result is kind of wellnown, but I'll say it anyway. Our idea that we had initially for a placebo um was to randomly design programming languages. So literally I had my graduate student sit around and roll dice laughing her ass off and then eventually we came up with a language called randomo right now. Yeah, it's wonderful. It was ridicul. But one of my colleagues said you'll never pass peer review with ridiculo. They'll just get pissed off at you. So which is probably true. But um I actually appreciated that comment. So the name is up for grabs then. You may have it. You may have it. Absolutely. I did not I did not trademark. Um so one of the things we found right away was uh in multiple studies it turns out to be the case that the programming language languages Java and Pearl for novices are actually no better than Random O. I'm going to stop for a moment. Now what's interesting is when we first did these studies we had no idea why, right? We just got these effects. We're like, "What the hell?" As we stopped laughing. Um, but it turns out we came up with an idea which ironically I figured out how to do from designing talking debuggers for blind kids called token accuracy mapping. And believe it or not, it's a DNA processing technique. I went to your DNA talk and I was thinking of it the whole time. Um, but anyway, the idea is you basically splice what the programmers put out as this sort of fake DNA sequence and you can get a map as to which tokens and symbols cause the problems, right? It gives you clues clues and it looks kind of like this. So on the left here, this is some data from our experiment. This is a repeated measures design. It's way past 1935, but nonetheless, it's very well known in the academic literature. You can look it up easily. Um, but here's an example of a token accuracy. Now, this is actually Quorum before version 1.7. And what happened was when we ran our study, we found out that we had messed up so many things on Quorum. And we basically found data on particular tokens. Like if you have a then at the end of an else, there's only a 25% chance in our replication that a novice would type it correct. Why? Who knows? But I'm sure as hell removing the then, right? And this is actually was interesting to me because I assumed at the the get-go that natural language would make sense to noviceses. Turns out not true. What matters is the specific word choices that you make and the specific ordering that they're in. Right? So if then makes sense to some of us, I think, but actually doesn't make much sense to noviceses, which is surprising to me. uh braces and stuff don't do any better, but Ruby did a great job because it just didn't have tokens for that stuff. So, guess what we did? We copied Ruby, right? Oh, by the way, except for the equal sign, that one, it turns out if you use a double equals instead of a single equals, it causes an eight-fold increase in their lack of understanding or accurate, whatever you want to call it, right? All right. So, what about other systems in quorum? because we're trying to build a language based on evidence knowing full well that there really isn't a whole lot of studies and that unfortunately we have to write a lot of them oursel as the literature catches up right well what about type systems well this one's actually more well known in the last three years largely because of Stefan Hannenburg who was doing this kind of stuff as I was playing around with kids and words now specifically the data from this shows pretty clear evidence that static typing actually does improve your productivity for people like you, right? Professionals or people that have a little bit of experience, right? But here's the catch. If you're a novice, you get a statistically significant negative impact, but it's pretty small, right? So maybe you see where I'm going with this. If you're a novice, you get a significantly negative impact, but it's kind of tiny. It's not that big of a deal. On the other hand, if you're a pro, you get a significant bump and it's much bigger. So who should we wait for? Probably the pros because it makes a bigger impact. So guess what we did? Oh, by the way, if this theory is correct, it implies that sometime in the novices's sort of experience, there should be a drop at some point in the type system errors. If this theory is true, which is a directly testable hypothesis, unlike a lot of the literature, right? I'll show you some stuff that just came out a few months ago, which blew my mind. So here's some example of some of the data from the static versus dynamic tests that we've done. This one is at Ixie. It's a top conference in software engineering. And static uh tends to do better than dynamic. How you analyze this is use something called a two-factor an ANOVA. It's a standard technique again way past 1935, but nonetheless it's well known in the literature. You can look it up. It's easy. On the right here is this conference oops. Although the data looks kind of messy, it's basically the same result more or less. Now get this. It turns out a few months ago these fellows in the UK, Aladri and Brown, they came up with a way to analyze out of the Do you know the Blue J group? Have you heard of that? Okay, not many. Um there's a group over there and they basically tracked compiler error for students. They gathered 37 million of them, right? I mean it's a lot of data, right? And take a look the green there's a typo in the slide that should say type there. I asked him. But this green line here is type systems. Notice at month nine, we see a drop. It doesn't prove the theory is correct, but it doesn't disprove it. It's at least plausible. But interestingly, if you look at over here on this graph, you'll notice there's this weird spike around the same time. This drove me crazy for months. I could not figure out why this spike was there. I pestered uh Neil, this he's the professor, um a hundred times. It turns out it was a bug in his analysis script. So, we don't know what that is. It's like a bug has something to This is a time measurement and some people didn't ever finish. So, it has they need to scale it in a certain way. So, we'll see what that shows. But anyway, just be aware that's what's going on there if you read the paper. So, he emailed me really apologetically and I was like, "It happens, dude." So, this is science. Well, it's true. Okay, so let me break down a actual action. It's called an action because um that's what the data shows uh in quorum for uh actions are just methods or functions that sort of idea. Now uh in quorum there's a whole bunch of different regions here. So if I look at the method up there get class sorted in package that has Pascal case, right? So why Pascal case? And that's because of David Binkley's paper from 2009 which showed compared to underscore you get a teeny teeny tiny increase in comprehension. It's really small but it's nonzero so we may as well use it. Um for these word choices like end, return, repeat, repeat while these come from a combination of our surveys and token accuracy map results because we have we literally have data on like that word and from a number of different perspectives. So, for example, if a novice uses that word and repeat, most of the time the data shows they get it right about 67% of the time as opposed to the word for which is usually around 18. I mean, this is really specific. This is not like we're making this up. There's a lot of data. Besides that, you can see I already talked about that equal sign there. Uh, but one of the interesting parts here is this uh text package key. Now, text and string actually both do pretty well in surveys. The word string is just fine. text is shorter. So we tend to reduce typing if we have the choice. But nonetheless, but interestingly the question is if we have this compromise for static and dynamic typing for novices versus professionals, what is the compromise? So in our case the what we think the data tells us as best as we can today we have to have the static declarations on the bind points otherwise all of you folks get a negative impact. Uh, by the way, in the studies, we videotape people and it turns out that if you watch people just sit there and do it, when they um the time difference you observe in static versus dynamic typing, they literally jump to another file and are looking for what type to pass. So, I mean, like it's pretty clear from the video tapes that that's what's happening. Um, but if you look at the inside of functions, the inside of functions, the impact for noviceses is pretty low and you can use some amount of type inference to get rid of it. We know this because if you look at static and dynamically typed language on the inside of functions, you can tell exactly which regions are problematic. So you can have a little bit of inference, but unfortunately we have no idea how much inference should be allowed. There's a number of options that you can make, but we don't know what it is. Yeah. Is there any evidence for the inclusion of undefined? Oh yes, actually there is the word. Oh, wait. For the inclusion? No, there is not. And I would love to see a study on that. I would love to see a study on that, but I don't know of one. So, um, yes, there is for the word, but that's pretty shallow, right? Okay. Now, a couple other regions I want to mention. If you look at the hash table mark here, it looks like there's some form of generics. Now, unfortunately, every study we've done on the syntax of generics has failed. We basically don't have a good design for the actual syntactic properties of generics. However, what we do have is a study by Hannibal showing that it increases productivity of all of you, right? The amount it increases is not that much. It's just a little, but it's non zero. And interestingly, if you look at the study by this fellow named Chris Parn, if you don't know his work, he's a totally awesome scholar. Um, he did basically an analysis of all of GitHub to look at what types of generics PE features people actually use. And the answer is not so many. Right? If they use generics, it's usually list string, like literally just list string. And then all the advanced features like question mark extends or all the stuff they have in Java, they're just not used that much. Now, that doesn't prove that you shouldn't have them in a programming language, but nonetheless, that they weren't adopted is an interesting data point, right? It's at least something. And then finally, this word returns on the top right there. Our data here is actually pretty bad. We don't have a good choice for the word. So the only thing we've been able to think of is to make return types quote unquote optional, which just means if you don't see one, the compiler ignores it. But if there is one, you're stuck with the word returns, which noviceses hate and always get wrong. So the point is evidence-based programming doesn't mean quorum is this magic like bastion of goodness. It means that we're actually using data that is refutable and we're using the best available evidence to date which is not that much in the literature unlike medicine which might have thousands and thousands of trials. Right? So what about other systems? Well, you can look at inheritance. There's only to my knowledge three studies ever on the topic of inheritance. Uh, one is by Walter Tiki, which is a great study, but somewhat inconclusive as terms of like like for example, um, we don't really have a good handle on multiple versus single versus mixin or any of those kind of questions. To my knowledge, it's never been studied. It's kind of important. Um, there are quite a few studies on compiler errors. Believe it or not, even professionals are profoundly impacted by the design of compiler errors. Google actually did a study on this not that long ago, finding that their engineers just got totally flabbergasted by certain kinds of errors. And if you want to know which ones, you can go read the paper. It's fascinating. Um, in terms of concurrency, there's a couple studies now. Probably my favorite so far is Chris Rossback's work out of UT Austin, although I think he's at Microsoft now. Um, and they basically compared transactional memory versus locks and they found that transactional memory led to fewer bugs in code. Now, of course, that's easy to say, but we all know that transactional memory has like technical problems as well depending on your point of view. So, this isn't a panaceia, right? It doesn't magically tell us which things are good and bad, but it's it's the point is that we're starting to collect evidence which never happened before, right? So from our list of thousands of papers, quorum to my knowledge takes into account every paper that exists in the literature so far on human factors of language design. So why should you care about this? So I want to I want to end with sort of a brief explanation of why I think uh not just noviceses, not just professionals, not just teachers, not just businesses, why I think people should care about this problem. Notwithstanding the fact that programming languages make up literally the entire foundation of our modern technology, right? And it hasn't been studied in depth. But I think programmers should care because there is solid evidence that these issues impact productivity, right? When I program in languages that make flaws, it ticks me off because they could look it up for some of the things that are known, right? So, you should know that. You should know that there is research on some of these things and you should care. There is solid evidence that language designers are not following historical scientific norms. So if you're on the internet and someone says this language is better, you should say do you have a randomized control trial as evidence. And if they say, "What's that?" then you should say bye-bye. Right? And I'm I'm completely serious. Right? In other fields, scientists would not stand for the evidence standard that they use in computer science. It simply would not pass peer review ever. But in computer science, it's fine. Why should teachers care? Well, there's solid evidence that these issues impact students. Now, some people that I've talked to say, "Yeah, but I don't care about students. Students aren't professionals." And that is true. However, sometimes the design decisions that you make for students can impact may not have any impact on professionals, but they might like help them and not have an impact. So, for example, if you all had to make the ultimate sacrifice and type repeat 10 times, I think you'd be fine. You know, you'd probably understand it, right? It wouldn't be that big of a deal. But for a novice, it makes a big difference. Like a really big difference. U besides that, languages change all the time. So educational institutions and students actually have to pay money when these things change. So when C++ 11 comes out, they have to buy new textbooks, which either the university or the students have to handle. That's a problem because education is already exceedingly expensive. And this probably isn't helping right now. Why should businesses care? Because productivity costs businesses money. And I'll talk about this a little bit more in a second. In fact, the language wars itself may be costing industry a fortune. and we have no idea how much it costs. So here's the key point. As programming languages change or as you use particular technologies, consider carefully whether the designer is following historical scientific norms. If they're not, it may be costing you or your organization time or money. That is the facts. So let's break it down with a thought experiment. Now I don't mean this as some sort of like serious economic analysis. I'm not an economist. uh and don't pretend to be. However, suppose just for the sake of argument that developers lost one hour per week, just one. How many of you have lost one hour per week to something that made you less productive? How many have lost five? Wow. How many have lost 10? How many just don't want to go to work anymore? Okay, so I have no idea how to measure this, right? I've thought about it, but I don't know how to measure yet. But if you made the assumption that developers lost one hour of productivity due directly to language design considerations, which I have no idea how to measure and I admit, and you use the Bureau of Labor Statistics average for software developers, which is 44.88 an hour, and you have 10 developers, that costs you about 23,000 a year, give or take. Now, of course, that doesn't mean you fire people. It might mean you're more productive, but whatever, right? If you have 10,000 developers, which is not even really that big, right? You could have a company like at JP Morgan or other places where that's totally feasible, that would cost you $23 million a year. It's a lot of money. And this doesn't even take into account all the other stuff that goes along with programming languages as they change. As new toolkits are introduced, uh the language might change. That might have to require training time. It might require that you go to conferences, read new books. That takes time, right? And that time costs money, right? Now, that doesn't mean that we have an easy answer to this problem. But the point is, while not intended as a serious economic analysis, productivity issues may be very expensive for industry. So, for example, in the medical studies, they went to randomized control trials because of death, right? Because they were literally killing people, right? We don't have that problem. We're not going to like if you have to type, you know, a for loop, it's not that big of a deal, right? Especially for pros. But nonetheless, um the issue we would care about is how much is it costing us, right? And it's probably not cheap. No one knows. There is no studies on it, but probably not cheap. All right, so that's what I've got. Um, usually at the end of these talks, there's always a couple people that want to know if they can contribute anything. And sometimes running experiments is hard, but there are a whole bunch of things that you can do to contribute to these projects if you think it's interesting. And if you don't, that's okay. Um, one is you could help us flesh out the standard libraries. We have a bunch. We have game engines. We have Lego robot libraries. We're working on a linear algebra library. We're working on some science stuff. All sorts of things. But nonetheless, um there is no way in hell that we could do all the cool stuff that you can do in many languages without people like you to say, "Dude guys, that library sucks. Please let me rewrite it." And that would be okay. Um do a small project in quorum. So a lot of times when you're in academia, you test more with students than with professionals. We do a little bit of both. Oh, okay. Um we do a little bit of both. But if you try a small project, email us and give us feedback. I need libraries that do this. I need these things. And then finally, if you really want to, you could come to our conference. It's invitation only, but we do let people apply. It's usually relatively small. Um, Quorum's a JVM language. You can call down to Java C++ if you want to, stuff like that. Um, and there's the website. Thank you very much.
- from
- Talks
- added
- 2026-10-10
- likes
- 0
similar
-
Bret Victor - The Future of Programming youtube.com
-
Faith, Evolution, and Programming Languages youtube.com
-
The Future of Programming youtube.com
-
LLM-assisted coding - A Systems Perspective youtube.com
-
The Programming Language Wars youtube.com
-
Programming is Writing is Programming youtube.com
Talks › Categories > Software Design: “by Andreas Stefik (Strange Loop 2015) [41:42]”