Saturday, March 19, 2022

Please, Stop Collecting Developer Opinions

Just recently, I was quite enthusiastic to read a software science paper, because its title sounded promising. I do not want to refer to this specific paper, because it is neither the goal to discredit one specific work nor one specific author. It is a set of similar papers that needs to be criticised.

Over the last years more and more questionnaires can be found at scientific venues in software science. The commonality of these papers is that authors ask developers about their opinions on some topic. Then, responses from a huge number of people are collected. Then, the results are analyzed.

So far, there is no problem.

There is nothing wrong with opinions. It is interesting to know what peoples' opinions are. It is especially interesting from a marketing perspective, because it says something about the perception of people.

The problem lies in the conclusions.

What a number of works in software science do and which is fundamentally wrong is to infer from subjective perceptions something about the perceived phenomena.

Let's assume there is a technology X that tries to make developers' life easier. Then, someone asks developers whether it makes their life easier and let's assume with high evidence the answer is yes. What can we conclude from it? 

We can conclude that developers think that it makes their life easier. It is also possible that developers just pretend that it makes their life easier. But we do not know whether it makes their life easier. Making any claim about how or whether the technology influenced a developer's life is not possible from the evidence gathered so far. 

In order to find out whether the technology X helps, no result from subjective perceptions would bring us closer to an answer. Whatever the result of the questionnaire is, the question whether or not the technology helps is still unanswered.

One could argue that a developer's life becomes better because he thinks that there is a technology that helps him - independent of whether or not it actually helps him. This kind of argument is comparable to a placebo argument that we frequently find in homeopathy. But it should be clear that this argument should not be used in software science, because it is more a meta argument: if something makes people think that it makes their lives better, then it is good. The argument is comparable to the question whether a free beer makes a developer's life better.

Of couse, this leads us (as always) to the need for studies. But the argument is not that questionnaires are no studies. They are. But the problem with them is that they purely depend on subjective perceptions. 

There are good reasons why you find whole textbooks about perception in psychology. Perception is not only subjective from the perspective that people can perceive the same phenomenon in different ways (because of differences in physionomy, differences in experience, etc.). Perception is influenceable. You can easily find a bunch of studies that show that perception can be influenced and the Pepsi versus Coke experiment [1] is just one example (again, whole textbooks are on that topic, there is no point in giving a longer list on that here).

So, what is actually the problem? When we study technology, we need to measure interactions with the technology in a non-subjective way. You can still ask developers questions. "In this scenario, what is the outcome?" could be an appropriate question. But it differs from a question "Do you think that technology X helps?".

We need to stop asking for subjective perceptions.

The implications of the statement are much more serious than we think. Community processes that can be found today for example in programming language design typically ask people about opinions. But it leads to far to discuss this issue here.

Please, stop collecting developer opinions.

Opinions are important. But they do not permit to draw any conclusion beyond peoples' opinions. And should a technical discipline have the focus on opinions? I think the answer is no.

Feel free to leave a comment.

Thursday, October 15, 2020

What Should Software Science learn from the Corona Crisis?

What Should Software Science Learn From the Corona Crisis?

The corona crisis does not only influence people's daily life, it also influences how people think about science. Suddenly, scientists are present in the news, scientific results influence new laws that appear because of the corona crisis, and the results of scientific studies become part of people's daily conversations. Actually, this new popularity of science is good.

And there are a number of people who doubt in scientific results. Actually, this is not that bad. Science requires doubt. Progress happens, because some people do not believe in commonly accepted theories and search for alternative explanations or new interpretations for given phenomena. 

What's bad is when people ignore results or invent new theories without having any evidence for them. And what's even more bad is, that there are people who follow such new theories without even demanding evidence. Such people can be fooled too easily and for other people fraud becomes a profitable business.

People's typical reaction on skeptics of the corona crisis is that education would help. If people would be better educated, their knowledge about science helps them to distinguish between serious interpretations of scientific results and rather wild guesses based on personal anecdotes. But while this statement is probably true in general, we cannot assume that every discipline provides such a profounding knowledge. 

Taking into account that the scientific foundation of software science is rather low it actually makes sense to think the other way around: what can software science learn from the corona crisis? So, why not trying to find some "lessons learned" for software science from the ongoing crisis?

It is the numbers that do matter

It sounds stupid to point this out, but the first thing to be learned from the corona crisis is, that it is the numbers that do matter. 

The first and rather obvious number that directly comes to one's mind is the death rate. But other numbers such as infection rates, etc. do matter as well for medicine. For other disciplines such as economics numbers do matter, too. There, monetary aspects such as the costs of the crisis do matter. In the end, it is not a single number that matters. Each single number plays its role, but a number of different numbers need to be taken into account in order to get the big picture. 

But the important insight is not only that numbers matter. The important insight is, that hardly anything else than numbers matters. Even if there is a person who has a strong believe in the effectiveness of some medicine, therapy or vaccine, it does not imply that such statement should be taken too serious. Even if someone ignores the huge, negative impact of corona on the economy, it does not imply that this negative impact does not exist. Rhetorical skills of people might strongly work on some people. But rhetorical skills hardly change the reality.

In the end, the effectiveness of some treatment, or the validity of an argument requires numbers: we want evidence for statements and not only some famous peoples' believes. Software science should demand evidence. Software science should demand numbers - numbers that do matter.

Numbers are dirty

The second lesson to be learned is, that numbers are rarely pure. Numbers are dirty. Empiricists are aware that measurements are rarely as pure as people would want them to be. Measurements imply errors in measurements, measurement tools have their drawbacks. Although people want measurement tools to be as precise as possible, we have to accept that every measurement tool has problems. 

And empiricists are used to the problem that people, who do not accept empirical results, discredit the numbers. In the corona crisis, the death rates are discredited. People doubt, that the number of deaths is valid - and they have good reasons to doubt in the perfection of the reported numbers. Obviously, there is no independent institute that can analyze for every single case whether a person died just with or from corona. We do not necessarily speak about intentional lies. We speak about cases where it is not clear whether there was a causal relationship between the virus and a person's death. And we have to accept that even corona tests can fail. It is the nature of measurements that there are error rates - and it is the goal to reduce such error rates.

We also see that the numbers are attacked on different levels. For example, we find people who doubt whether the reported death cases are actually true, i.e. people argue that there could be additional corona cases that were intentionally not reported. Or some people argue that the reported number of infections are too low, because some governments are not interested in reporting high numbers. And even if numbers are accepted, people who are not willing to accept empirical results start new interpretations. For example, people argue that a high infection rate is not the result of an ongoing pandamia, but rather the result of a high number of tests. Or a high death rate is not the result of failing countermeasures, but rather the result of an extremely aggressive virus.

In the end, we have to accept that all numbers have their problems. This does not mean that we should blindly trust in all reported numbers. It is important to see how the reality is mapped to numbers and it is important to understand potential problems. But just stating "the reported numbers are wrong" is rarely a constructive criticism. There is a need to understand how problematic a number is, how measurements could be improved, etc. And it is always necessary to question the relevance of reported numbers. And it is important to identify people who discredit numbers for rather personal reasons and who hinder that way the process of knowledge gathering.

For software science, the lesson learned is that we should not be too quick to discredit reported numbers. We need to understand the process of data collection (and interpretation) and need to understand how large possible errors of certain measurements techniques are. This means that finally we need to identify relevant measurements for our discipline and we need to define measurement techniques in order to get valid measures. And we should be cautious with people who discredit numbers for the sake of discrediting numbers.

It is not a single study that matters, it is multiple of them

Another lesson learned from the corona crisis should be that scientific knowledge usually does not arise from a single study. 

Up to now, there are hundrets and hundrets of studies on corona from the field of medicine. And our knowledge on corona is the results of a combination of a large number of these studies. This does not means that each single study is fantastic. There are in the meantime a number of studies which are today considered invalid. And there are studies who just reproduce results that already have been already reproduced by others.

But the essential lesson learned is, that people in a mature discipline study the same phenomenon over and over again from different perspectives. Different experimental designs, different treatments, different measurements, different measurement methods, etc. -- the knowledge on the field consists of multiple tools and efforts in order to get the big picture.

Such effort is required in software science as well. Instead of celebrating novel ideas in our field, we should appreciate more studies that study given phenomena in depth. We should collect multiple studies on the same phenomena. We should encourage people to study phenomena, although there exist already some studies on such phenomenon.

The Need for Education and Demystifying Science

These corona times teaches us, how necessary it is that people understand non-subjective reasoning and how necessary it is that people distinguish between fact and fiction. Unfortunately, this requires education. It is not enough to argue for or against a statement by adding a phrase such as "scientific studies have shown" to it. Science is not a a magical process. Science just means to be as much non-subjective as possible. Science tries to run, collect, summarize and interpret studies without any agenda in mind. Education demystifies science and statements such as "there is a scientific study" start losing their authority -- which is good, because there is the need to understand studies and not only to accept an author's interpretation of a study.

People should doubt in the results of studies. Such doubts must not be some naive scepticism. It requires knowledge about the underlying procedures and it requires knowledge about the lines of reasoning built upon collected numbers. The necessary willingness to doubt in results also requires knowledge and recognition of valid results. Knowledge about methods teaches us where the limits of doubt are.

Infortunately, this is probably the biggest issue for software science. Actually, it is not clear whether software science provides to its actors enough knowledge for the mentioned kind of reasoning. There are even reasons to believe, that software science education, which is massively influenced by or based upon math, is counterproductive for understanding the results of empirical studies: when you are familiar with proofs by contradiction or with counter examples that disproof a general statements, it is hard to understand why a single case in an empirical discipline does not destroy a whole theory. When you are used to counter examples, it is hard to understand why a single person, who suffers from covid-10 for a second time, does not automatically falsify an immunity theory.

Summary

There is a lot that can be learned from the corona crisis. Software science can learn a lot from the corona crisis. We as software engineers or software scientists should not just read newspapers today and pretend that the process of knowledge gathering for corona is completely different to what needs to be done in our field. 

We should demand numbers. We need to provide such numbers. Our lines of argumentation should rely on numbers. And we need to accept the impurity of numbers - and use education as a weapon against wild speculations and naive scepticism in our field.



Wednesday, September 9, 2020

How much distrust do we need, how much trust can we afford in software science?

How much distrust do we need, how much trust can we afford in software science?

While summarizing results of identifier studies for a magazine I had to make a decision: I had to decide how much I trust in some experiments. In other words: how much distrust is needed when reading papers, reports, etc.? For example, if someone just wrote "I did an experiment and technique A turned out better then technique B", should I just take these works for granted and assume this is some evidence?

Example: Shneiderman's Experiment on Identifiers

The problem happened to me when I tried to summarize identifier studies. One of the earlier studies on identifiers was mentioned by Shneiderman and Mayer and I asked myself whether I should take the available, very short description of the experiment as a form of evidence into account: 
"Two other experiments, carried out by Ken Yasukawa and Don McKay, sought to measure the effect of commenting and mnemonic variable names on program comprehension in short, 20-50 statement FORTRAN programs. The subjects were first- and second-year computer science students. The programs using comments (28 subjects received the noncommented version, 31 the commented) and the programs using meaningful variable names (29 subjects received the mnemonic form, 26 the nonmnemonic) were statistically significantly easier to comprehend as measured by multiple choice questions." [1, p. 231]
That's it. Almost nothing more is said about the experiment in the given source. I could have said "well, the paper appeared in a peer-reviewed journal, hence this is evidence", but I did not feel that way. The problem was not only, that concrete numbers were missing in the description. The problem was also that I did not had a precise idea what was done in the experiment.

I wanted to give an impression of what was done in the experiment and then report means (and differences in means) and in case there are interaction effects, I wanted to report them as well (not in terms of statistical numbers, but rather in terms of text).  I know, means are problematic, but my goal was to summarize results for a magazine - the audience should not be bothered with p-values, etc. But I think it was also important to give people an idea what exactly participants did in the experiment and  how conclusions were drawn from it. I did not know what programs or how many of them were given to the subjects, how exactly the different treatments looked like, etc.

Ok, I was not satisfied with the description. But I also did not want to make it too easy for me and just say "I should ignore the text", because in the end the description still came from a peer-reviewed journal. So, I tried to find more about the experiment. I was digging at Ben Shneiderman's webpage, but was not successful. But in his book Software Psychology [2] I found some more text: 
"One of our experiments, performed by Don McKay, was similar to Newstead's, but the program did not contain comments. Four different FORTRAN programs were ranked by difficulty by experienced programmers. The programs were presented to novices in mnemonic (IDVSR, ISUM, COEF) or nonmnemonic (I1, I2, I3) forms with a comprehension quiz. The mnemonic groups performed signifianctly better (5 percent level) than the nonmnemonik groups for all four programs." [2, pp. 70-71]
After this paragraph, the book contains a figure that illustrates a difference in "mean comprehension scores" for four different programs. The score between the programs seems to vary from 3 to 5 for mnemonic, respectively 2 to 3.5 for non-mnemonic variables.

The second citation in combination with the figure gave me some more trust that the experiment revealed something. But it still puzzled me what the programs looked like that were given to the subjects. I also wanted to see more in more detail what variable names were used. But I also wanted to know what questions were given to the participants. And the first citation mentions multiple choice questions that were used (the second just speaks about a quiz). How many alternative answers had the subjects? How much time did they have for reading the code? What were the raw measurements, what the means, the confidence intervals? I just had the figure which does not show confidence intervals nor do they give a precise understanding of what the mean is. And finally: how was the data analyzed and what were the precise results? Was just a repeated measures ANOVA used? What about the second factor (programs)? Were there interaction effects?

Finally, I spent a lot of time on the experiment (mainly for searching for a more detailled experiment descriptions and for comparing both descriptions, whether they do match).  And I finally decided for myself, that I should not take the experiment as a form of evidence into account. In the text for the magazine, I just wrote that "there was once an experiment which is today rather historically intersting, but inappropriate as a form of evidence".

I felt bad. 

I had the feeling that I did not give Ben Shneiderman the appropriate credit for his efforts. But I really felt that it is my duty to be much more sceptical with his text - despite the fact than Ben Shneiderman is one of the leading experimenters in software science.

On Scepticism, Evidence and Trust

As a scientist (well, in fact as an educated person), you should not trust too much, you should not just believe someone (no matter who he is) and you should not stick with your own fantasy or your own personal and subjective impressions. I.e. when you are confronted with a statement, the following should hold: A statement ...
  • ... is not more important, because it follows a current hype,
  • ... does not become true just because the author of such a statement is an expert,
  • ... is not more valid, because it was articulated by an authority.
  • ....is not more important, just because you believe in it or the statements comes from you.
This kind of argumentation is far from being new. For example, Karl Popper wrote:
"Thus I may be utterly convinced of the truth of a statement; certain of the evidence of my perceptions; overwhelmed by the intensity of my experience: every doubt may seem to me absurd. But does this afford the slightest reason for science to accept my statement? Can any statement be justified by the fact that K. R. P. is utterly convinced of its truth? The answer is, ‘No’” [3, p. 24]
Of course, you cannot endlessly play this scepticicm game. You cannot ignore everything on this planet and just say that you feel sceptical about it. In the very end, you need to take some evidence into account. This evidence might be damn strong or just weak (and weak evidence does not mean that you heard an anecdote somewhere). 

But even strong evidence implies trust to a certain extent. You must trust that a study was executed, you must trust that the resulting numbers were measured, you must trust that not further numbers were measured that were withheld by the authors, you must trust in the validity of the analysis and you must trust in the seriousness of the interpretation. You must trust that the goal of the study's author was to find out something. 

Unfortunately, whenever some kind of trust is required, it is a door opener for fraud. Trust can be exploited. People can intentionally lie when trust is required.

Reporting Experiments - Setup, Execution, and Analysis Protocol

Before speaking about the problem with trust, let's take a look what can be known about an experiment.

In an ideal world, there is a guarantee that a study was executed, that this study follows a well-defined protocol, and that the study was analyzed in a way that matches the study's design. Such protocol consists of three parts: the setup protocol  (which describes what and how something should be done when it is replicated), the execution protocol (which describes the special circumstances in which the experiment was actually executed) and the analysis protocol (that gives the results of the study in statistical terms).

The setup defines the subjects that are permitted to participate (such as "professional software developer with skills X and Y"), the dependent (such as "reaction time") and independent variables (such as "programming language") and the hypotheses that are tested. Furthermore, the protocol contains the experiment layout (such as "AB test", etc.),  the measurements techniques (such as "reaction time measurement with stop watch") and the different treatments given to the subjects (such as Java 1.5, Squeak 5.0). Furthermore, it describes how and under what circumstances the different treatments are given to the subjects (such as the programming tasks given to the subjects, the task descriptions, the used IDE, etc.). In case the measurement techniques require some aparatus (such as some software used for measurements), this is also contained in the setup protocol. And in case, the aparatus cannot be delivered as part of the protocol, a precise description of the apartus is given.

The execution protocol describes the selection process for the subjects, the subjects that were finally tested, and the specific conditions under which they were tested. These special conditions could be the time interval in which participants were tested, the location where the test was executed, the machines used in the experiment, the concrete IDE (incl. version), etc. And finally the execution protocoll contains the raw data.

The analysis protocol describes how the possible effect of the independent variables on the dependent variables is determined. Since probably some statistics software is used for the analysis, this software is mentioned as well. The analysis protocol describes the results of the experiment in terms of statistical values. For the statistical values, corresponding reporting styles should be used such as APA (although this is very uncommon in software science). Each test comes with a measurement for the evidence (aka. p-value) and effect size (such as Cohen's d, eta squares, or just the means and the differences in means as well - latter ones are no effect sizes, but it is often more useful to have measurements that mean something to the readers instead of abstract things such as Cohen's d that most reader won't be familiar with).

On double-checking results

The information above is required because it permits readers to double check the experiment. It can be checked, whether the layout followed a standard-layout, whether the measurement technique is state-of-the-art, whether the tasks given to the participants were appropriate and whether the analysis follows the experimental design. And in case the reader doubts that the statistical results are right, he can recompute them. It even permits him to apply alternative statistical procedures.

In a real scientific world, there would not be the need for all this, because if an experiment was published in a peer-reviewed journal, you can trust that the reviewers did all this for you. Of course, this is no 100% guarantee. Even in disciplines such as medicine that have a very high research standard, studies are retracted in journals (see for example a recent case with a COVID-19 study). But the situation in software science is different. 

Taking the terribly low number of experiments in software science into account, there is good reason to doubt that an average reviewer in software science is able to do a serious review (just because quite few people are familiar with experimental designs and analyses). As a consequence, we cannot assume that a reviewer double-checked an experiment. Hence, it makes sense today not to trust on published experiments in software science, but to double check them. It does not necessarily mean that that authors intentionally lied. They might have just done some errors. Not intentionally, but just accidentally.

Unfortunately, we run here into a problem: double-checking costs time. Even if someone is well-trained in experimental analyses, it takes time to do the recomputation from the raw data (in case the data is available). But stats are just part of the game. There is the need to check whether the data collection followed the procotols, etc. But we cannot double-check everything. At a certain point we have to stop and say "I just have to trust". But we should make explicit on what we need to trust and what was actually double-checked. And we should give readers a fair chance to decide on his own, what he should trust in and what not.

Why not Making Chains of Trust more Explicit?

But what about the average developer who is interested in what evidence actually exists in software science? He is probably not trained enough to double check experimental results. But that implies that the developer will not take any evidence into account that exists in software science. And that implies that the results of software science will be in vain. This should not be the consequence, otherwise our discipline will never get out of this situation where countless statements without any evaluation exist.

Hence, it makes sense to give developers all essential information about an experiment but also the information about whom and what needs to be trusted. I.e. we should provide developers information such as "I, Stefan Hanenberg, recomputed the results of the experiment X and the results of the analysis are Y. I.e. if you cannot do your analysis on your own, you need to trust me that the computation of the analysis is correct". 

Probably it makes sense to make even more information available such as "I got the measurements and I repeated the analysis, but I was not able to access the tasks given to the subjects. I.e. I only confirm that the results of the experiment match the given data, but I cannot confirm that the data followed an appropriate experiment protocol, hence we need to trust the author of the experiment about that".

And maybe, it makes even sense to make some ratings about the resulting chains of trust. An experiment result such as "we need to trust the author or the experiment" seems to be less trustworthy than "we need to trust that a valid setup protocol was followed, but the results match the given raw data" which is less trustworthy than "we received everything about the experiment and we confirm that the experiment followed an appropriate design, was executed in an appropriate way and the reported results match the raw data". And the best case would be probably: "we received everything needed from the experiment, the results are as described from the author and the experiment was executed by others and they received comparable results".

Why Could that Help?

Our discipline suffers from the problem that a number of statements ("object-orientation is good", "functional programming is good", "UML improves understandability, etc.") are hardly or badly evaluated. But even if there are experiments available, we should make explicit what parts of the experiments are trustworthy and what parts are not - because in the very end, we want to rely on strongly trustworthy results and not just on "we trust some single person".

By making explicit what results in our field do not just depend on our trust in the authors, we make explicit where we could or need to improve our discipline.

References


  1. Ben Shneiderman, Richard Mayer. Syntactic/semantic interactions in programmer behavior: A model and experimental results. International Journal of Computer and Information Sciences 8, 219–238 (1979). https://doi.org/10.1007/BF00977789
  2. Ben Shneiderman. Software psychology: Human factors in computer and information systems, Winthrop Publishers, 1980

  3. Karl Raimund Popper. The Logic of Scientific Discovery. Routledge, 2002. 1st English Edition:1959.

Tuesday, May 19, 2020

Reporting Standards in Software Science Desperately Needed

Reporting Standards in Software Science Desperately Needed

  
If we are really interested in achieving something in software science, there is a need for reporting standards. I really mean this statement. But just recently I made the experience how urgently needed such standards are: I summarized an experiment and became aware how much time it took to extract relevant information from it. Some information was missing, some was confusing, etc. In case a standard such as CONSORT whould have been applied, it probably would have cost me minutes to summarize an experiment - instead of many, many hours where I finally even needed to contact the author, because some information was missing.
  

Recent Experiences While Summarizing Research Results

I recently summarized research results from experiments. The goal was relatively simple: Just collect the results from some studies and summarize them in a way that an ordinary software developer is able to understand them. I think such work is needed, because for example the study by Devambu et al. has shown that most developers judge the validity of claims in software construction based on their personal experience and not because of independent studies [1]. But taking into account that experience is limited and that subjective experiences are quite error-prone, it makes sense to give developers information about studies that exist and that give evidence for some claims. And what's even more important: Give developers studies that contradict given claims.

The topic of my summary was "identifizers", i.e. I wanted to summarize studies who checked what the influence of identifiers on code reading or code understanding is. Yes, I know. No big deal. Everyone of us knows how important the choice of good identifiers is. But I really wanted to know what was actually measured by researchers. And we should know something about the effect sizes.

Most of the studies were done or at least initiated by Dave Binkley. I read most of his papers already in the past and since I am well-trained in reading studies I assumed that it is no big deal to give a quick summary of some of them. And there was another reason why I focussed on his papers: From my experience and in my opinion his studies are well-conducted and I trust in the validity of the results, i.e. I trust that the numbers were collected in a way as described in the papers, I trust that the analyses of the data and I trust that the writings do not try to over-sell results: I think his research has to goal to find answers. His papers are not written for the sake of writing papers, but for the sake to improving the knowledge in our field.
 

My goal for the summary

More precisely, I wanted to summarize papers in a way that gives a 1-2 sentence description of the experimental design, another 1-2 sentences about the dependent and independent variables and a few sentences about the main results. And maybe some more sentences about what can be learned from the study. The goal was not to bother readers with stuff that is needed for scientific writings. I.e. I wanted to skip information about whether the experiment followed a crossover design, whether e.g. a latin-square was used or what statistical procedure was applied. 

Actually, I think it is necessary if readers who are not too deep in scientific writings get results in an understandable way. I.e. if an AB test has been applied, I think it makes sense not to write about statistical power, p-values, confidence intervals or effect sizes, but just to write that "a differences was detected" (in case a significant result was achieved) and then to report means and mean differences. And in case multiple factors are tested, my goal was not to write about interaction effects, etc. but just to explain interactions in a way that an average person can get the meaning quickly. 
 

Shouldn't someone else do the job?

Quickly is the point here. If we want developers to understand results of studies, they have to be communicated efficiently. And the typical research paper at a conference or in a journal does not seem to have the goal to communicate results efficiently. Authors of conference papers are given a certain number of pages they can fill. And authors are actually forced to fill this number of pages. If for example a conference such as the International Conference on Software Engineering (ICSE) has a page limit of 12 pages, you will hardly find a paper at that conference that does not have 12 pages. This has something to do with the review process (which should not be discussed here, although there is an urgent need to discuss it). Ok, so you want to communicate scientific results for a broader audience. But how?

Actually, scientific journalism in other disciplines does this job: people who are trained in writing (for a popular market) summarize results in a way that people are able to understand them. This is important, because people should be informed about what knowledge exists - especially taking into account that people pay for the generation of this knowledge (because a lot of scientific work is paid from tax money). But for software science this kind of journalism does not exist. Yes, there are a bunch of magazines that address technical things. You find books on new APIs or new technology that explain how to apply it. But this is something different. These writing explain how industrial products could be used. They do not explain what we actually know about them. It would be great if there would be people who summarize research results - but we currently have to live with the fact that this is actuall not done in our field.

So, back to the studies.
 

Giving a quick summary took damn long time

Again, I really love Dave's work. I think his studies are great. His writings are great. But it turned out that just writing a quick summary took much more time than expected. When I now explain what happened to me and why I had troubles to summarize the paper, this should not and must not be understood as a criticicm of Dave's work. Really not. Dave's work is definitively a shining example of how good science in our field should be. Dave's paper is just an example for what troubles people could have reading scientific papers. And I assume my papers suffer from the very same problems.

One of the papers I started with was Identifier length and limited programmer memory [2]. I remembered that this study compared 8 expressions with different lengths and that subjects were asked to write down a part of the expression. So, I wanted to write sentences such as:
 "The experiment gave A subjects B expressions to read for a time C (D subjects were removed for some reasons). Each expression consisted of E parts and the authors used the criterion F to distinguish between short and long expressions. After reading, a part from the expression was removed and subjects had to complete it. The average time for reading short expressions was T1 and for long expressions it was T2, so the (statistical significant) differences was T3, respectively it took people G percent more time to read the long compared to the short expressions."
I am aware that these sentences are quite a simplification of the results. Especially, I do not mention all independent variables and I do not mention the applied statistical method. By reporting the means people do not get an idea of the size of the confidence intervals, etc. But, again, the goal was to give a quick (but still informative) and not a complete overview. Why do I think that this kind of summary is informative? Well, I think it contains the most relevant information. The number of subjects gives an idea how large the experiment was (and people are mad about this idea of "being representative" - that's another point that needs to be dicussed, but not here, not now), the dropout rate gives an idea how much the data says about the relation between the "originally adressed sample" and the actual data used for the analysis. And the average times give people an idea how large such differences are. Yes, there are effect size metrics such as Cohen's D or eta square, but if someone does not know these things, such numbers would rather confuse him.
 

Sample size and dropout rate

Doing the first step (number of subjects) seemed relatively easy, because the number 158 is already mentioned in the abstract, so I directly started searching for the dropout rate. Suddenly it took some time to understand what exactly happened to the data. The paper does not have an explicit section such as "experiment execution" or something. But there is a section "Data preparation" where I found the following:
"[...] the data for a few subjects was removed. For example, one subject reported writing down each name. A second subject reported being a biology faculty member with little computer science training. Finally, the time spent viewing Screen 1 was examined. It was decided that responses with times shorter than 1.5 s should be removed because they gave the subject insufficient time to process the code. This affected 18 responses (1.4% of the 1264 responses). In addition, excessively large values were removed. This affected 6 responses (0.5%) each longer than 9 min." [2, p. 435]
Ok, but what was the actual data being used? The second sentence seems to describe that the data of a whole subject was removed. But what means "this affected 18 responses"? Does this mean that the data of 18 subjects was removed? Or just 18 answeres? And what about the other six? Does it mean that 24 answers, i.e. three subjects were removed? Or was each single response treated individually? I felt suddenly slightly reluctant to write down a sentence such as 158 subjects participated, because I was not able to find precisely what data was skipped. But, ok, I lived with the problem - and just reported that 158 subjects participated. Actually, this step alone took me quite a bit of time, because I reread the paper more than once because I assumed I missed some relevant information.
 

How large is the effect of expression length?

The main reason why I looked into the paper was, that I wanted to know whether expression length was a significant factor and in case it was, how large the effect was. The paper report on a significant average difference of 20.1 seconds between long and short expressions in reading time, i.e. longer expressions took longer. But how much longer did they take in comparison to short expressions? 

I started searching either for effect size measures or at least some descriptive numbers such as means or confidence intervals or something. I was really convinced that I must have missed it somewhere. So I re-read the paper over and over again - and did not find the number. The only things I had was the following:
"It was decided that responses with times shorter than 1.5 s should be removed because they gave the subject insufficient time to process the code. This affected 18 responses (1.4% of the 1264 responses). In addition, excessively large values were removed. This affected 6 responses (0.5%) each longer than 9 min." [2, p.435]
So, should I just report that the average reading time was between 1.5 and 9 min which would mean that 20.1 second is "between factor 14 and 4 %"? That does not sound meaningful. Again, searching just for this single number (that I finally did not get) took me some time. The same is true for a second variable: syllable. It is reported that each additional syllable costs the developer 1.8 seconds. But what does the first syllable cost?

In fact, I felt more uncomfortable with the variable syllable because there are multiple treatments of this variable and I would be much more interested in the precision of 1.8 seconds.
 

How exactly were the results of the study computed?

What puzzled me as well was the question, how the results were achieved: what statistical procedure was used? And what tool was used? The paper just says that linear mixed-effects regression models were used. Ok, but with what tool? And what exactly were the input variables for the regression models?

Going back to the question on the effect of the variable length, the paper says that "the initial model includes the explanatory variable Length" [2, p. 437]. Length? In a regression? The paper uses length, which is a  binary variable (the paper distinguishes between short and long), so in principle, it is just a simple AB-test or did I miss something? Or was bunch of variables added to the initial model and just length was the one that was significant?

Actually, it turned out that I had many, many more problems. And it took me quite a lot of time just to identify that some of these problems were real. I should mention that because I had so much trouble, I contacted Dave who sent me the raw data set within hours so I was able to analyze the data on my own in order to get the results from the experiment.
 

Why standards such as CONSORT are urgently needed in software science

Finally, I got the raw measurements and was able to recompute some numbers and everything was fine. But why did I feel something is really problematic?

Again, I am well-trained in reading studies. But if it took me hours to understand what was in the paper, I assume that it took many more hours for people who are not trained in these things. So, how can we even assume that someone will take studies into account if it takes many hours to read them? The study by  Devambu et al. indicates that we should blame developers for not knowing what is actually known in the field. But if understanding a single study takes many hours, it actually makes sense that people do not read them. Why? Because developers have more to do than just spending a whole day on reading a single paper. And in case essential information is finally missing, the whole days was spent in vain.

So, how come that essential information are hard to find in studies or are even missing? Again, I do not blame the authors of the here mentioned study for forgetting something. But the paper was published in a peer-reviewed journal. How is it possible that it passed the peer-reviewing process while some essential information is missing (again, we need to speak about the reviewer process at some point, but not here)? 

I am happy that the paper was published, because otherwise the whole body of knowledge in our field would be even less - and it is already inacceptable low (see the study by Ko et al. [3]). But what would have reduced the problem?

Here come research standards into the game. If our field would be disciplined enough to apply a relatively simple reporting standard such as CONSORT [4], things would be easier. Such standard implicitly contains a summary in each paper that permits you to find information quite fast. For the review process, it is a relatively easy thing to check, whether a paper fulfills the standard. I.e. authors can double-check whether the relevant information is contained and reviewers can do this double-checking as well. 

Applying such a standard would have another implication: if for example a conference would apply such a standard, many papers could be directly rejected because they do not fulfill the standard. The problem identified by Ko et al. (and there are in fact many, many more authors who documented that evidence is hardly gathered in our field) would vanish: scientific venues would publish just papers that follow the scientific rules. This would reduce the problem that readers are confronted with tons of papers whose content cannot be considered as part of our body of knowledge. 

Yes, there is the other problem which makes it hard to imagine that we finally get to the point that the software science literature would contain scientific relevant studies: people must be willing to execute (and publish) experiments which are able to conflict with their own position. But this is a different issue I discussed somewhere else.

Yes, research standards are urgently needed. At least reporting standards. Urgently.

References

  1. Devanbu, Zimmermann, Bird. Belief & evidence in empirical software engineering. In Proceedings of the 38th International Conference on Software Engineering, ICSE 2016, Austin, TX, USA, May 14-22, 2016, pages 108–119, 2016. [https://doi.org/10.1145/2884781.2884812]
     
  2. Binkley, Lawrie, Maex, Morrell, Identifier length and limited programmer memory, Science of Computer Programming 74 (2009) [https://doi.org/10.1016/j.scico.2009.02.006]
     
  3. Andrew J. Ko, Thomas D. Latoza, and Margaret M. Burnett. A practical guide to controlled experiments of software engineering tools with human participants. Empirical Software Engineering, 20(1):110–141, February 2015. [https://doi.org/10.1007/s10664-013-9279-3]
     
  4. The CONSORT Group, CONSORT 2010 Statement: updated guidelines for reporting parallel group randomised trials, 2010. [http://www.consort-statement.org/downloads/consort-statement]

Thursday, May 7, 2020

Before Doing Science in Software Construction Something Else is Needed: Critical Thinking

Before Doing Science in Software Construction Something Else is Needed: Critical Thinking


Why is Science Needed in Software Construction?

Software construction is a huge, multi-billion market where new technology appears almost every day (in case you doubt that software is a multi-billion market, just take a look at the 10 most valuable companies on this planet today). Such new technology comes with multiple claims and the most general one is, that the new technology makes software development easier and hence cheaper.

Taking the size of the market into account, there are good reasons to doubt whether all technology on the market exists for a good reason - beyond the reason that new technology increases the income of companies or consultants who propagate this technology. Taking the size of the software market into account, there are reasons to believe that a lot of technology exists although its promised benefit neither ever existed nor will ever exist.

In the very end, one has to accept that most claims associated with a certain technology are not the result of non-subjective studies. Instead, they are the result of subjective perceptions or impressions of people who either have strong faith in a new technology, who really hope that the technology improves something, or who just love a new technology ("faith, hope, and love are a developer’s dominant virtues" [1, p. 937]). Finally, some of these claims are just the result of marketing considerations: Claims that are made and spread just because they increase the probability of success for the new technology and not because they are true.

That non-subjective studies are rather rare exceptions in the field of software construction is a sad, but well-documented phenomenon. For example, Kaijanaho has shown that up to 2012 only 22 randomized controlled trials on programming language features with human participants were published [2, p. 133]. Another example is the paper by Ko et al. who analyzed the literature published at the four leading, scientific venues in our field. The authors came to the conclusion that "the number of experiments evaluating tool use has ranged from 2 to 9 studies per year in these four venues, for a total of only 44 controlled experiments with human participants over 10 years" [3, p. 137].

So, what's wrong with this situation? The problem is, that new technology causes costs. Costs for learning this technology, applying it, and maintaining software written in it. And there are additional, hidden costs. First, there are costs because new technology supersedes existing technology. Such existing software becomes often rewritten which means that investments done in past need to be repeated the future. And in case existing software is not newly written, there are additional costs for maintaining the old technology. And old technology causes larger costs because once a technology is no longer taught and no longer applied it becomes more expensive to maintain it simply because there are no longer people on the market who are able to master the old technology. An extreme example for this was the Y2K problem, whose costs were to a certain extent caused by forgetting the old technology COBOL.

But there is another, tragic problem. The problem is, that in case a new technology would appear that solves a number of problems we have today, such technology could not be identified. The claims associated with this new technology would just be lost among all the other claims that exist for today's technology or claims that will be associated with competitors.

We must not forget that the goal is not to find excuses to stick to old and inefficient technology. The goal is to make progress. But progress does not mean just to apply new stuff that appeared recently, but to apply technology that improves the field of software construction.

So, what we need are methods to separate good from bad technology. We need to separate knowledge from speculation and marketing claims. And we need to teach such methods to developers to give them the ability to separate knowledge from speculation. This does not mean that we need developers who execute studies. But we need developers who are able to read studies and who are able to identify trustworthy studies from bad ones. In the end, we want a discipline that relies on the knowledge of the field as a whole and not on speculations of individuals.

The Scientific Method

The alternative to subjective experiences and impressions is the application of the scientific method, which is actually the alternative to subjectivity and not just one alternative among others. This does not imply that the term scientific method describes a clear, never-changing and unique process of knowledge gathering. Instead, it is a collection of things that can be done, should be done or must be done. And this collection changes over time, because not only knowledge in a certain discipline changes because of the scientific method. The method changes as well.

It is not surprising that the scientific method is often critically discussed in the field of software construction which is more an expression of the immaturity of the field instead of the community's willingness to generate and gain non-subjective insights. Just to give an impression: even at international, academic conferences on software construction, there are discussions whether not the scientific method makes any sense at all. At such places, there are discussions about the need for control groups, the validity of statistical methods or the validity of experimental setups. All these discussions exist despite the fact that there are tons of literature available from other fields on these topics (which give very clear answers to these topics). One could argue that this immaturity just exists because the field is quite young. In fact, this statement can be easily rejected. In medicine, which is typically considered as one of the old fields, most of the experimental results that we accept todays as those ones that follow valid research methods, are just done in the last 30-40 year.

The fundamental part of the scientific method is, that there are people who are willing to test the validity of hypotheses. This implies that they are willing to accept results although they conflict their own, personal and subjective impressions or attitudes. But this means that they not only accept their own experimental results, but they also accept results from others. Although this seems quite natural, it has one important implication. It means that people established some common agreement what a valid research result is and what not.

Scientific Standards

Let's discuss the very general idea of research standards via an example. Let's assume there are two programming techniques A and B and one would like to test the hypothesis that it takes less time to solve a given problem using technique A than it takes using technique B. So one person tests 20 people, 10 solve a given problem using A, 10 solve it with B. Then the time for both groups are measured and then compared. This is a standard AB-test where not only the experimental setup (randomization of participants, etc.) but also the analysis for the data (t-test, respectively U-test) is well-known since decades. But the general question is, whether or not one should take the results of the experiment into account as a valid result.

It turns out that especially in software construction people complain a lot about such a standard approach. And in case technique A is more  efficient than B, a larger number of people who prefer technique B will find reasons either to ignore the result or to discredit the experiment. Actually, there are quite plausible arguments against the experiment and the most general one is the problem of generalizability: one either doubts that the number of subjects is "representative" in order to draw any conclusion from the experiment. Another doubt is, whether the given programming problems represent "something that can be found in the field" or whether the problems are any "general programming problems at all".

We should not be too ignorant to reject such objections directly, because there is some truth in them. But we should also not be too open minded to take such objections too serious, because of the following reasons: there is no experiment in the world that is able to solve the problem that underlies these objections. No matter how many developers are used as subjects in the experiment, one can always argue that the number is too low. And no matter on how many programming problems the techniques are tested, there are other programming problems on this planet that were not used in the experiment. 

In order to overcome such situation there is a need to have some common understanding of the applied methods: there is the need for community agreements. If people agree on how experimental results are to be gathered, there is no need to doubt in results that come from experiments that follow such agreements. In other disciplines, the problem was identified as well (some longer time ago) and corresponding scientific standards were created. Examples for such standard are the CONSORT standard in medicine [4] (which mainly addresses the way how experiments are to be reported) and the WWC-standard that is used in education [5] (which not only covers the way how experiments are to be executed, but which also handles the process of how experiments should be reviewed).

The need for such community agreements is obvious and we argued already in 2015 that such community agreements are necessary in software construction as well [6]. Today we find movements towards such standards. An example for this is the Dagstuhl seminar "Toward Scientific Evidence Standards in Empirical Computer Science" that takes place in January 2021 [7].

On the Selection of Desired, and the Ignorance of Undesired Results

Such movements are good and necessary. However we should ask ourselves, whether the field of software construction is ready for such standards. Because the introduction of research standards entails some serious risks that should be taken into account. But before discussing these risks, I would like to start with some examples.

In the last years, one situation occured over and over again to me. A collegue contacted me and asked, whether there is one experiment available that supports a certain claim. The collegue's motivation is typically that she or he tries to find a way to argue about the need for some new technology and from her/his perspective this motivation would be stronger if there would be some matching experimental results. At that point I usually start a conversation and ask what if there are experimental results that show the opposite. At that point I usually get the answer that such experiments would be interesting, but wouldn't help in the given situation. In order words: an experimental result (in case it exists) is ignored in case it contradicts a personal intention.

Something else happened to me in the last years which is related to an experiment I published in 2010; an experiment that did not show a difference between static and dynamic type systems [8]. Today it seems quite clear that the experiment had problems and it would have been better if the experiment was never published. In the meantime, other experiments showed the positive effects of static type systems (such as for example [9]): Taking the sum of experiments into account, the question of whether or not a static type system helps developers can be considered answered (so far). But what happened is that people, to whom it is helpful that that no difference between static and dynamic type systems was detected, have the tendency to refer only to the first study in 2010 but to later ones. For example, Gao, Bird and Barr explain relatively detailed the results of the 2010 paper, but do not mention the latter one [10]. Again, it seems as if only those results are taken into account that match a given intention - and results that contradict such intention are ignored.

Finally, another situation occured more than once or twice. A collegue created some new technology and asked me for advice in order to construct an experiment that reveals the benfit of the new technology. After some discussions (which often last for hours) we typically come to the point that the collegue is really convinced about the benefit of the technology in a certain situation, but thinks that in a different situation the technology could be even harmful. Often, this collegue is in the situation that a PhD needs to be finished and "just the last chapter - the evaluation" needs to be done. And what happens next is often that an experiment is created that just concentrates on the probable positive aspects of the new technology - the (possible) negative aspects are not tested.

The commonality of these examples is, that people today have the tendency to select only those results that do match their own perspectives or attitudes. In other words: even if strong empirical evidence, i.e. a number of experimental results, exists for a given claim, people still have the tendency to search for singular results that contradict such claim if people do not share this claim.

This is comparable to people who advocate homeopathy and select those rare experiments where homepathy showed a positive effect - and ignore the overwhelming evidence we have about homeopathy.

The Required and Currently Missing Foundation is Critical Thinking

Probably there is a reason for such a behavior and I assume that such reason has something to do with people's attitude in our field. In our education, from the very beginning people are involved in ideological warfares: procedural versus functional versus object-oriented programming, Eclipse versus IntelliJ, GIT versus Mercurial, JavaScript versus TypeScript, Angular versus React, etc. "Chosing a side" seems to play an essential role in software construction. And it actually makes sense to a certain extent. If I master a technology, it is beneficial for me if this technology becomes the leading technology in the field. If I master a technology that no one uses and that no one is interested in, my technological skills are not and maybe will never be beneficial to me. Consequently, people advocate the technology they use and they try to find reasons why this technology should be used by others as well. And in order to achieve this, all kinds of arguments will be applied and it does not matter whether an argument is actually valid as long as supports my intentions. This behavior becomes stronger as soon as people start developing their own technology. If someone writes as part of this PhD a programming language, there seems to be the tendency to defend this language. 

The idea of defending a self-created technology or to defend a technology just because one is able to master it seems quite natural. But actually, this is probably the core of the problem. We need to communicate from the very beginning that the goal is to make progress. And that progress means that we are willing to identify problems. And in case there is strong evidence that a certain technology has serious problems, we must be open minded enough to take alternative technologies into account. We must be able to accept and apply critical thinking.

Of course, this must not lead to the situation that people switch technology directly after some rumours appear about some better technology - in fact, this would be closer to the situation we have today where a large number of people accept new technology for the sake of being new. Just to throw everything away in order to apply something new is closer to actionism than to critical thinking. Critical thinking also does not mean that we find ad hoc arguments against some technology. Critical thinking must not mean that we encourage wild speculations. It just means that we are willing to accept different arguments. Critical thinking means that we are willing to collect and accept pros and cons. It means that we are willing to give up our own position.

This willingness is the very foundation we need in our field. Because it does not matter if we define research standards in our field and enforce people to follow such research standards as long as people are not willing to accept results that conflict their own positions. Otherwise, research results will be either just ignored or people generate and publish only those results are match their own attitudes.

Once we have achieved this kind of critical thinking and once we are able to give this idea to students, we can go the next step towards evidence in order to give people the ability to differ between strong, weak and senseless arguments, arguments that are backed up by evidence and those ones that are not. Then, we have researchers who are willing to define experiments whose results might contradict the experimenters' positions. This would be the moment where science could start in our field. This would be the moment when we are ready to apply the scientific method.

References

  1. Stefan Hanenberg, Faith, Hope, and Love: An essay on software science’s neglect of human factors, OOPSLA '10: Proceedings of the ACM international conference on Object oriented programming systems languages and applications, October 2010, pp. 933–946. [https://doi.org/10.1145/1932682.1869536]
  2. Antti-Juhani Kaijanaho, Evidence-Based Programming Language Design A Philosophical and Methodological Exploration, PhD-Thesis, Faculty of Information Technology, University of Jyväskylä, 2015. [https://jyx.jyu.fi/handle/123456789/47698]
  3. Andrew J. Ko, Thomas D. Latoza, and Margaret M. Burnett. A practical guide to controlled experiments of software engineering tools with human participants. Empirical Software Engineering, 20(1):110–141, February 2015. [https://doi.org/10.1007/s10664-013-9279-3]
  4. The CONSORT Group, CONSORT 2010 Statement: updated guidelines for reporting parallel group randomised trials, 2010. [http://www.consort-statement.org/downloads/consort-statement]
  5. U.S. Department of Education’s Institute of Education Sciences (IES), What Works Clearinghouse Standards Handbook Version 4.1, January 2020. [https://ies.ed.gov/ncee/wwc/Docs/referenceresources/WWC-Standards-Handbook-v4-1-508.pdf]
  6. Stefan Hanenberg, Andi Stefik, On the need to define community agreements for controlled experiments with human subjects: a discussion paper, Proceedings of the 6th Workshop on Evaluation and Usability of Programming Languages and Tools, October 2015, pp. 61–67. [https://doi.org/10.1145/2846680.2846692]
  7. Brett A. Becker, Christopher D. Hundhausen, Ciera Jaspan, Andreas Stefik, Thomas Zimmermann (organizers), Toward Scientific Evidence Standards in Empirical Computer Science, Dagstuhl Seminar, 2021 (to appear) [https://www.dagstuhl.de/en/program/calendar/semhp/?semnr=21041]
  8. Stefan Hanenberg, An experiment about static and dynamic type systems: doubts about the positive impact of static type systems on development time, Proceedings of the ACM International Conference on Object Oriented Programming Systems Languages and Applications, Reno/Tahoe, Nevada, USA, ACM, 2010, pp. 22–35. [https://doi.org/10.1145/1932682.1869462]
  9. Stefan Endrikat, Stefan Hanenberg, Romain  Robbes, Andreas Stefik, How do API documentation and static typing affect API usability?, Proceedings of the 36th International Conference on Software Engineering, May 2014, pp. 632–642. [https://doi.org/10.1145/2568225.2568299]
  10. Zheng Gao, Christian Bird, Earl T. Barr, To Type or Not to Type: : Quantifying Detectable Bugs in JavaScript Proceedings of the 39th International Conference on Software Engineering, 2017, pp. 758-769. [https://doi.org/10.1109/ICSE.2017.75]

Sunday, May 24, 2015

Rejection Letter For the Paper 'A Falsification of the Aristotelian Theory of the Free Fall and an Alternative Theory'

Rejection Letter For the Paper 'A Falsification of the Aristotelian Theory of the Free Fall and an Alternative Theory'

(a pdf of this text can be found here)

Cynicism is definitivly a problematic thing if being using in academic argumentations. Well, the text below is actually not meant to be cynical. It only tries to describe a non-exceptional situation in the current scientific selection process: the rejection of a paper that is based on measurements. The kind of argumentation in the reviews is not invented. It is an extraction of multiple reviews I got over the last few years. The text gets its cynical flavor because it refers to a work whose result and impact is known to us: the (hypothetical) experiment by Galilei on the free fall.
It should be clear that I do not and will not (not even try to) draw any parallels between my work or any other's work with the work by Galilei.
....have fun reading the text...
Stefan Hanenberg
stefan.hanenberg@gmail.com
version 0.1, Essen, 2015-05-23

Dear Mr. Galilei, 

thank you for submitting the paper 'A Falsification of the Aristotelian Theory of the Free Fall and an Alternative Theory' to the special issue 'Physics and Stuff' of our journal 'Software Technology Usage in Productive Industrial Development'.

Unfortunately, I have to tell you that the paper is rejected based on the common proposal of all reviewers.

The reviews are attached to this notification in order to help you improving your paper and improving your future work.

With kinds regards,The Editor

First Review


Overall merit: 1. Reject 
Reviewer expertise: 1. Expert


Summary


The paper gives a short introduction of the aristotelian understanding of the free fall and discusses potential problems with this theory. Then, the author runs a small controlled experiment where two cannonballs are dropped from some rather peculiar tower. From the experiment's measurements the author concludes that the aristotelean theory of the free fall must be wrong. In a second experiment, the author dropped even more cannonballs with different weights and proposes (based on the measurements) a different relatively trivial model that is from the author's perspective a better theory of the free fall.

Review

I am quite positive about empirical studies and results in general. Still, I cannot hide that I am not convinced at all about the here proposed experiment and the conclusions drawn by the author.

First, the description of the aristotilean theory and its relevance is hardly described. In fact, the author just says that it is a theory that exists since some centuries and plays a major role in our current understanding of the real world. However, it is unclear why this theory should be relevant at all, or why a change in the theory should be necessary (taking into account that this theory exists since centuries!). Hence, neither the background of the theory nor the possible impact of changes to that theory are clearly described which makes it hard for readers to understand why it should be even interesting to read this paper et al.

Second, the experiment is far from being convincing for a large number of reasons.

1) The author uses cannonballs of different weights and drops them from a tower. As the author mentions the tower is not an ordinary one but seems to have some peculiarities (why do such towers even exist in Italy?). As a consequence, it is quite plausible that the experiment results (whatever they are) are highly influenced by the choice of the tower itself and not by the theory being tested. Hence, the reviewer urgently asks the author to rerun a different experiment with a different tower. 

2) Next, the author choses cannonballs. As a consequence, he completely ignores that cannonballs serve a special purpose. Cannonballs are explicitely designed in a way that their behavior with respect to being shot or being dropped is very similar - no matter what their weight is. Such statements by cannonball designers are even completely ignored in the related work section. Definitively, well known papers such as "Why I think writing software for cannonball designers is a better option than having no job at all" by Fubar et al. '63 which is a fundamental and ground-breaking paper for the whole cannonball industry must be mentioned. Hence, it is clear that just because of the used subjects in the experiment it was not even possible to show anything else than just similar results - the experiment does not falsify any theory but just gives another indicator for the maturity of our cannonball industry.

3) The measurement process is not described. While the author mentions that the cannonballs had different weights (in the first experiment 1 kg vs. 10 kg) it is completely unclear how the measurements are performed - the precision of the measurements are completely unclear and the author does not discuss how the measurements were performed. Even worse, the author just describes the tower in terms of its height (without mentioning any other measurement) and - again the 56 meters are given without describing any other measurement. Taking into account that the tower seems to have some problems with its fundament, the height measurement is obviously not enough to describe the essential parts of the experiment. The time measurement is hardly described: the author just describes on page 3 that the time measurement started from the moment when a cannonball was dropped until the cannonball hit the floor and that the author used some clock that shows even miliseconds*10. We need to take into account that even the definition of a second changed over time (mean solar day second vs. period of the Earth's orbit around the Sun vs. atomic clocks). From that we can conclude that none of the time measurements is trustworthy -- as the author should have noticed, there are small deviations between the different measurements. This shows clearly that the measurements are not trustworthy at all. Hence, we must not conclude anything from the measurements.

4) The sample size is much to small for any serious study. Taking into account that for the first experiment - only 2 kinds of cannonballs have been used (again, all measurements were performed from the same tower) only 10 time measurements were collected it is completely impossible to generalize from the measurements to anything else. However, such kind of generalization is done in the second experiment. In the second experiment the author concludes from 10 different - yes, again cannonballs - and 20 different heights - again without giving a precise description of the measurements - that all measurements can be described with a formula t(h)=sqrt((2*h)/9.8). Again, such a formula cannot be derived from the measurements (for the reasons explained below).

5) The resulting formula appears completely unmotivated -- where does it come from and why should the free fall described in that way?

6) The analysis of the experiment is somehow obsure, especially when taking into account that a precise number is the result of the formula: not a single measurement matches the expected result from the formula! The author should have noticed that not a single measurement fulfills the formula. Instead of mentioning this obvious thing, the author tries to rescue his experiment with some statistical tests. The applied test (some so-called significant test) is not explained in detail and the rather mysterious results of this test are not explained. The author should have noticed that even an obvious test (the arithmetic mean of the results) differs from the formula's results.

7) The external validity of the experiment is in fact zero. Only cannonballs have been dropped in order to argue that the aristotelian theory does not work. I strongly advice to repeat this experiment with multiple other objects to be dropped from the tower. It seems obvious to drop additional things such as water, sand, or even complete ships in order to increase the experiment's external validity.

8) The related work section is far from being complete. Again, fundamental works about cannonball constructions are missing, no work is mentioned about the used tower, and not even different works on time measurements have been cited.

Hence, I conclude from the review above that the paper has to be rejected: the motivation is unclear, its relevance is unclear, the measurements are unclear and as a consequence, no conclusions must be drawn from these measurements. Additionally, the external validity of the experiment is not given. Although the proposed alternative model for the free fall is interesting the paper does not give any valid trust in the validity of the model.

Second Review


Overall merit: 1. Reject
Reviewer expertise: 1. Expert

Summary


The paper describes what happens when cannonballs are dropped from a tower. The author performs a number of measurements and states that these measurements conflict with an older theory of the free fall. Finally, the author describes his personal formula for the free fall which does not seem to be consistent with the older theory.

Review

While the general idea of the paper is interesting and the author's conclusions are quite innovative, the paper has a restricted perspective: the whole argumentation is based on quantitative measurements. Probably the most important information is missing in the paper: The design process of the experiment is completely unclear. Why was this extraordinary tower used for the experiment? Why were cannonballs used? Why was the measurement based on time? What were the reasons to come up with the final formula proposed in the paper? How can the formula be explained?

According to this, the general comment to the paper is that the paper lacks of any qualitative analysis that is relevant to the studied topic.

a) What additional observations were made in addition to time, height and weight? The pure use of quantitative data is a too restricted perspective on any aspect of daily life. The chosen cannonballs are only explained in terms of weight: it is clear that additional characteristics of cannonball are essential, too (What are they made off? Who was the producer? Has their functionality been tested before? How can the surface of the cannonballs be described? Were they comparable?). With respect to height, it is unclear how the tower can be described best. The author mentions (in addition to its height) only one special characteristic of the tower in one single sentence. All other aspects of the tower are completely ignored. With respect to the time measurement, it is unclear why the authors tried to measure time in such a complex way and not only asked people whether they saw differences in the free falls of the cannonballs. Additionally, the author not even tries to describe the different ways how the cannonballs fell down (although the measurements do show differences!). Hence, the most essential information -- the different ways how the cannonballs fell down -- are missing in the paper. Because of the resulting missing qualitative analysis it is not possible to find explanations for the differences in the way how cannonballs fall down from a tower.

b) The chosen experimental design reveals some obvious weaknesses. In the first experiment, the author uses two cannonballs in multiple measurements. Hence, the indiviual influences of a single cannonball is very high and it cannot be expected that the resulting measurements imply anything meaningful: it is well-known that AB experiments have the problem of unbalanced groups and this effect is even stronger in the here proposed experiment because of the use of the same cannonballs. In order to get rid of the problem, it is more desirable to have multiple different things to be dropped from the tower. Again, qualitative studies could help in order to find out what kind of different things could be dropped from the tower. An additional qualitative study could show what additional items could have been carried on top of the tower.

c) Because of the missing qualitative data, it is impossible to replicate the experiment. In case someone wants to replicate the experiment, it is necessary to understand what kind of tower could be used in the experiment and what kind of cannonballs could be dropped. Because of the special characteristics of the tower it seems even impossible to replicate the experiment, because it ia rather unlikely that such towers can be found somewhere else in the western hemisphere.

While I think that the provided quantitative data has some value it is still necessary to provide additional qualitative data and to run an additional qualitative analysis.

Minor comment:

The paper needs proofreading by a native english speaker.

Third Review

Overall merit: 1. Reject
Reviewer expertise: 1. Expert

Summary: 

The author describes an experiment that tries to falsify an existing theory: the aristotelian theory of the free fall. Based on two experiments, the author comes to the conclusion that the theory must be wrong and the author proposes an alternative theory.

Review:

The starting point of the paper is quite unusual. While most of the works that can be found today in physics address real world problems such as alternative sources for energy or the construction of large machineries somewhere in Europe (Switzerland), the paper addresses a more basic topic: the free fall. But it is unclear how this could be able to provide any usable or useful insights that could be applied today. Hence, the relevance of the work is unclear.

The line of reasoning in the paper is problematic. The authors try to falsify a theory (free fall) by some measurements in an AB experiment that actually do not differ: the author does not measure a difference between group A and group B. However, the literature explicitely says that measuring no differences is no indicator that there are no differences. It could only means that the experiment is problematic. Hence, the p-value of .99999 cannot be interpreted as no difference between A and B. The same is true for the second experiment where multiple measurements are compared (minor comment: the author forgot that a correction is needed because of the cummulated alpha error).

The author ignores completely that the theory of the free fall is no longer relevant - take into account that our industry is in the meantime able to construct things such as airplanes. Even more plausible comparisons are not being done. For example, birds do land with different speeds which directly contradicts almost everything that can be found in the paper. Hence, the proposed theory is not only a pure artificial one, it is even possible to find direct contradictions with it in reality (landing birds). Hence, the general idea of falsifying the aristotelian theory is not only irrelevant, it is wrong.

Minor comment: The author should have noticed that the proposed theory reveals results that are not measurable with the clock he used -- nor with any other clock that exists today (because all of them are not precise enough). As a consequence, it is clear that any arbitrary measurement inherently falsifies the proposed theory.


Thursday, November 13, 2014

10 Feet Steel Plate (Parabel on Software Construction)

->This document is originally a pdf <-

This is the a story about “situations in software construction” that the author frequently uses to argue about problems in software construction. The intention is to show by an analogy that the discipline of software construction has serious problems. Although analogies are often inappropriate to argue for or against something, the author has the personal feeling that for software construction the partially absurd and frustrating situation becomes even more obvious with the aid of analogies. 
Stefan Hanenberg
University of Duisburg-Essen, Germany
stefan.hanenberg@uni-due.de
version 0.1, Essen, 10/11/2014

The 10 Feet Steel Plate 

1 The Story

Vince wants to build a house. After speaking with his bank he is convinced that he has the financial resources to build a house. However, Vince is not a house builder, so it is hardly astonishing (after he was thinking for a while whether he should still try build the house on his own **1) that he asks an architect for help. Vince never built a house before. Hence, he has no experience in what can go wrong and how expensive all his different wishes for the house are. Vince feels slightly unsecure, because he is not able to judge whether the advice given by the architect is meaningful (because of Vince's absense of civil engineering capabilities).

The meeting with the architect works quite well. They speak about the number of rooms Vince requires and speak about the budget for the house. They speak about the number of bathrooms, the number of floors, the size of the living rooms and the size of the kitchen. They speak about the kind of stone that will be mainly used and the construction of the roof. Vince is absolutely aware that the more he wants, the more he finally has to pay and because of that he asks all the time the architect how expensive every single decision will be.

After a while (and Vince already has some faith in the architect) the situation changes.

Architect: "Ok, it looks like we have clearified almost everything. I just would like to make one final proposal that you might take into account. I was thinking that we could install a 10 feet steel plate between the first and the second floor."

Vince is really surprised. He has never seen a house with such a steel plate before. However, since everything the architect said so far appeared meaningful to him, he is curious about the proposal.

Vince: "Tell me more about it. What exactly would be the benefit?"

Architect: "Well, in case you want to increase the number of floors in some years, this steel plate probably allows it."

Vince feels even more insecure. He never thought he could have the desire to increase the number of floors in some years. And he has actually never seen a house in the neighbourhood where additional floors were added after the initial construction. However, the architect might be right. Maybe in some years he might be interested in that and if he does not take care for this upfront, it might become more expensive later on.

Vince: "Ok, I am really not sure whether I need additional floors somewhere in the future. But just in case: How expensive is it to construct the house with the steel plate?"

Architect: "I do not know."

Silence. This was definitively not the answer Vince expected.

Vince: "Well, but could you guess, how much it might be?"

Architect: "Seriously? No. I have never seen a house with such a plate. Even if I might be able to determine how expensive the plate is, I have really no idea how much is required to install it."

Vince: "Ok. Maybe we can clearify this later. But tell me, how many floors will be later possible with this ten feet steel plate?"

Architect: "I really don't know. But it is plausible that at least some floors will be possible in case the fundament is able to carry the weight of the plate and the additional floors."

Vince: "But the fundament will be able to carry the steel plate plus one floor, correct? I mean, you proposed to install this plate between the first and the second floor."

Architect: "In fact, I do not know yet. It might be possible that a completely different fundament is required. And it is also possible that we have to make a number of changes to the first floor so that the steel plate does not simply crush it."

We do not know how the story went on, but we do know that Vince is living since years in his own house that was designed by a different architect. And although not a single house in town has a 10 feet steel plate installed, there are continous rumours that there are such houses somewhere else.




What follows is my interpretation. Of course, there could be different ones, and I assume there are people out there who not only think that the story cannot be applied to software construction, but that the story is rather a good example why there are no or hardly any problems in software construction. I would love to read such interpretations and explicitly welcome any comments or additional interpretations.

2 Interpretation and Discussion

Would anyone of us take the architect's advice serious? Would anyone take the idea serious enough to be even discussed? Whom of us would just have bumped out the architect just because of even articulating such crazy ideas?

The interesting part of the story is, that probably all of us agree that the architect's proposal is completely stupid. We would argue that just having a new idea --- the steel plate between the first and second floor --- does not imply that the idea is meaningful. Nobody would blame Vince for bumping out the architect.

2.1 From House Construction to Software Construction

As soon as we switch into the role of software engineers, our judgement of the situation changes. Let's replace some words in the previous story. Instead of building a house, we speak about building a piece of software. Instead of the 10 feet steel plate, we speak about the new technique on the market that appeared just recently and that should become one of the central parts of the new software.

Such a new technique has different kinds and forms. It might be the new programming language that just appeared, comparable to the situation in the 90s when Java appeared, comparable to the situation some years later when PHP appeared, or comparable to the more recent revival of the approximately 20 years old programming language JavaScript. In the 90s, it could have been the new document format such as XML or more recently the format JSON. It might be the new IDE, the new API, the new framework, the new code generator, the new markup language, the new middleware, the new architecture, etc. In the 70s it could have been a newly released relational database system or thirty years later one of the NoSQL database systems.

What all these techniques have in common is that they suddenly appear on the market, cause some interest among developers or managers*1, and are taken into account for real software production (comparable to the house that is about to be built). This is similar to the architect's spontaneous idea. However, it still feels like the story (and the story's transfer to software construction) is absurd.

2.2 Is the Analogy Absurd?

One direct reaction on the analogy is, that the idea of the 10 feet steel plate is obviously totally crazy while the new software construction technique is not -- the analogy is absurd. No single house was ever built with such a plate. Hence, it is clear that such a house should not be built.

Well, actually no single software project was built with the new programming language either, which does not seem to imply for software engineers that they should not be the first to apply it. The argument “there is no single example for such a house” that seems plausible in house construction and that argues against the steel plate does not seem to work in software construction -- it looks like the analogy reveals something about the very different lines of reasoning in house construction and software construction.

We as software engineers have the tendency to say that the argumentation does not hold, because our knowledge on house construction is sufficient to determine that the plate does not make any sense while our knowledge, experience and intuition in software construction rather advices us to take one of the new software construction techniques into account. This issue requires a much deeper discussion of the relationships between history, experience and knowledge. This will be done later in section 2.11.

Another direct reaction of software engineers is that they say they are not naïve and do not blindly apply new techniques just for the sake of applying something new. But this is what the architect is doing by proposing an untested technique for being used in a final product. Instead, software engineers try out new techniques very carefully before considering them appropriate for being used in the construction of productive software systems.

2.3 No Naïve Application of New Technologies?

Let's assume for a second that Vince is our close friend and it really looks like he takes the installation of the 10 feet steel plate into account. We would probably try to argue out Vince of even thinking about the plate. We would do that, because we are not only afraid this house could financially ruin our friend; we are also afraid that the plate could do some serious harm to our friend and his family when the plate falls on their heads.

Maybe Vince is such an adventure seeker that he ignores our concerns. Probably we would ask him to build some kind of small model first. We do that because we want to convince him that the general idea is crazy. But even when the first model holds, we would probably ask him to build some small house first with such a plate -- maybe in the end we would advice him to build the house he desires with the plate but not to move into it, but to see first whether the steel plate actually does any harm to the house. We do that because we are not aware of any book titled “Why 10 feet steel plates cannot be used in house construction” and we are not aware of any building we could just show our friend. This implies that we cannot give a direct reference or resource that reflects the craziness of the 10 feet steel plate. Our advice does not mean that we take the steel plate seriously into account (again, based on the knowledge argument, see section 2.11). We do that because we want the test to fail: We want to prevent Victor from doing a serious fault.**3

When we now consider a typical argumentation in software construction, we see similar approaches but with completely different intentions. As said before, software engineers often argue that they do not blindly apply new techniques. Instead, they first try out new techniques in a smaller context. They do that because they want to find out whether the new technique could work. When they are convinced that the technique could work, they apply it maybe in another, maybe larger context. When it works there, too, they take the technique into account to be used for the development of productive systems.

Again, it is interesting to compare this argument with the (hypothetic) advice to our friend. We asked him to apply the technique, because we want him to see it fail. In case the test does not fail, we ask him to apply another test. We use the idea of testing to show the invalidity of the new idea. Even if one of the tests in house construction does not fail, we do not think that the appoach could work in actual house building. This is because we are aware that the reality is much more complex and cannot be simply immitated by a simple model. And we know that one, two or three simple models, i.e. try outs under non-controlled conditions, do not permit any serious and stable insights that could be directly applied to reality.

In software construction the idea of testing is used to show the validity of a new idea. Talking to software engineers reveals that their testing of new techniques is not a massive and critical analysis of a new technique. Instead, it is rather some test of plausibility whether the technique could work.**4 Hence, our advice that is directed against the installation of the 10 feet steel plate is quasi-inverted in the domain of software construction.

2.4 The Need for New Ideas

The reason for this inversion of argument is the difference between house construction and software construction with respect to how new techniques are considered. While in house construction new ideas are considered rather conservative, the software industry is rather open for new ideas. In house construction, new (and potentially innovative) ideas are only applied if the resulting risk is rather low. This might be because the consequences of failing in house construction are terrible for the builder (who probably does not have financial resources to build just another house) while the consequences in software construction are...well, we do not know, because we would need to know first, how large the software project is with respect to the company's resources.

However, software companies also often emphasize the need for new and innovative ideas, otherwise no helpful techniques would be applied at all. And not using new techniques has two bad consequences. First, the company becomes old fashioned. If it becomes known on the market that the company applies only old and conservative techniques, the company becomes uninteresting for potential new clients. The other point is that such a company becomes uninteresting for new potential employees, too. And attracting potential employees is necessary in order to get new developers.

Both arguments might be right, but interestingly, this falls rather in the domain of marketing: software companies assume that applying non-new techniques (and ignoring in that way any bleeding-edge technologies) would lead to some kind of negative reputation that negatively impacts the company. This implies the software market somehow desires bleeding-edge technologies -- which is an interesting argument, because it is at least not obvious why the software market should work different than other markets. In order words, it looks like a statement such as “building houses for a predictable price” that works for house building does not work in software construction. This corresponds to an often heard argument by the software industry, which says that they finally do apply new and innovative techniques in order to reduce costs.

2.5 Cost Reduction by New Technology?

When speaking about costs, there seems to be a direct mismatch between the 10 feet steel plate and the application of the new software construction technique. The 10 feet steel plate causes incredible costs: even without knowing any details about house construction is seems clear that the steel plate itself will cause most of the costs for the house. At the same time, it is rather unbelievable that a new software construction technique could cause such costs.

Having said this, it is worth to think about some common argument heard from the modeling community who states that before a piece of software is actually built, it should be modeled first. Maybe this argument is right (maybe not), but it is obvious that it causes initially additional initial costs, because the development of the actual software starts much later because of the additional modeling step (that might turn out valuable later on).

While for the 10 feet steel plate some costs are obvious (at least for buying the plate; here, the steel price could be used for a first approximation), other costs were completly unknown: since such a plate was never installed into a house, it is unclear what additional steps need to be done. Things to be answered are how the plate would be laid on the first floor, what kind of machines might be necessary for that, how many workers might be necessary, etc. The (probable) changes to the first floor and the (probable) changes to the fundament also need to be considered.

For software construction techniques some costs are directly visible, because money needs to be transferred to someone else (who could be the producer of the new application server software). These costs could be hidden as well (because getting the new technique under control might require some man months). Some of these costs might not appear directly but somewhere in the future (when the only person who was nearly able to handle the technique left the company and it turns out that additional people require some training before doing any actual work) -- hidden costs that might be also valid for the house with the 10 feet steel plate (in case it turns out later, that the fundament finally gets damaged).

The interesting thing about the new applied software construction technique is that there is typically relatively few knowledge about it, especially when we speak about the potential costs they cause. Of course, each of these techniques come with a number of claims and promises, comparable to the additional new floors that can be installed on top of the plate. The additional new floors seem from first glance directed to the idea of scalability --- each of us knows a story about some technique in software construction that promises that the application of it makes the software more scalable. But in fact, these new floors are just an arbitrary promise. Statements such as “makes the software more readable or understandable” matches the “additional floors” as well as “increases flexibility or reduces maintenance costs”.

In fact, we do not know to what extent certain choices in software construction techniques cause what costs.**5 It is unclear whether the application of a new programming language will in the end cause a dramatic increase of costs or whether the applied framework is in the end responsible that our software project completely fails. Although we have the tendency to think that we are sure that the 10 feet steel plate costs too much, we are only less sure what the actual costs of the new software construction techniques are.

At that point, we find another objection: the objection that we are more sure that the new software technique will cost less than the steel plate, because there is actually some software market where the new software construction techniques are advertised while there is no such market for the 10 feet steel plate.

2.6 Market as an Indicator for Applicable Technology?

One obvious difference between the 10 feet steel plate and the new software construction technique is that the steel plate appears as the architect's spontaneous, crazy idea. This idea does not seem to have any background or foundation (at least the architect has not spoken about it). The 10 feet steel plate has not yet produced. There is no market where the installation of 10 feet steel plates is advertised to be used in house construction. There is no producer of 10 feet steel plates who creates such plates for the purpose of being used in house construction. The idea of the plate comes out of the blue.

The situation is different with new software construction techniques. As long as we do not speak about things such as process models, at least programming languages, APIs, etc. are things that have been produced already (before taking into account that they should be applied). There were people involved, time spent, investments being done in order to create this artifact. There are webpages where the artifacts can be downloaded or shops that sell these artifacts.

At least the existence of some market gives some trust that the provided products are more than just spontaneous, crazy ideas.

However, it is not clear whether this first impression is the same after a closer look. Maybe the architect has some friend who owns a company that actually is able to produce such a plate. Maybe this friend was so convinced that the architect will be successfully advertising the plate that he already produced some. And finally, we need to take into account that the architect was quite honest in the meeting with Vince: Instead of massively advertising the plate, he honestly told him, that there is hardly any knowledge about it. From that perspective, the presense of the market is not an indicator for the validity of using 10 feet steel plates in house construction. It is rather a question whether some company somewhere thinks that the plate could be sold -- although the plate might turn out to be not only useless but even harmfull. Applying this thought to the software market, there is no reason why we should not think about it in the same way. That someone produced some new technique is just an indicator that someone thinks he can sell it. It is no indicator for the usefulness of the technique -- nor whether the technique is usable at all.

2.7 Salesmen, Innovation and Responsibilities

The problem of speaking about the market and potential innovation is that the architect acts, from the perspective of a professional salesman, unprofessional. Instead of praising the new innovative product, he is honest to the client -- no wonder that he finally neither sells the steel plate nor any house. From the company's perspective (in case the architect comes from a compary) the architects behaviour is harmful.

However, from the client's perspective the architect's behavior is appropriate. His initial consulting services were satisfactory and he finally made a new proposal and gave an honest estimation of the usefulness of the new technique. Based on the architects estimation the client was able to determine on his own whether or not he is willing to take the risk and apply the technique. The only problem left was that the architect's proposal was from the client's perspective so unbelievable that he completely lost his faith in the architect.

From a software developer's perspective the situation is considered differently. Because of the architect's honesty it will not be possible to actually try out whether the new technique could be applied, i.e. whether such a house could be built. And after that, we will not soonly be able to find out whether it is possible to increase the number of floors on top of it. The software developer rather considers the situation as a missed opportunity -- and confuses something. First, serious house builders would not just play with a new technique and try it out for a new house (that is actually habited by people). Even if one or two of the most trivial tests succeed, house builders would still not simply apply a new technique, because they are aware that some simple tests hardly say anything about the validity of the new technique. House builders are aware that many different factors determine whether a new technique is appropriate -- and that some simple tests cannot replace serious studies about the appropriateness. It looks like software engineers have a completely different perspective on this risks. This might have to do that hous builders are aware that in the end someone has to pay for a failed test.

The other fact is, that the existence of something new does not necessarily imply some opportunity -- and does not imply innovation.

Doing something new without knowing the potential risks and without knowing how large the potential harm is, is irresponsible. The architect acts irresponsible because he does not keep away the potential harm from our friend. Proposing a certain technique (while being aware that the technique cannot be applied or at least not without a very high risk and very high costs) potentially causes harm. If the architect would have been succefully advertising the steel plate, it probably would have been Vince's ruin (and could have been worse in case the steel plate would in the end fall on our friend's head). In the end, taking a risk is not for free -- but people need to be aware of how large a risk is in order to decide whether or not they are willing to take it. Here, the unfortunate situation is, that most of the risks in software construction are yet unknown. Software developers have the tendency to say that as long as these risks are unknown, they can be ignored. Other disciplines (such as house construction) have rather the tendency to say that an unknown risk must not be taken.

2.8 But Finally, it Works!

Putting the previous concerns aside, software engineers often argue that finally they are able to deliver a running piece of software with the new software construction technique while the architect is not able to deliver a house. Hence, finally it works! The right analogy would have been that finally the house would have been built including the 10 feet steel plate.

Unfortunately, this often heard argument has two problems. First, it is untrue with highly probability. And second, in case it is true, it has to do with something software engineers typically don't like to speak about: budget.

There is the often found statement that software projects have the tendency to fail quite often. However, although this statement seems to belong to the general knowledge of software engineers, it is still unclear to what extent this statement is true. Mostly, people refer to the CHAOS report that reports that between 15% and 40% of software projects fail. However, there are people that argue (for very good reasons) that the numbers of the CHAOS report should not be taken too seriously, because it is unclear where they come from. However, there are other sources available that report a cancellation rate between 10% and 15%**6. Let's be positive with the software developers and let's just assume that 10% cancellation rate is correct. Again, let's go back to house construction. Would a serious house builder accept a 10% cancellation rate? No way!

But even if we would accept the wrong argument “Finally, it works”, we still need to ask, what the budget constraints were. Again, if there are no budget or time constraints at all, there is no risk at all, because a project can go on and on forever.

Again, it is possible to build the house with the 10 feet steel plate. The only problem is, that in the end it is getting expensive.

2.9 But Software is Unique!

Coming back to the cancellation rates of software projects, developers typically argue in another very specific way: Software is unique and therefore it cannot be compared to anything such as the construction of a simple house. The problem with the analogy is not, that we build a piece of standard software (the house) with one single new element (the 10 feet steel plate). Each piece of software is an indivual act of creativity of software engineers and the requirements for a new piece of software are so unique that they cannot be compared to anything else. The cancellation rates do not express that people have troubles making software because of the application of unknown techniques. The cancellation rates express that a lot of software is so unique that it is hard to predict whether the desired piece of software is doable at all.

This argument is interesting, because it says a lot about how software developers think about themselves. It reflects the perspective of a young and wild discipline where its members frequently address problems that appear unsolvable and that are finally solved as the result of hard work. And in case the problems are not solved, it is the result of unsolvable requirements.

However, when we finally take a look into the software market, we see a lot of investments into (and actually built) software that does not appear that completely new. We see web applications where people register themselves and where personalized information are shown, games on mobile devices that do not seem to be much different to games we have seen before, or migrations of existing APIs in other programming languages. Yes, from time to time we see new and innovative products and techniques. But this is what happens in house building as well.

At least, it would be interesting to think about whether the high cancellation rates are an indicator that software construction spends too much of the time on techniques that cause later on (when the product has not yet been delivered) serious troubles. In order to apply the 10 feet steel plate analogy again -- it is valid to question whether the high cancellation rates are more an indicator for how often equivalents of 10 feet steel plates are tried out in software construction instead of using the cancellation rates as an argument for the uniqueness of software construction.

2.10 Software is Much More Complex Than Anything Else!

This cancellation argument is often used not only to argue for the uniqueness of software. It is also often used to argue that software construction is just much more complex than anything else. Already in simple programs there are a large number of threads, communication processes, events, computations, etc. And a single error somewhere could already break the whole software. Just because software developers can make a large number of errors it makes their life much harder than the architect's life who just needs to compose four walls, a door, some windows and a roof. A side effect of this argument is the implication that house construction is rather a simple task.**7

Again, this argument contains some naivity of software engineers with respect to other disciplines. For unknown reasons software engineers do not only have the tendency to declare their own discipline as rocket science, they have also the tendency to consider other disciplines as trivial. As described before, a large amount of software we find on the market is not so new and completely different than any other piece of software that already existed. This does not directly imply that software is not complex, but it implies that at least a larger number of people were able to cope with software's complexity.

Both previous arguments (cancellation because of software's uniqueness, cancellation because of software's complexity) have the same tendency to invert a given argument -- something that already happened in section 2.3: In section 2.3 it was discussed that simple, successful tests in software construction are considered as a proof for the validity of a new technique and the legitimation for the application for such techniques. In other desciplines such simple tests are not more than small indicators for new techniques -- without any implication about its applicability. Here, we find that a high number of cancellations of software projects should be used as an indicator for the complexity of software construction and the uniqueness of software construction. Other disciplines would argue about this as an indicator for immaturity.

However, it is time to come back to the very first argument: the statement that the discipline of civil engineering and software construction cannot be compared because of the differences in experience and knowledge in both disciplines.

2.11 What about Experience and Knowledge?

A already mentioned in section 2.2 there is another argument against this analogy that should be discussed here. While there is a very long history of house construction and a long history of knowledge from civil engineering, these statements do not hold for software construction. Software construction is still a very young discipline. Our knowledge cannot be compared to civil engineering at all.

This reaction is be not completely wrong: obviously, houses are constructed since centuries while software is being constructed only since decades. As a consequence, we would directly think that the knowledge in both disciplines is different. But in fact, there are two different facets used in the same argument. The first one is related to age in terms of years (which probably somehow correlates to experience) and the second one is related to knowledge. Discussing both issues is not trival and requires some space.

2.11.1 Age and Experience

With respect to age and experience, we need to agree that software construction is quite young -- let's say approximately 60 years. However, with respect to experience, we just have to take a look at our mobile devises, to ask our bank's webpage about our savings, or to start our car (and getting directly feedback from some automated procedures about the car's condition and getting a bluetooth connection with our mobile that directly starts acting as a navigation system). A lot of software has been already constructed. While it is true that software construction is relatively young, it is at least not obvious whether there is not already “a lot of” experience in sofware construction: At least it is not directly obvious whether this experience is “less” than the experience in house construction.

The problem comes directly from the word experience, which is often used in different meanings. Often, a phrase such as “experienced craftsman” means that someone has already done a lot of work in a certain domain. We assume that such a person is aware of lots of problems in his domain and has found (or was taught in) means to solve these problems. This means such a person is able to detect a problem, i.e. in a given situation he is able to see the similarity to something he went through in the past. Solving a problem means to apply some tricks that helped before and that quite often actually do help. In summary, we assume that such a person went through some learning curve over some time.

Having said this, we are aware that it is not necessarily the case that a person who spent many years in a certain domain actually has much experience. Experience implies “having gone through different situations and problems and solved a number of them”. It is possible that a person has spent many years in a given domain although he has not gathered much experience. On the other hand, it is also possible that a lot of experience is gathered within a relatively short time frame.

Here, again a number of people will complain that the analogy of the experienced craftsman does not hold, because the way craftsmen are trained is to a certain extent the transfer of experience -- since this happened over centuries, this cannot be happen in software construction. Well, again, this argument is maybe partially true. But continuing the discussion here would take too far away from the original question. It is only intended here to say that “age and experience” are two different things. And it should be emphasized that “long time” does not imply “much experience” while “short time” does not imply “hardly any experience”. It should be only mentioned that it is at least not directly obvious whether experience in software construction is “less” than the experience in house construction.

2.11.2 Experience and Observation

Although the previous argumentation seems somehow plausible, there is one thing missing. As argued before, observations (i.e. the craftsman who sees a problem) are an essential part of experience -- and observations are finally the connecting part between experience and knowledge. However, the way how observations can be made is quite different.

When we see a house, most of us see a difference between a ruin a a newly built house (although we are wrong from time to time). When a house collapses (and it is for everyone in the neighbourhood observable when a house collapses) we imply that there was something wrong with the house. When we see a number of similar houses collapse, we assume that this is related to the commonalities between those houses. Such observations (which might be wrong) and theories (the commonalities between the houses that cause the collapse, which can be wrong as well) can be done by everyone who just understands the concept of a house: Such observations can be done by non-experts.

In addition to those observations by non-experts, there is an incredible amount of observations possible by experts. Even if something completely new is tried out in house building, experts have been trained in doing certain observations in order to check whether there is something seriously wrong. They would be able to detect much earlier than non-experts whether the house's fundament is in trouble. Or whether the walls begin to have problems to carry the weight of the roof, etc. Just to make sure: so far, this text speaks about observations and not about applied knowledge. I do not mean here that the civil engineer applies his knowledge about statics in order to compute whether the house is doomed to collapse. He just observes certain phenomena (such as cracks in the wall) where his experience tells him that his might result into serious problems.

If we compare this to software construction, a simple observation such as “a house collapses” is not that easy. From time to time our webserver is not available. Or there are frequent problems with our WiFi. From time to time, we need to restart certain programs. But it is very often the case that it is not possible to connect the observation directly with a certain software or a certain characteristic of software. For a non-expert it is relatively hard to make observations that are directly related to the software. Again, it would be possible to argue here against this. Windows-users knowing the blue screen do observe that something went seriously wrong which seems comparable to “observing that a house collapsed”. However, in that case it is only the environment that crashed which might have something to do with the software or not. Or Linux users are able to observe that a program stopped with a segmentation fault -- an observation that is (somehow) directly related to the software itself.

The software examples above might or might not be applicable but there is something wrong with the analogy. Even of we see a piece of software crashing from time to time (and if we do think for some reason that this is related to the software itself), it does not help us to judge what part of the software is responsible for that, because software is hidden behind a user interface. We do not see directly the different parts the software consists of. Hence, we cannot judge what parts of the software are problematic. Experienced users who frequently use a certain software are often familiar with problems it causes. They know that certain interactions with the software (long tables, certain input frequences, certain functions) lead to troubles and therefore do not do those interactions. In fact these are observations. But these observations (again) do not permit to identify those elements that actually cause the error. A spontenous reaction here is to say that for example the usage of tables in a text processing application is the element in the program that is problematic. But, we do not know whether the table implementation is a modular unit and we cannot conclude what part of the software needs improvement in order to fix the problem. Compared to the house with the 10 feet steel plate (under the assumption that the house has been actually built), we might see that jumping around in the second floor increases the cracks in the fundament or the walls in the first floor. We stop jumping around (comparable to stop using tables in our text processor), but we cannot conclude from it what might be the reason that actually causes the cracks.

When we focus more on the general idea of the discussion (the relationship between experience and observations), the main argument above was that experience depends on observations. But as argued before, we have troubles doing observations in software construction. While it is hardly possible to do such observations as non-experts, it is even hard for experts to do such observations. While some characteristics of a software are observable (the size of the software in terms of hard disk memory or memory consumption at runtime**8) it is unclear what can be concluded from these observations or what additional observations are possible or necessary. For example, even if we know that a piece of software requires 10 mb space on disk or 50 mb space at runtime, it is unclear what exactly this says about the software and its ingredients. While it is obvious that the civil engineer is able to detect a crack in the wall, it is unclear what comparable observations could be done for software.

One could argue that software is more a mystery than something else. However, the statement here is that we have not learned so far what could be observed. And as a consequence, our knowledge on software construction is quite limited.

2.11.3 Observations and Knowledge

In addition to experience there is the concept of knowledge -- again, both concepts are not directly related to each other. It is possible to extend knowledge by reading a book while most of us would not assume that reading a book increases experience (well, at least no as long as we do not do some additional exercises that are possibly proposed in the book). At the same time, it is possible to increase experience without increasing knowledge. Sport might be a good example for that. An amazing football or soccer player might be able to do great things on the field; because of hard training the player is doing the right things at the right time. He does not need to think about how he currently moves in order to get some advantage on the field -- he just does the right things by intuition. This does not imply that he is aware of any of those things he is doing, nor does he needs to be aware of any causal relationships between actions such as catching a ball and feinting a movement at the right time in order to get some benefit.

Yes, in house construction a lot of experience has been gathered over the centuries. But what's much more important is that over the centuries a lot of knowledge has been gathered. The origin of this knowledge might have been experience in the beginning based on observations. But knowledge is more. Instead of doing the right things by intuition, knowledge gives us some a conceptual framework to understand a situation and to judge whether a certain action would work in such a situation.

When we think about the 10 feet steel plate we do not only think after the house has been built that the plate is responsible for serious harm. We think before building the house that it is a bad idea to install the 10 feet steel plate. Every software engineer has already made some experiences with certain APIs or architectures that are badly implemented. In these cases they might refuse to apply the same technique in later projects. This is comparable to a civil engineer who actually has installed such a plate once and came to the decision not to apply the technique again**9. But (again) the difference is that we know upfront that the steel plate is a bad idea.

Civil engineering has gathered a lot of knowledge over the last centuries. The different characteristics of stones and woods being used in house constructions have been studied (in addition to millions of other things) by experts and an amazing theory has been created as a result of these observations. All of these observations depent on the ability to do observations. When we want to study whether a certain stone is applicable for house construction, we need to study what weight the stone is able to carry, to what extent the stone can be damaged by water or wheather, etc. And all of these studies depend on the observation when a stone gets cracks**10. And civil engineers are aware that this knowledge is not the result of a singular test. When a stone needs to be tested (and people are aware that two natural stones are not directly identical), this is done by a larger number of tests, each with a large sample size. This implies that as long as the idea of cracks had not be found, none of these studied was possible. And without any of these studies, no general theory of house construction would have been possible: The maths, that expresses these theories would not exist. And without such maths on house construction, there is the high probability that people propose things such as the installation of 10 feet steel plates. And the result is that the cancellation rates in house construction would probably be comparable to nowadays cancellation rates in software construction.

Hence, we conclude that there is the urgent need to start doing observations in software construction. As a first step, it is necessary to identify what is needs to be observed -- similar to the cracks in the wall. The authors opinion is that human effort (in terms of working time) is a first step towards this direction. Based on these observations it is possible to gather knowledge that might finally lead to some (tested) theories about software construction.

2.12 Wait! Software Construction does not Follow any Laws of Physics!

Wait! There is another often heard spontaneous reaction on the previous arguments. The typical reaction is: “House construction depends on the laws of physics. But software construction does not!”. It is probably true that no direct implications are possible from the Newton Physics to software construction. Well, no direct implications of Newton Physics to medicine is probably known either -- which does not prevent medicine to apply scientific methods to study the effect of drugs, etc.

The main problem is, that the laws in software construction are not yet found. It is unclear what the general laws are that underly the understandability, readability, or maintainability of software. No general theory about software comprehension has been found yet. But this does not mean that such a theory does not exist. It just means that this has not been found yet.

There are people that insists on the non-existence of these theories. And from time to time these people are even found in academia. The astonishing thing about this argument is, that this would mean that software construction is a purely probabilistic process that cannot be directly influenced by some external interventions. If this is true, it implies that none of the new programming languages, IDEs, architectures, etc. would have any effect on software construction. This idea is not naive, it is crazy -- maybe even more crazy than building a house with a 10 feet steel plate.

3 The Meaning of the Story

The meaning of the story is, that in software construction we find ourselves probably more often than we want in situations where we seriously take 10 feet steel plates into account: crazy new ideas that in the end do harm on the original goal (the construction of something new). The goal of the story and its discussion is to illustrate that an obviously absurd situation stops being absurd as soon as it is considered as an anology to software construction. It should emphasize the author's opinion that there are lines of reasoning that we directly reject in our daily lives for some good reasons but that we do not reject as software engineers. This is an indicator that there are serious problems in software construction.

Our goal as software engineers should be in the future to identify those 10 feet steel plates upfront, i.e. before we actually build a certain software that crashes because of the newly applied technique. This should be done with the goal in mind that in the end we want to improve our discipline.

I often heard the argument that this story tries to kill innovation.**11 The main argument is, that although it might be possible that a new technique turns out to be dangerous or just useless, it is necessary to try it out. Otherwise it would not be possible to achieve any improvements or progress. Trying out new things is essential not only for the discipline of software construction but for the whole society.

I totally agree with it. There is a need to do improvements and to make progress. And I see the need to try out new techniques. However, I do not only see the duty to try something new, I see also the duty to deliver finally a product and the responsibility to deliver the product with a reasonable effort. I do not think that playing with new techniques is a reasonable and responsible step into the right direction. New techniques should be seriously studied and tested before taking into account in productive software systems.

Another goal of this story is to invite people to think more about to what extent currently applied techniques in software construction are completely unknown with respect to their risks. The name of this story, the 10 feet steel plate, might also be used as a metaphor in software construction as a technique that probably implies a very high risk and that comes with a lot of promises where it is unclear (and maybe even unprobable) whether it is able to keep any of these promises.

The right steps to prevent us from building software with 10 feet steel plates is to start doing observations. Not singular ones, but a large number of observations in controlled environments. These observations will lead to knowledge. And without such knowledge software construction will remain a mystique field full of wild ideas that in the end lead to houses constructed with 10 feet steel plates and an amazing number of houses that simply crash.

4 Acknowledgement

I would like to thank all the students that participated in the discussions about the 10 feet steel plate and mostly those students who were reluctant to accept the story as a valid analogy and who argued about the absurdity of the analogy.

Footnotes

**1 It needs to be mentioned that Vince is a software developer and software developers have from time to time the tendency to build things on their own from scratch.

**2 It is not completely clear where this interest interest comes from, but at least some first approaches exist to find possible explanations (see the great survey by Meyerovich and Rabkin for such possible explanations [3].

**3 The goal is not to repeat here the arguments of experimentation and falsification in software engineering. In case the reader is still interested in that, a different article discusses that in more detail (see [2]).

**4 Probably the reader has now the tendency to think that this discussion finally leads to the killing of all inventions. This is definitively not the goal. I come back to this objection in section 3. 

**5 The essay by Andreas Stefik and me speaks about the problem of unknown costs and unreliable statements about possible benefits of programming languages in the programming language community (see [4]).

**6 In fact, this question has not been studied so far in detail, but at least there are studies such as the one by Emam and Koru [1] that give first hints about how often software projects actually fail.

**7 The funny thing about this argument is, that in the 90s the software community was massivly interested in the works by Christopher Alexander, an architect that actually built houses. At that time, software developers rather spoke about the parallels between house construction and software construction -- which in the end led to the works on software patterns.

**8 In fact, observing the memory consumption at runtime is far from being trivial, but it would lead to far to discuss that here.

**9 This would mean that a simple test falsified the idea of the steel plate and hence he concluded the non-applicability of the technique.

**10 Again, this explanation is slightly too trivial, because a crack does not need to be something that can be seen by eyes but by some other instruments. Again, in order to keep the discussion simple, this is not discussed here in more detail.

**11 In fact, this is not only a reaction on this story but also a reaction on the demand to apply human-centered usability tests to software techniques in general.


Literature

[1] Khaled El Amam, A. Günes Koru. A replicated survey of IT software project failires. IEEE Software, 25 (5), pp. 84-90, 2008

[2] Stefan Hanenberg. Faith, hope, and love: An essay on software science's neglect of human factors. Onward 2014

[3] Leo Meyerovich and Ariel S. Rabkin. Empirical analysis of programming langauge adoption, OOPSLA'13, pp. 1-13, 2013

[4] Andreas Stefik and Stefan Hanenberg. The programming language wars: Questions and responsibilities for the programming language community, Onward 2014, 2014.