Random input software testing

August 23, 2011

The usual approach to testing software is to create a specific problem and see if the software gets the correct answer.  Although this is very useful, there are problems with it:

  • It is labor-intensive
  • It almost totally neglects to test the code that throws errors
  • There can be unconscious bias in the test cases created

One alternative is to create problems with random inputs.  The talk I gave on this at useR!2011 was “Random input testing with R”.

There was a question at the talk concerning full coverage — basically that a fully random distribution is not going to be efficient at covering the whole space.  I don’t have any particular experience with this, but here are my thoughts:  If you have a space you are concerned about such that you can keep track of how many times each point in the space has been hit, then you could dynamically change the distributions of inputs to increase the chance of the lesser hit points being hit.

Past versions of the Portfolio Probe software have benefited from this technique on an experimental level.  Future versions will be subjected to more thorough tests of this sort.

Subscribe to the Portfolio Probe blog by Email

Leave a Reply

  1. […] Patrick Burns talked about random input software testing and had a great analogy: if writing test suites is like digging ditches, random input testing is like digging in the sand (ie, fun). (I do random input testing, but with my users providing the inputs.) [slides (with transcript)] […]

Related posts

  • July 18, 2011

    Ticker Sense posted about the mean correlation of the S&P 500. The plot there -- similar to Figure 1 -- shows that correlation has been on the rise after [...]

  • July 11, 2011

    If a particular prediction comes true, how surprised should we be? The prediction The page that sparked my curiosity tells of a prediction made a year ago that the [...]

  • June 30, 2011

    Winsorization replaces extreme data values with less extreme values. But why Extreme values sometimes have a big effect on statistical operations.  That effect is not necessarily a good effect.  [...]