Acceptance Testing For Continuous Delivery
transcript
my name is Dave Farley I'm a newly independent consultant and software developer this talk today is a new talk I want to give everybody fair warning that it's actually a significantly more technical talk than my common conference talks of late so if anybody's afraid of design concepts and code now is the time to go and leave and leave and find something else because that's what this one's about really I wanted to talk about acceptance testing for continuous delivery this is a topic that's close to my heart this is this is one of the things that I what one of my contributions to the book amongst others and I think this is really important and often continuous delivery gets kind of we think a lot about the automation of deployment and the configuration management and all of those things are important there are lots of things that are important in making continuous delivery work but the automation of testing is is crucial in my opinion and I want to talk about some ways of approaching that that make it easier to live with because it's hard to do this stuff well I want to talk a little bit so what's the role of acceptance testing if you've seen any more presentations before this is this is my picture a deployment pipeline and really what we're talking about is a series of validations that any change will go through to find out whether there's something that's broken so we're going to you know have a in multiple levels of validation as we go through the process all I would probably say that these bits are they are what I would count as acceptance testing so I'll talk more about what I mean by the definition of that the bit that I'm gonna I saw a bit I'm gonna focus on today though is is largely despit which is acceptance testing in the purest sense and what I mean by that is that it's the automation of tests around asserting that the acceptance criteria for your stories are met so I think of this is kind of the automated definition of done and executable specification of the behavior of the system that kind of stuff so what's an acceptance test so it's asserts that the code does what the users want of the system a unit test asserts that the code does what the developers think it ought to do and accept the assess is different to that it's focused on on the the behavior of the system from a user user on it it's an automated definition of done that the asserts that the code works in a production like test environment so you're not just testing the behavior of the system you're testing it in a production alike simulation a simulation of the production environment you want at your no C so you're a sort a certain kind of more broadly that the software is deployable is configured correctly in all of those other things as well as just the the interactions with the system it provides feedback on stories and kind of closes that feedback loop and often it gets talked in these context acceptance test-driven development BDD specification by example executable specifications are those sorts of ideas so as I said I think the shortest form of the definition from my perspective is an acceptance test is an executable specification of the behavior of the system and if you think of it that way it's a valuable way in to to the to the approach so where does that sit so this is again a slide that's kind of pitch from other presentations I I think of software development as a series of interlocking feedback loops and the establishment of these feedback loops is crucial to making it work effectively and the feedback loop that i'm talking about today is really this it's the it's not the kind of inner TDD cycle that's you know that's covered that's fairly well covered but this one I think is often less well covered the outer loop is kind of the the the feedback loop from ideas to getting valuable software in the hands of a user it's a really important feedback loop probably the most important one but I'm just talking about this one today so what's so hard why is the why is this difficult why do people struggle with doing these sorts of tests so so part of the problem is that tests break when the the system under test Changez and that's because the the tests are too coupled to the system tests are often complex to develop people complain about the cost of developing and maintaining these tests and the history is kind of littered with bad practice and poor implementation this presentation has a a series of little things in it so I'm gonna put slides up and I'm gonna stamp on what I think are the anti patterns so so one at one of the little games you can play Z skis kind of anti pattern bingo when the slide comes up see if you can spot the anti pattern before up with the stamp on e so hoons who owns the tests so I think it's e we want to have lots of these tests we want to have probably if it's a significant system tens of thousands of these tests running to assert that that behavior is correct that the common pattern that I use on the teams that I work on is that we specify story we define the acceptance criteria for those and part of our definition have done is that there is a minimum of one acceptance test for each acceptance criteria in the stories that means that you build up coverage in these sorts of tests very quickly it also means that you have probably over testing but it's a nice easy way of kind of closing the feedback loop making sure you're building the coverage with that I would it be too smart into predictive about about how you design your testing strategy so it's very nice that anybody can write the test if you've got a if you're a technical business analysts great if they can write write the tests if you've got a testing professional they have to not have insight but what I think is crucial is that the developers own the test through the life cycle this is one a lesson that I so I've I've I've spent probably in the order of about thirty five years a little bit more doing this kind of thing and I have never ever seen a separate testing team how many automated tests work this is a toxic idea what happens is that the development teams write lots of tests sorry the development teams write lots of code a separate testing team write some tests the developers team to change the code break the tests the testers start trying to fix those meanwhile the development team of who of whom there are usually many more carry on writing code and changing stuff and the testing team never ever ever catches up the test spend all of their life broken and are of little value this is a really simple problem to fix is that you make the developers responsible for keeping the tests passing so as soon as the developer makes a change they notice that they've broken the test and it's their responsibility to fix the test there and then even if they didn't write it and it's fine if they didn't write it it's fine if somebody else wrote it but they own the responsibility of fixing the tests and I'm not a big believer in hard and fast rules but this is one of this is the close very close to what I get to an investor all these days in terms of the way that I sell the teams that I work on so developers only acceptance test that's kind of lesson one if you like it's a really important thing to get this work thing with you and it's quite unusual in my experience in organisations it's not it's not the way that we set things up pretty often so the rest of this talk I'm gonna it's what I'm gonna focus on this this is kind of my agenda for the rest of this talk I'm gonna go through each of these things so what are the properties of a good of good acceptance tests a good acceptance test focused on the what not the how you focus on what it is that we're trying to assert not how it is that we're interacting the system under test that's a really important distinction this is a an important separation of concerns that leads you to a certain approach to designing the the tests and the infrastructure that support the tests need to be isolated from other tests we want to run many of these tests we want to run them in parallel perhaps you want to run them again and again lots of different ways you want to run them they need to be isolated so that we have maximum flexibility we want to be able to set the system up into a known good State and be able to assert that the things that happen to are the things that we intend they need to be repeatable we want to do this again and again we want to get the same results each time we run the test we want it I struggled with this one so I put use it uses the language of the problem domain I struggled with it because the rest of these things they're nice properties this is kind of a bit restrict prescriptive I nearly reword it to say they should be easy to write but that's not strong enough using the language of the problem domain is a key enabling a Pro approach to making this stuff work if you want to get the abstraction right you want to be able to express that the semantics of the test in the language of the problem domain in the language that the business understands it makes it more accessible for everybody makes it easy to maintain and crucially it drives the abstraction the technical abstraction that's going to allow us to implement this in a way that's going to be robust in the face of change we want to be able to test any change we want our build pipeline our deployment pipeline to accept any change and you evaluate that and and and us to be comfortable that what we're going to see is going to work and we want them to be efficient as I said we want to run thousands probably tens of thousands of these tests in a sensible amount of time and get a result a result so let's look at the first of those let's imagine that we have a system and this system has a number of different channels of communication with the outside world the the details of what these are is irrelevant but most system you know most systems of any complexity will be a number of different mechanisms technical mechanisms through which we all interact with the system and there's a number of you know groups of interested parties that you know domains subdomains that are interested in this let's imagine that we write a bunch of tests against such a system okay and let's imagine that there's a change in one area of the system which invalidates a bunch of these tests now if we've written if we conventionally many organizations do and we've conflated the concerns of what it is that we're trying to test with how it is that we're going to do the testing we're writing you know the typical kind of record and playback sort of style UI test is a good example of this where it can flex the concerns of what it is that you're trying to do with how it is that you're interacting the system so which buttons to push which fields to fill in which checkboxes to select that kind of thing if we've conflated those things then the only way in which we can fix this change is to go to each individual test and fix you there like most problems in software development this is fixed by introducing a level of indirection I did go be animation happy in this presentation forgive me it gets worse so what we can do instead is these channels here that are represented we think we can kind of make concrete we can represent those channels as part of our testing freh structure and the tests can interact with the system through these channels so now if if our fix API changes and our test abstraction of the fixed API break it breaks all of these tests we've only got to fix this part to fix all of these tests that broke in so so this is kind of the device driver pattern so we're rewriting a device driver that abstracts the how we're going to interact with that system on the test from how the test cases are going to to express their needs at this level these things are talking in business concepts these things are placing orders creating users well whatever it is that they're doing they're creating accounts whatever it is that they're meant to do at this level it's talking the protocol that the system understands but this thing is the only thing that's responsible for kind of encoding that interaction in this circumstance we can fix it in the one place it fixes all the tests so we kind of had a win with its simple and err for us to make tainless actually if you kind of follow this through what tends to happen is that there's there's a bunch of components and stuff that kind of sits in here that forms part of our test infrastructure so there are the concepts in the problem domain they're rather the encoding of these channels of interaction that become common amongst their test cases so we're starting to build at the idea of some test infrastructure that's specific to our problem domain and specific to that you know the system that we're working on but he's abstract and and and valuable and shared amongst a variety of different test phases another important thing is to kind of separate the idea of deployment from testing traditionally it's fairly common to think of every test should control its dark conditions and should start the app in the unit's application I'll give you a clue here so so there's an anti-pattern on here can you guess yeah it's the first one we don't want every test to start which has stuff to to settle the application in status and conditions because that's inefficient if we if we've got a big complex or application we want to be running these tests in lifelike circumstances it's going to be costly for do this for every test so we want to separate the decision of when we deploy things from where we test things we want to isolate the tests differently acceptance test deployment should be a rehearsal for production release that's absolutely true what we want to be doing through the process of continuous delivery is by the time we get to the point of release into production we want to have a very high level of confidence that our release is going to work you want to have validated that so we want to be using the same mechanisms for deploying the system into our test environments that we will use when it's ultimately put into production including any data migration or data provisioning or anything else all of those things and this is our first chance the acceptance test stage is classically the kind of first chance of rehearsing that as we go through our deployment pipeline this separation of concern that I'm talking about gives us a great opportunity for optimizing the process because it gives the opportunity to parallelize the test we can deploy the application once and then run many tests against do so we just incurred that penalty once there that cost once we can share it in different environments if we got this we can kind of go parallel that multiple environments what we do this it gives us many chances and of course low as the tests start up overhead so the next thing I want to talk about is test isolation so I would suggest that any form of testing is about evaluating something in controlled circumstances I think that's pretty close to a decent definition of what testing is all about if you wanted to evaluate something in control circumstances you need to isolate it from other things you want to limit the impact of other tests running alongside it from from this thing so there are a number of different ways in which this is true so we want to isolate the system under test we want to isolate test cases from one another so we can run them in parallel and we want to isolate test cases from themselves so we want week if we run the test case once and then run it straight afterwards we want to have repeatable results we don't want it to have left some stuff in the database that means that the second running of the test is broken isolation is kind of a vital part of your testing strategy and I've seen many organizations in the past suffer when they try and create a complex network of tests you if you run this test that sets what the start conditions for this other test and you build this complex rat's nest this tip dependency tree that you can never maintain you break one thing in the middle and you have no no idea what's going wrong so isolation is a vital part of your tent your testing strategy so first let's talk about isolating the system under test this is kind of a typical picture in many enterprise systems so let's imagine that you're working on on system B this is the system under test then system B is kind of in a chain of dependencies between between other system system a talks to it and system C's downstream and it's fairly common in large enterprises to have an integration testing when you kind of glue all of these things together and you saw it's a requirement that you're going to test testing to see in this circumstance this is problematic so if your work if your your mode of testing system B is to put inputs into system a then you haven't really got control of the inputs you've got you you can't really put the weird cases in unless you really know what's going on in system a to be honest if you're working in this system you probably don't really know much about system A or C either and so you don't really understand the inputs that you need to put in there you you know this kind of assumes that you have global knowledge of all of the systems in depth understanding and that's just not going to happen much it ends up that this is not really in a predictable deterministic State it means that you can't really know what state this thing is in and whether whether you're your test is going to pass or fail it's like it's like trying to eat spaghetti through a letterbox with chopsticks it's it's you're at your Europe a very long way divorce from the thing that you really care about so that's that's another anti-pattern so what we really want is something that's more like this so you want budget groups sorry sorry you know nearest animation was a mistake so what we really want is something more like this we want we want to really have control of our system now we do want to be interacting with our system in realistic ways so you know we want this to be interacting through the real communications channel to our system the API that's exposed the user interface that's exposed you know whatever it is we want these interfaces to be real interfaces but we want to attack we want we want control so we don't special access for the tests but we do want close access to our tests so that we can exercise that system as it would really be exercised we can put wacky things in here we can simulate the the upstream system braking and having exceptions or sending garbage downstream we have Falken we have much greater control over the state of this system and so their ability to test it and assert against it the trouble is is that when you've got a system like this well the reason that people do this kind of thing is that they're worried about these bits they were worried about the interfaces between the systems what about the stuff that you don't you don't expect what about the you know something changes in here that uses this into the interface itself doesn't change what something changes in here that changes the way this interacts with this and and so invalidates some your tests in some way and that's that's a valid concern so I think that what we're really looking for is something more like this we want to test those cases too we want to test these the interfaces but actually when you do that the amount of testing that you need it for these systems just to assert that their interfaces are good is much much smaller than the whole thing so now you've got much better control over the state of this system much stronger assertions that you can make against this system and you're defending yourself about against changes in the interfaces from from these systems this this this is a strategy there's a lot of devil in the detail of this strategy this isn't this is not a slam this is not a simple solution but it's a in my opinion it's a better solution and it certainly means that you have a better opportunity for testing this thing and there will be time as we're still you may get broken by changes in here that you didn't expect but that's a better situation than trying to look after the whole world and have the scope of your testing extend forever to the band the bands of the internet so now it's also test so test isolation so I'm going to make an assumption now that most of us are talking about multi-user systems this isn't always the case and if you're not talking about this test isolation is relatively simple because you then you do spend up milk multiple versions of the application as a way of doing that but let's assume that we're talking about multi-user systems those the systems are mostly I've spent the recent years of my career working so what we want is that we the test should be efficient we want to run lots of them what we really want to do is we want to deploy once and run lots of these tests so we must avoid any dependencies between tests dependencies in terms of state shared state persistent state any of those sorts of things so what we're looking for is a way of isolating each test from from the others that's going to allow us to run many of these tests in parallel against the say miss instance of an application so we can start it at once run a bunch of tests get some results and each test is independent a great way of doing this is using functional isolation so you look within your prone problem domain and identify the boundaries within your problem domain that makes sense and carve up your tests along those lines so if you're testing Amazon create a new account or and a new book every time that you want to run a test if you're testing eBay create a new account and a new auction every time you're running a test if you're you testing get up create a new account and a new repository every time you run a test that gives you the isolation you can set up the the account or the the marketplace whatever the auction whatever it is in precisely the state that you that's important for this given test case and use it just for this the implementation of this test case and then it's done you don't you don't need it anymore it has a weird side affecting that you tend to make the creation of accounts and marketplaces and stuff like that efficient but it's outside the other kind of isolation that's important tempura Louis isolation and we want we want repeatable results I want to be able to run this test and they don't want to be able to run it again and and and I want to get predictable results that's not always it's not always obvious that you can do that so let's let's imagine a little bit of code here so there's a a test here I'm gonna put my glasses on so I can read it here so here it doesn't really matter the problem domain here very much so we we're placing an order for a book so we've got an idea of a store we gonna create a book which happens to be called continuous delivery we're going to place an order for the book and we're going to assert that the orders been placed so the kind of the bit that we're talking about when we're talking about the isolation here is this is the problem bit if this is the book that we're dealing with so let's imagine that what's going to happen is that we're going to access execute that line and we're going to create this book called continuous delivery now the next time that we run this test continuous delivery already weak exists so the test it's going to fail probably or at least it's not going to be in the same state that it was the first time because we maybe we've changed the state and we change the price of the book is part the test something that may have changed we can't guarantee that things in the right State so another alternative is the first time so this book here is called continuous delivery in the test but actually what we do is that we take continuous delivery one two three four so we alias the thing in this scope of the application and there's next time we run we alias it as something else and now we've got our isolation based on our functional entities so it's a powerful practice to alias your functional isolation entity so so that you can have this repeatability you can run test many times and get repeatable results each one of these is isolated because it's an actually independent case I said repeatable lots of times the tests need to be repeatable so what do we mean by that so let's imagine that we've got a basis and this is our system under test and our system under test is talking to an external system we've already said that we'd like to cry and isolate ourselves on this alternative a external system if we're good software developers and we you know we're following the principles of design and we didn't just learn too we didn't only learn software development from software development for the dummies or by watching any visual basic demonstrations in the 1980s then we're not going to do this we're not going to have our primary logic talking directly to an external system we're going to have something that sits between us and the outside world to give us an option to the things to change that's good that's that's nice that's that's that's a slightly better design it's a better separation of concerns I would say that this so this thing is focused on presenting a nice interface whatever your definition of nice is to your your model here of this outside world thing the next level is that there's this kind of you know there's a the actual communications that's going on with the outside world I would suggest that's a separate concern from these because this is focused on giving some nice interface here this is just the the communication channel to this thing so if you have that kind of architecture with this little plug in here you have the opportunity to have an alternative plug in so you you want to test this interface because this interface is doing useful work it's mapping between your concepts and the outside worlds concept this thing is just wire protocol stuff that's just boring that's that's easy to that's easy to do so you can separate these concerns and you can kind of define this through configure raishin and then you can have something different in your external system to your production system in this instance we've got a stub of our external system so we're faking this interaction with our external system the real communication channel is still being exercised here the real thing is you know it's still talking as though it would in the real world it's just not really talking to the external system so this is really about where that sits in terms of the testing for structure so let's you know we've got our interface is that system under test these are external stub test infrastructure and there's a bunch of test cases what I'm talking about current is is this thing being enrolled as part of the test infrastructure there is some back-channel of communication to this stub here so I can collect results from my test interactions and submit inputs through this via my test infrastructure whatever that means this means that I have within this test infrastructure in abstracted so I my test cases can talk to the thing I can express these things I can say what I want of the external system so I can simulate different behaviors and so on through the real communication channel into my external system uses the language of the problem domain so I've kind of been leading you gently in the direction of talking about domain-specific languages so what I'm talking about here is abstracting a DSL allows us to solve many of these problems if we start approaching this and encoding our test cases in a domain-specific language and developing our domain-specific language it gives us the right level of abstraction it gives us a readability it makes it easy to create the test cases easy to maintain the crit of the test it allows us it gives us a natural separation of the watt from the how and it gives us the ability to do the test isolation as I've just described in a variety of ways in this infrastructure that I've been drawing pictures of so here's a simple example of such a DSL this is a real one this is from projects that I worked on there's a couple of examples here and you can see that you know you can even this is that this is a thing called an internal DSL so this is a DSL that he's actually executing in this case in the java language so it's written on top of java run we execute these things using j-unit but their sis complex acceptance tests interacting in the system and there's this is kind of a very stylistic current form of java which is the dsl bit you can I think I think it's fair to say that you can imagine somebody that's not a program and that's familiar with this understanding what these test cases are these are about as clear as a you know a templated Word document or Excel spreadsheet defining manual test cases this is what I mean about the ideal executable specification there are organizations that use some things like these as the the support documentation for the system so you can you get to make a support call you call up the support pers and say this is going wrong and they can search through the tests and say well the test says it behaves like this and they won't know that that's the case because that is the test that's run against that version of the software and so it's a very strong assertion it's not it's not a a a weak link between a document that may be deeper that may have drifted may be different to the version that's in production it's the test that was run against that version of the software in production this is a very strong link between in terms of specifying what it is that we want of the system and getting the results that but I think you can imagine it's relatively easy to write this without knowing how the system works in detail from a technical level or how the test interacts with it the bit that I'll ask this here's a clap the the typical setup that's kind of creating a and we called them instruments in this case so that was a market that you could trade in it's creating a user and so on that so it's doing the aliasing stuff that I was talking about earlier in order to to isolate this test from any other tests and from running it multiple times itself underneath the covers in the implementation of the DSL in this case there's been quite a lot going on so we can specify simple things at a high level just the stuff that we care about but there's also a lot of default parameters in here it's actually a very powerful thing we can go into a great deal of detail to specify precisely the state that we want the system to be in when it's under test or we can just take the defaults and and write stuff very quickly and move on this is probably this this is kind of nice it's arms-length it's abstract there is there's still a little bit of more coupling than needs to be and actually later in the life of this system kind of evolved a bit differently so you can see here we're talking about the trading UI on the deal ticket and the fix api so this was a trading system so you could either go through the user interface which is the trading UI here or you could go through the fix api if you're not familiar with fixie doesn't matter it's just an application program interface that's common in the finance industry and that's so that's okay but it's kind of coupling the channels a bit so later we kind of evolved to this and this was even nicer so now we've got an annotation which defines the channel that you could talk through so we could the fix API the deal to key at the public API and this test could be executed against all of them so we're getting to the stage where it's quite easy to write these tests very quickly this is very very loosely coupled with the system you can imagine very significant dramatic changes in this case the difference between a rich web user interface and a fix API or a binary API which is the other one that's listed here and it doesn't matter to the test so this is very loosely coupled but very fully specified in terms of our ability to control the environment and and and set the system under test open to the state that we desire it to be and you can kind of you know you encode stuff like the domain you build you domain over time to evolve your domain model your domain DSL your domain-specific language over time to match the testing as it grows as you grow the the system is something this is an evolutionary approach we want to be able to test any change so test cases need to be deterministic times a particular kind of problem when it comes to determining Surman ism which is often which is often ignored I think that broadly there are there are two approaches you can either ignore time so you can every time you your system talks about time you just ignore that value and pretend it's not there and you don't bother validating it or you can deal with it so ignoring time you can filter out time based values in your test it's nice and simple and it works for some cases mostly often though that's not enough if you certainly if you're talking about a complex enterprise system you probably can't under inaudible test cases that you're interested in and be able to evaluate that involve time so the alternative is to control time so I my approach is treat time as though it's an external dependency like any other it's an external system treat the supply of time information as you would any external system and stub it control it strategize it so that you can change you this is very flexible it's a very nice way of doing it it allows you to simulate time-based scenarios you can have long-running scenarios in your test that execute very quickly so in in the exchange where we some of these examples come from we we could run daylight savings examples we could test daylight savings changes through our system by fast forwarding time so that it crossed over a daylight savings boundary we could run clearing examples where so fastly in finance you clear three days later so you could kind of run the clock forward three days later and and test all of those scenarios so it gives you a lot more control it's slightly more complex in terms of the infrastructure so he's an example of of such a test so you can see I'm selling my book so let's imagine you're borrowing your book and and you know the books not overdue you've only just borrowed it so time travel one week books still not overdue time travel for weeks the books now overdue okay this is simple test easy to understand and it's quite a nice way of expressing it sorry that you're quite right that should say effort assert true this is the first time I presented that thank you for being my system test sorry now III I confess this one's a PowerPoint only test case the others were real so what we're talking about so again here's my pretty picture of a system under test and here's here's something some behavior of the system that's time dependent classically this is the kind of thing that we'll do we'll make a system call to get the current time so another way that we could do it again we're using one level of indirection so let's introduce the idea of a clock so we'll have a clock that sits between us and the system time so it's probably it's doing the same thing under the covers it's going to do the simply system but now our system under test only ever accesses time through this mechanism so here's an example of a clock we've got a clock which by default is a system clock which is just gonna do what he was doing before look at the system time but it also allows us to to change that it's it's public you could probably not write the code neater than that you could put accesses or whatever from your technology we're just trying to keep it simple sorry it does yeah yeah but it's but that's but still I want to be able to change that and not change the real time you know real underlying time of the system and all that kind of thing so here again part of our test infrastructure under control of our DSL as I demonstrated in the previous example we've got some stuff so here we're going to on initialization instead it we're going to change the clock so instead of using the system clock we're gonna override that and we're going to use the test clock and now when we call time travel we're gonna have a new time we're going to pause that time from the nice two string representation that's pretty in our DSL and we're going to change the time in our clock now we've got control of time and the our DSL can do the fast forward thing move time on and test all of these complex scenarios the other thing that's kind of it's that it's kind of interesting is tests that need a special environment so they're so I talked a lot about sharing environments and being able to run lots of these things in parallel so they don't crackle crunch into each other there is some kind of test that's hard to do if you're doing what I've just described if you're time-traveling you don't really been one to running tests that aren't time-traveling against a system where you're time-traveling because that's going to confuse things so you probably want to select separate those things a good way of doing this is you tag the tests with whatever the difference is and then you use an allocation strategy at the point at which you'll you're deciding to run the test so I deploy those two environments of your choice so for example a time travel test could be treated in a certain way a destructive test where you're you're killing bits of the sea to see how the system works under failure or a test that depends on specific hardware or something like that it could all be tagged and the thing allocates your test set to your your test infrastructure can use those tags and determining where to put them here is the nice visualization from an ex colleague of mine mark price off precisely that in play so the in this case we've got a series of sequential secret parallel tests so that's the kind of the norm that's what I was talking about you have one infrastructure running a whole bunch of tests in parallel and here we have a bunch of time travel tests where each one of these things each one of these nodes will start up a new instance of the application and that will be owned within the scope of the test so we can time travel it under the control of the test and then we have some sequential tests mostly these were destructive tests where we will be killing different bits of the system to see that the system was robust in the face of change and which also you couldn't do while you were doing other things it's kind of a trivial visualization but it's kind of shows the dynamism that's possible with these simple strategies so the last of my properties of a good acceptance test is that they need to be efficient we want to have thousands or tens of thousands of these things we want the ability to be able to evolve our test coverage as we're adding new requirements and new stories to our process we want to add a new a new test case for each acceptance criteria of those stories so we're gonna we're gonna end up with a lot of these things so if your production environment looks like this when I say apron these tests in a production like environment this is probably asking a bit too much if your Amazon or Google you don't have to lots of hardware infrastructure Amazon or Google size in order to do this kind of thing so what what what matters so you know typical interaction looks like this then clearly the selection of your hardware if you're testing which structure is gonna look something like this if it looks like this it's gonna in the selection of your horizontal at least what we want to do is that we want to be representative of the key concepts if you've got a distribution boundary that's important in your system it's probably important to represent that depending on the nature of your system it might be important that that's using the right kind of network card from low latency finance that mattered to in my world maybe it doesn't but you could simulate that with a VM it doesn't matter but you need to represent the stuff that's important in your deployed environment if the hardware is different in this machine to this machine you need to represent that if that's significant to your application you want to be testing in scenarios that our lifelike because that's the stuff that will catch you if you don't the other thing that the other thing that's important about efficiency is to make sure that each test is running efficiently so that you don't want to be wasting you would only spend inordinate amount of time in each test I would suggest so the kind of mental model that I have of an ideal cycle time is the kind of mental experiment is I want I imagine I disastrous situation going on in production you know something really Bad's happening and we're losing money okay in production you know or users or whatever it is something Bad's happening I want to be able to still go through my deployment pipeline at this point I don't want to be bypassed this is the time where I want to be sure that my change is safe I don't want to be making the situation worse so I want my cycle time to be short enough that I can get a result and get that emergency fix out in a timely manner but still be confident that this the change is good so I don't want to be spending days running these tests ideally I don't want to be spending hours so generally my rule of thumb is that I think that a reasonable compromise between having enough time to run a reasonably complex involved test somewhere in the order about 40 minutes or something like that maybe an hour for these things that's kind of my rule of thumb I tend to try and operate at that kind of level for these sorts of tests one of the things that tends to trip you up with this in terms of inefficiency that's very common is putting weight statements in your tests particularly when you're dealing with asynchrony I would suggest it's a really important point when you're dealing with asynchronous systems and if you're not you probably should be looking at doing that because they're much more efficient and easy and have fewer problems but one of the things is to look for a concluding event so you you place an order and there's some other event that happens later after that successful and catch that so you don't you don't want to make your system synchronous just to make your test work because asynchronous systems are faster and more efficient and easier to reason about but you do want to you do want to your test cases to be sink rest because you want it to do a step and then move on at the level of the DSL EQ line in your DSL is synchronous you don't list to be a complex coding exercise you want anybody to be able to understand the sequence of events and that what's the flow through your test cases so here's a trivial example imagine that this is the implementation of a place or the method within the DSL so this is the thing that test cases will call and what you want to do is you want to separate the constitute the school so you going to send an asynchronous place order message which is passed from the params that were passed in and we're going to wait for an order that the order was confirmed or fail on a timeout as a separate step that way we're presenting it in the test case as though it's a synchronous call even though it's not just as an aside it's kind of useful in the DSL is one of those little green blobs that I showed earlier in my diagram to kind of have the demo you know key domain concepts in this idea the idea of the order is a useful concept so there's probably some stuff that you'd like to remember about that and it's probably material to a bunch of different interfaces I mean you know as part of the implementation of the DSL you can roll up into this in terms of your implementation so if you really really really have to do a poll and time out mechanism in your tech do it in your test infrastructure not in your test cases but this first way you're looking for the concluding event is a much stronger strategy and please don't put weights and expect your tests to be reliable or efficient they will be neither if you do this if you've done all of this if you've taken this approach then it makes this next bit easier so what tends to happen over the life of a project is you start off with a relatively simple test infrastructure and then your 40 minutes duration starts getting eaten into it into the intervening no time at all is taking a day and a half or two hours and that's not good enough so you want to throw some hardware at it you want to parallelize it so let's imagine we have something in our artifact repository we've got some coordinating function that's running our acceptance tests we're going to deploy that release candidate to our acceptance test environment then we're gonna spin out a bunch of parallel tests hosts that are going to run all of these tests in parallel against this thing this is kind of nicely widely scaleable to go on in the limit of performance of your system so you can kind of turn up the dial you know if you want your test results in half an hour buy some more hardware buy some more VMs go to go to Amazon and bison you know buy some of their their services you can scale this stuff up when I left del max oh she's a while ago now we had something like I think he was back 25 thousands except of these sorts of acceptance tests covering all aspects of the system we got a result in about 40 minutes if you run them serial if you run into end it took over 30 hours so it's an effective strategy this parallel lines paralyzing things so I'm going to round one ups with some some quick anti patterns and forgiving them to put my glasses on so it's easy to read here and they don't they so don't use rely you I record and playback systems they're too brittle they're too tightly coupled to the system under test and it doesn't work don't record and playback the production data this has a role but it's not acceptance testing you want to control the state of the system under test if you're just recording back production data if that's your only approach to testing you're only testing the common scenarios that happen to have occurred in the recording of that set of test data you're not getting the weird outlying circumstances you're not testing the exception cases and mostly it's the exception cases that you want to simulate because those are the ones that are usually going to kill you don't dump production titer in your test system because you want it scalable you want to be nimble you want to be able to deploy this all out all over the place you want to have flexibility to be able to scale this up very far and if you've got vast amounts of data you're not going to have to do that don't assume the some nasty automated testing product trademark he's gonna he's gonna do what you need design your test strategy you you decide what it is that you want out of your test strategy and evaluate things in that context don't have separate QA team cute having human beings who are specialists in testing is a very good thing for many teams I'm not saying don't have QA people but don't have them managing writing the acceptance testers as separate to the developers developers own the acceptance tests that every test start don't need I'll do that either don't include systems out of control don't put waiting stress I've talked about all of these so do you ensure that your developers own the test can you tell what I care about that I've said that quite a lot do focus your tests on the whatnot how do you think of your tests as executable specifications do make acceptance testing probably the definition of doing do you keep the tests isolated from one another do keep your tests repeatable do you use the language of the price for the problem domain I recommend that you try this DSL I do this is not language specific doesn't have to be Java I've done it in Python Ruby Java fitness all sorts of different ways it's the idea that's important it's the separation of concerns that's important in this in this pattern not anything else sorry pattern there's probably a bit strong use but do stub external systems limit the scope of your testing to the thing that you care about testing production logic environments because that will catch things that you don't expect do you make extractions appear as synchronous at the limb of the test case and do test for any change we want to be able to have a very high degree of confidence that at the point at which we push the button we're probably likely to be okay we never write enough tests to prove that our software is good but we can write enough tests that if one fails it tells us that it's not good falsifiability is an important part of this whole strategy and do keep your testing for efficient we want to have lots of them we want to be able to make them easy to write easy to maintain and cheap to run and with that I'm done so cool yes sometimes so some of the more complex bits so some of the infrastructure yes there is a the the stuff the examples that I was using for from Lmax there is an open-source version of the kind of test infrastructure it's a really really simple project it doesn't do very much you have to do quite a lot of stuff to build your DSL on top of it but it does something a listing for you the gives you the defaulting for properties that kind of thing is called simple DSL mix on the Lmax open source website that's worth a look but but some of that's unit tested so me some he's not we tended not to unit test as we grew the evolve the DSL but we did tend to to unit test the infrastructure I'm I'm not strongly against cucumber or the other things in that class my personal experience is in the use of it internal DSL I think they have an eye the nice property is that if you want the developers to own the tests then you're in the tool set of the developers you're in their world you use that you're using the languages that they're familiar with it makes it easier you do want to be able to run the you want you do want the developer develops to be able to run any of these tests in their local development environment and so and so I think that's a wing I'm sorry I'm not anti cucumber or anything I am anti a separate QA test team writing cucumber tests yes yes so generally generally I would advocate writing the test first it's it's a nice it's a nice way of focusing your mind about what it is that you care about the story if you do treat these as executable specifications it's not nice to treat these as your definition have done you know so so once the test passes once it passes all of it was all of those tests pass you don't know what else is there to do another thing is that by doing that debrief isolation but you introduce the risk of the first context of the internet blows up into 89 and that could be a massive maintenance and in the light of what you said later about paralyzing I would I would agree that it's not always the right thing to do however in the sorts of systems that I work on it's always been the right thing to do sure I'm talking I'm on the whole talking about large-scale enterprise systems big complex systems essentially it's down to can you can you start it up in under a second if you can't start it up in under a second it's probably too expensive to do it for every test case big and complex maybe and that's a calculation you have to do but maybe for in the systems I'm talking about you know it's certainly an anti-pattern really what I'm to worry what I'm talking about is making sure that you've got katroo you've got your control of the environment by definition is sweet I'm sure people know that they say do food doesn't work yet so generally what with hindsight I'm not quite sure mind it might my picture shows it but the the idea is is that what you're talking about is that the plug-in bit is kind of the external bit so it's so you are communicating with the application through its natural interfaces so you're not changing anything that's kind of run the runtime property of the thing this thing is really simulating the external bit not being the internal communications does that answer your question that's where I think the functional isolation thing comes in so so I don't like the idea of resetting the database that you know in the life of the test we're going to do that I want to blow it away and redeploy the application and do whatever it does from a known good starting position including any data migration routes because I want to rehearse the deployment so I don't want to be I don't want the tech the the state to be in an artificial state that's under control of the test I don't want to be going under the covers of the system in order to to set the state up I want to be going through real external interfaces of the system to get the system to the state that I want it to be yeah no sorry no in between every deployment in the testing environment so if it's this complex system I'm gonna I'm gonna start off from an you know maybe there's some reference data in my test data sample I don't know but I'm going to start off with the same position every time and I'm gonna use the deployment tools that are using production to migrate that to the version the new version of the schema or whatever to get into the state and I'm going to use my DSL to get the system into the state that I want it to be in in order to be to be ready for the test yes it is and and that's what that's what the functional isolation is about to try and prevent the need for you to to reset the data so I'm going to isolate the the scope of my test you know within the functional entities that make sense in my problem domain and that should give me the isolation that I need so are they keeping up with it with the development team are their tests broken all the time you know fix that problem so so if they aren't then I'm I'm wrong in that circumstance and fine but whatever works ultimately but my experience is that that I've never seen network they're always lagging the tests are always broken breaking and so the way to close that loop is that they're still important they still they still add value they can be delivering the vast majority that they can be creating the vast majority of the tests but as soon as those stories are released or you know the the developers have implemented the the behavior that fulfills those tests the development team own them and they have responsibility for fixing any breakage is that's the point I'm trying to me no it's it's one of those cultural things you need you need to win people over and get them the one thing I would say that the teams that I've worked with at work like this wouldn't go back so so the teams I've worked with this regularly grumble about how how difficult and onerous is to stay on top of these testes keep all the things working it's hard work it's difficult but the safety net that it provides is so valuable that they wouldn't go back to doing any other way thank you [Applause] you
- from
- Talks
- added
- 2026-10-10
- likes
- 0
Talks › Categories > Testing: “by Dave Farley (PIPELINE Conference 2015) [01:02:34]”