• Nederlands

Approach

CRISP-DM

CRISP-DM, Cross-Industry Standard Process for Data Mining, was developed in the late 1990s by a consortium of NCR, SPSS, Daimler-Benz and OHRA and published in 2000 as CRISP-DM 1.0. It describes a data science project as a cycle of six phases and has been the most widely used methodology in the field ever since. Twentynext works with the ten steps of IBM’s Foundational Methodology for Data Science, which builds on it.

Get in touch →
Framed print of a six-segment cycle diagram on a light office wall

Methodology

The six phases of CRISP-DM

The original CRISP-DM has six phases. In practice we work with the ten steps of IBM’s Foundational Methodology for Data Science: the same six phases, complemented by four steps that CRISP-DM does not name as phases of their own: Analytic Approach, Data Requirements, Data Collection and Feedback. And the arrows do not only run forward: six return loops are part of the model. Below you will find, for each phase, what happens, what it delivers and where it goes wrong in practice.

1

Business Understanding

step 1 of 10

The project does not start with the data but with the decision it is meant to improve. We establish which business goal the project serves, how success will be measured and who will work with the outcome. Skip this phase and the model can succeed while nothing changes in the organisation.

+ · steps 2–4 of 10 · added by the IBM methodology

Analytic Approach

Which kind of analysis fits the question: predicting, explaining, detecting or optimising. In the model this step comes before data exploration, while in practice you often only know which analysis fits after a first exploration. Knowing that, you plan the loop instead of being caught out by it.

Data Requirements

What data that analysis needs: sources, format, period and quality.

Data Collection

How that data comes in, what is missing and what may not be used. If collecting reveals that the requirements are off, they are revised: the loop back to Data Requirements.

2

Data Understanding

step 5 of 10

The first confrontation with the real data: how complete, how current and how reliable is it, and does it mean what everyone assumes it means? Assumptions fail in this phase, and that is the point: better in week three than after go-live. If data turns out to be missing, the loop takes you back to Data Collection.

3

Data Preparation

step 6 of 10

Usually the largest item in the project budget: joining sources, fixing errors, building features and recording every operation so it stays reproducible. What is done well here pays off in every later phase; what is missing here you pay for later, and then the loop back to Data Collection fetches new data after all.

4

Modeling

step 7 of 10

Only now does the modelling start. We train and compare several candidates against the measure from phase one, from simple to complex: the simplest model that answers the question wins. Modelling and preparation alternate: every intermediate result feeds back into data preparation. In our projects the model itself is rarely the risk; the phases before and after it are.

5

Evaluation

step 8 of 10

Two questions, in this order: does the solution answer the business question from phase one, and do the results hold up on data the model has not seen? This is where the decision is made: to production, back to Modeling via the return loop, or stop. Stopping is an outcome too; better here than in production.

6

Deployment

step 9 of 10

Where most data projects run aground: from working experiment to a system that runs in day-to-day operations, with monitoring, documentation and an owner. This is what Twentynext is built for: our Service & Maintenance practice keeps what we build running.

+ · step 10 of 10 · added by the IBM methodology

Feedback

After go-live the real measurement starts: is the system being used, are the predictions still right and is reality shifting under the model? The answers loop back to Modeling: adjust, redeploy, and the system does not quietly grow stale.

The order is not strict: as soon as the findings give reason to, you return to an earlier phase. What that looks like in practice is in our cases and projects.

CRISP-DM describes how we run a single data project. If you are looking for the route by which your whole organisation starts making decisions based on data, start at Data-driven working.

Overview

The six phases of CRISP-DM in one table

What each phase has to deliver according to the original methodology, CRISP-DM 1.0. The four steps Twentynext adds are described above, with the phases.

PhasePurposeDeliverables according to CRISP-DM 1.0
1. Business UnderstandingEstablish which business objective the project serves and how success will be measuredBusiness objectives and success criteria, situation assessment, data mining goals, project plan
2. Data UnderstandingCollect, describe, explore and quality-check the available dataInitial data collection report, data description, data exploration, data quality report
3. Data PreparationBuild the dataset the modelling will run onDataset and its description: selection, cleaning, derived attributes, integration, formatting
4. ModelingSelect techniques, build models and assess them against each otherModelling technique, test design, models, model assessment
5. EvaluationCheck whether the result answers the business question and whether the process was soundAssessment of results, approved models, process review, list of possible actions, decision on deployment
6. DeploymentPut the result into use and keep it in useDeployment plan, monitoring and maintenance plan, final report, project review

Literature

The original publications, for anyone who wants to read the methodology at its source.

  • Chapman, P., Clinton, J., Kerber, R., Khabaza, T., Reinartz, T., Shearer, C. & Wirth, R. (2000). CRISP-DM 1.0: Step-by-step data mining guide. SPSS Inc.
  • Wirth, R. & Hipp, J. (2000). CRISP-DM: Towards a standard process model for data mining. In Proceedings of the Fourth International Conference on the Practical Application of Knowledge Discovery and Data Mining, 29–39.
  • Shearer, C. (2000). The CRISP-DM model: The new blueprint for data mining. Journal of Data Warehousing, 5(4), 13–22.
  • Rollins, J. B. (2015). Foundational Methodology for Data Science. IBM Analytics.

Frequently asked questions

What is CRISP-DM?

CRISP-DM stands for Cross-Industry Standard Process for Data Mining. The methodology was developed in the late 1990s by a consortium of NCR, SPSS, Daimler-Benz and OHRA, published in 2000 as CRISP-DM 1.0 (Chapman et al., 2000), and has been the most widely used methodology for data science projects ever since. It describes a data project as a cycle of six phases, from understanding the business question to putting the solution into use.

Which six phases make up CRISP-DM?

Business Understanding, Data Understanding, Data Preparation, Modeling, Evaluation and Deployment. The order is not strict: as soon as the findings give reason to, you return to an earlier phase.

Why does Twentynext work with ten steps instead of six?

We follow IBM’s Foundational Methodology for Data Science, which builds on CRISP-DM. Four steps that sit inside CRISP-DM’s existing phases get a place of their own there: which kind of analysis the question calls for, which data that requires, how it comes in, and what happens to the results after delivery. Those four often decide the outcome.

Which projects is CRISP-DM suitable for?

Complex data projects of any size. The methodology starts from the business objective, only ends at implementation, supports several iterations and is not tied to any industry.

What happens after go-live?

That is when the Feedback step begins: measuring whether the system is used, whether the predictions still hold and whether reality is shifting under the model. Our Service & Maintenance practice is set up for exactly this: monitoring, maintenance and further development, so the system keeps performing in production.

Work with us

Realise your project together?

The people who build it also run it afterwards. Eindhoven, since 2014.

Martijn van Grieken

Martijn van Grieken

Director Data & AI

Get in touch