Breast Cancer Detection with Machine Learning: A Classifier Running Live
Altius Digital 5 min read
Give a model thirty measurements of a tumour and it will tell you, with 97.8% accuracy, whether it is malignant or benign. This is one of our own research projects, and still one of our favourite demonstrations of what a small, well-built model can do.
The model above is real: it runs in your browser and scores genuine records from the dataset, one after another.
The data
The Wisconsin Diagnostic Breast Cancer dataset from the UCI Machine Learning Repository holds 569 cases. For each one, measurements were computed from a digitised image of a fine-needle aspirate of a breast mass: the radius, texture, perimeter, area, smoothness, compactness, concavity, concave points, symmetry and fractal dimension of the cell nuclei, each as a mean, a standard error and a “worst” value. Every case is labelled by a pathologist as malignant (212) or benign (357).
The model
The approach is deliberately simple. Each of the 30 measurements is standardised, a logistic regression learns a weight for each, and the weighted sum is squashed into a probability of malignancy. Simple models are easy to inspect: the “worst” radius, texture and concave points carry the most weight, which matches what a pathologist looks for.
We tested feature sets and regularisation strengths against each other, and measured everything with 10-fold cross-validation repeated five times, so the model is always scored on cases it has not seen:
- 97.8% accuracy (±1.9%) on the 569 records
- 0.996 area under the ROC curve, so it ranks malignant cases above benign ones almost perfectly
Ten measurements alone reach about 94%; the full 30 are worth the extra complexity. A non-linear model (a support vector machine) scored about the same, so the simpler, explainable model won.
The machine learning behind it, in plain English
Supervised learning
This is a supervised classification problem. We have examples where the answer is already known (each case has a pathologist’s diagnosis) and we want a function that maps measurements to that answer for cases it has never seen. The measurements are the features, the diagnosis is the label, and the 569 labelled cases are the training data.
Standardisation
Area is measured in the hundreds; smoothness is a fraction below one. Fed raw into a model, the big numbers would drown out the small ones. So every feature is standardised: we subtract its mean and divide by its standard deviation, so each one sits on the same scale, centred on zero. The model then learns which features matter, not which ones happen to have large units.
Logistic regression
Logistic regression is the workhorse of classification. It learns one weight per feature plus an intercept, multiplies each standardised measurement by its weight and adds them up. That sum, called the logit, is then passed through the sigmoid function, an S-shaped curve that turns any number into a value between 0 and 1. The result is read as the probability that the tumour is malignant. Above 0.5 we say malignant, below it benign. The dial in the demo shows exactly this probability.
Because the decision is a weighted sum, you can read the weights. A large positive weight on “worst concave points” means more concave points push the prediction towards malignant. That transparency is why the model is popular in medicine and finance, where you have to be able to explain a decision.
Training: maximum likelihood and gradient descent
Training means finding the weights that make the known answers most likely. The model starts with random weights, measures how wrong it is on every training case using the log loss (also called cross-entropy, which punishes confident wrong answers heavily), and nudges every weight in the direction that reduces that loss. Repeat until the loss stops falling. This is gradient descent, and it is the same idea that trains neural networks, just with far fewer weights.
Regularisation and overfitting
With 30 features and 569 cases, a model can overfit: learn quirks of the training cases that do not hold in general. L2 regularisation fights this by adding a penalty for large weights, so the model prefers many small, sensible weights over a few extreme ones. The strength is set by a parameter called C. We tried several values; C = 0.3 gave the best cross-validated results, which is the model you see here.
Cross-validation: measuring honestly
A model will always look good on the data it was trained on. To measure it properly we use k-fold cross-validation: split the cases into 10 groups, train on 9 and test on the 10th, rotate so every group is tested once, then average. We repeated that five times with different splits, and kept the proportion of malignant and benign cases the same in every group (stratified folds). Every figure on this page is an out-of-sample figure.
Beyond accuracy: the errors that matter
Accuracy alone hides which mistakes a model makes, and in medicine the two kinds are not equal. Missing a cancer (a false negative) is far worse than flagging a benign tumour for a second look (a false positive). Across the cross-validation, at the standard 0.5 threshold, the model:
- caught 201 of 212 malignant cases (sensitivity 94.8%) and missed 11
- correctly cleared 355 of 357 benign cases (specificity 99.4%)
Lower the threshold to 0.3 and it catches 206 of 212 (sensitivity 97.2%) at the cost of 13 false alarms instead of 2. A real clinical system would set that threshold with clinicians, not by maximising accuracy. The ROC curve plots exactly this trade-off across every threshold, and the 0.996 area under it says the model separates the two classes almost perfectly whichever point you choose.
Why not a neural network?
We also trained a support vector machine with a radial-basis kernel, a model that can draw curved boundaries between classes. It scored within a fraction of a percent of the logistic regression. When a simple, explainable model matches a complex one, the simple one wins: it is easier to audit, cheaper to run (the demo’s arithmetic is a 30-term sum) and harder to fool. Deep learning earns its place with images, text and millions of examples, not with 569 rows of clean numbers.
What you are seeing above
The demo picks real records from the dataset, draws their measurements, runs the model and shows its verdict next to the pathologist’s diagnosis. The weights were trained with scikit-learn and exported; the arithmetic in your browser is the model itself, not a recording.
What it is not
This is a research project and a demonstration of method. It is not a medical device and it is not how cancer is diagnosed. Real diagnostic systems are built under clinical governance, with far larger and more diverse data, and they support clinicians rather than replace them.
What the project does show is the discipline that matters in any machine-learning work we do for clients: clean data, a model no more complex than it needs to be, honest measurement, and a clear line between what the model knows and what it does not.