Find the C-arm position
Find where the C-arm stood, from a radiograph and a CT.
Step by step
-
Step 1
Read the three panels
Left is the radiograph to match. Middle is the CT rendered at the slider pose, a digitally reconstructed radiograph or DRR. Right overlays them.
-
Step 2
Press Register
The optimiser takes over and the middle image chases the left one. Stop halts it wherever it has got to.
-
Step 3
Watch it land
Grey overlay means matched. Fit gives the correlation and the true error; Reveal true pose prints the answer.
-
Step 4
Then try to break it
Drag a slider far off, or press Perturb pose, and register again. New radiograph deals a fresh unknown pose.
Try it
Starting the renderer…
Fit
- Cross-correlation
- 0.0000
- Iterations
- 0
- Rotation error
- 0.00°
- Translation error
- 0.0 mm
- Time
- —
Cross-correlation over the run
How it works
Each step on the left is one line of the loop. The panel beside it holds the detail, for anyone who wants the numbers.
-
Step 1
Load the CT as a 3D texture
The scan is uploaded to the graphics card once, and stays there.
138 by 124 by 184 voxels at 2.6 mm, one byte each. Hounsfield units become linear attenuation by μ = μwater (1 + HU / 1000), clamped at zero and quantised to eight bits, which is what the shader samples.
-
Step 2
Render a DRR at the current pose
One ray per pixel, from the source through the patient to the detector.
A digitally reconstructed radiograph, a DRR, is that line integral. A fragment shader intersects each ray with the volume box, marches it in 384 equal steps and accumulates μ × ds. This is the same forward model nanoDRR (opens in a new tab) implements in PyTorch, where it is differentiable end to end.
-
Step 3
Score it against the radiograph
One number says how well the two images agree.
Normalised cross-correlation over every pixel. Subtracting the means and dividing by the standard deviations makes the score blind to brightness and contrast, so only the anatomy lining up can raise it. 1.0 means identical.
-
Step 4
Estimate the gradient
Which way should each of the six numbers move?
There is no automatic differentiation in the browser, so central differences: nudge each parameter up and down by 0.25° or 0.5 mm. That is thirteen poses per iteration, drawn into one tiled framebuffer and read back in a single call, because reading pixels back from the GPU is the expensive part.
-
Step 5
Step the six numbers
Move downhill, then render again.
Adam, with the learning rate decaying from 0.9 to 0.06. Adam rescales each parameter by its own gradient history, which is what lets degrees and millimetres share a single step size without one of them dominating.
-
Step 6
Coarse first, then fine
Find roughly where it is, then sharpen.
130 iterations at 64 by 64 pixels, then 90 at 112 by 112. The coarse pass is cheap and smooths the score surface, which widens the range of starting poses that converge; the fine pass recovers the last fraction of a degree.
What this is not
The contribution of XVR is a neural network that regresses the pose directly, trained per patient in minutes, which removes the need for a good starting guess and makes the method fast enough for the operating room. That network is PyTorch and CUDA and does not run in a browser. What runs here is the differentiable-rendering half, the part xvr register uses to refine a pose, started from a deliberately perturbed one. Drag the sliders far enough and the failure mode the paper is about appears: the correlation surface has local maxima, and gradient descent alone settles into the wrong one.
Run it on your own scans
The demo above is fixed to one CT. To register your own data you need the real thing, on a GPU. This notebook installs xvr on a free Colab runtime, takes a CT volume and a radiograph you upload, and runs xvr register on them.
Open the notebook in Colab (opens in a new tab) Download the notebook
- What you need. A CT or MR volume as NIfTI (
.niior.nii.gz), one radiograph, and the geometry of the machine that took it: source-to-detector distance, detector pixel spacing and image size. Those are in the DICOM header of most fluoroscopy images. - Runtime. Set Runtime → Change runtime type → T4 GPU before running. It works on CPU, slowly.
- A starting guess.
xvr registeris a local optimiser, exactly like the demo above. It needs an initial pose in the right neighbourhood, or a patient-specific network fromxvr trainto supply one. The notebook covers both. - Your data leaves your computer. Colab runs on Google's machines. Do not upload identifiable patient data without the approvals that requires.
What you need to know
Three lectures from the study track cover what this page assumes. Each link opens the course there and starts on that lecture.
- Multiple View Geometry, lecture 2 (Daniel Cremers, TUM). Rigid-body motion and SE(3): the six numbers the optimiser solves for here are the ones this lecture builds.
- Multiple View Geometry, lecture 3 (Daniel Cremers, TUM). Perspective projection. A DRR is this projection with attenuation integrated along each ray, instead of a surface being sampled.
- Uncertainty in registration with MCMC (William M. Wells III, MISS 2016). Registration posed as an optimisation over a similarity measure, which is what Register does on this page with normalised cross-correlation.
Credit
- Method. Gopalakrishnan V and colleagues. Rapid patient-specific neural networks for X-ray to volume registration. Nature, 2026. doi:10.1038/s41586-026-11045-x (opens in a new tab)
- Reference implementation. eigenvivek/xvr (opens in a new tab), MIT licence. The differentiable renderer behind this forward model is eigenvivek/nanodrr (opens in a new tab), also MIT, by the same author. This page is an independent browser reimplementation of the forward model and the optimisation step, not a port of either.
- CT scan. Task03_Liver, case 100, from the Medical Segmentation Decathlon (opens in a new tab), licensed CC BY-SA 4.0 (opens in a new tab), obtained through the 3D Slicer sample data collection. Cropped to the body, resampled to 2.6 mm isotropic and quantised to 8 bits for the web; the derived volume is shared under the same licence.
Questions
What is a DRR?
A digitally reconstructed radiograph: an X-ray image simulated from a CT scan rather than taken with a machine. Send a ray from the source to each detector pixel, add up the density it crosses on the way through the CT, and that sum is the pixel. The middle panel above is a DRR, redrawn every time you move a slider.
What is a voxel?
A pixel with a third dimension. A CT scan is a stack of slices, and a voxel is one small box of that stack holding a single number: how strongly that box absorbs X-rays. The volume on this page is 138 by 124 by 184 voxels, each 2.6 mm on a side, which is about 3.1 million of them.
What is PyTorch?
A Python library for building and training neural networks, and the one most medical imaging research is written in. Two things about it matter here. It runs the arithmetic on a graphics card, and it does automatic differentiation: hand it a calculation and it works out the derivative of the answer with respect to any input, which is what lets a pose be optimised directly. xvr and nanoDRR are PyTorch. This page is not, which is why it has to estimate its gradients by nudging each parameter and re-rendering.
What is a cross-correlation?
A single number saying how similar two images are, from −1 to 1. Subtract each image's mean brightness, multiply the two images together pixel by pixel, add up the result and divide by both standard deviations. That normalising is the useful part: it makes the score ignore overall brightness and contrast, so a simulated DRR can be compared with a real radiograph even though the two are made in completely different ways. 1 means identical, and the optimiser above climbs towards it.