Find the C-arm position

Find where the C-arm stood, from a radiograph and a CT.

Step by step

  1. Step 1

    Read the three panels

    Left is the radiograph to match. Middle is the CT rendered at the slider pose, a digitally reconstructed radiograph or DRR. Right overlays them.

  2. Step 2

    Press Register

    The optimiser takes over and the middle image chases the left one. Stop halts it wherever it has got to.

  3. Step 3

    Watch it land

    Grey overlay means matched. Fit gives the correlation and the true error; Reveal true pose prints the answer.

  4. Step 4

    Then try to break it

    Drag a slider far off, or press Perturb pose, and register again. New radiograph deals a fresh unknown pose.

Try it

Fixed radiographthe image to match
Moving DRRrendered at the current pose
Overlaygrey where the two agree

Starting the renderer…

Pose, six degrees of freedom

0.0°
0.0°
0.0°
0.0 mm
0.0 mm
0.0 mm

True pose: hidden

Fit

Cross-correlation
0.0000
Iterations
0
Rotation error
0.00°
Translation error
0.0 mm
Time
—

Cross-correlation over the run

How it works

Each step on the left is one line of the loop. The panel beside it holds the detail, for anyone who wants the numbers.

  1. Step 1

    Load the CT as a 3D texture

    The scan is uploaded to the graphics card once, and stays there.

    138 by 124 by 184 voxels at 2.6 mm, one byte each. Hounsfield units become linear attenuation by μ = μwater (1 + HU / 1000), clamped at zero and quantised to eight bits, which is what the shader samples.

  2. Step 2

    Render a DRR at the current pose

    One ray per pixel, from the source through the patient to the detector.

    A digitally reconstructed radiograph, a DRR, is that line integral. A fragment shader intersects each ray with the volume box, marches it in 384 equal steps and accumulates μ × ds. This is the same forward model nanoDRR (opens in a new tab) implements in PyTorch, where it is differentiable end to end.

  3. Step 3

    Score it against the radiograph

    One number says how well the two images agree.

    Normalised cross-correlation over every pixel. Subtracting the means and dividing by the standard deviations makes the score blind to brightness and contrast, so only the anatomy lining up can raise it. 1.0 means identical.

  4. Step 4

    Estimate the gradient

    Which way should each of the six numbers move?

    There is no automatic differentiation in the browser, so central differences: nudge each parameter up and down by 0.25° or 0.5 mm. That is thirteen poses per iteration, drawn into one tiled framebuffer and read back in a single call, because reading pixels back from the GPU is the expensive part.

  5. Step 5

    Step the six numbers

    Move downhill, then render again.

    Adam, with the learning rate decaying from 0.9 to 0.06. Adam rescales each parameter by its own gradient history, which is what lets degrees and millimetres share a single step size without one of them dominating.

  6. Step 6

    Coarse first, then fine

    Find roughly where it is, then sharpen.

    130 iterations at 64 by 64 pixels, then 90 at 112 by 112. The coarse pass is cheap and smooths the score surface, which widens the range of starting poses that converge; the fine pass recovers the last fraction of a degree.

What this is not

The contribution of XVR is a neural network that regresses the pose directly, trained per patient in minutes, which removes the need for a good starting guess and makes the method fast enough for the operating room. That network is PyTorch and CUDA and does not run in a browser. What runs here is the differentiable-rendering half, the part xvr register uses to refine a pose, started from a deliberately perturbed one. Drag the sliders far enough and the failure mode the paper is about appears: the correlation surface has local maxima, and gradient descent alone settles into the wrong one.

Run it on your own scans

The demo above is fixed to one CT. To register your own data you need the real thing, on a GPU. This notebook installs xvr on a free Colab runtime, takes a CT volume and a radiograph you upload, and runs xvr register on them.

Open the notebook in Colab (opens in a new tab) Download the notebook

What you need to know

Three lectures from the study track cover what this page assumes. Each link opens the course there and starts on that lecture.

Credit

Questions

What is a DRR?

A digitally reconstructed radiograph: an X-ray image simulated from a CT scan rather than taken with a machine. Send a ray from the source to each detector pixel, add up the density it crosses on the way through the CT, and that sum is the pixel. The middle panel above is a DRR, redrawn every time you move a slider.

What is a voxel?

A pixel with a third dimension. A CT scan is a stack of slices, and a voxel is one small box of that stack holding a single number: how strongly that box absorbs X-rays. The volume on this page is 138 by 124 by 184 voxels, each 2.6 mm on a side, which is about 3.1 million of them.

What is PyTorch?

A Python library for building and training neural networks, and the one most medical imaging research is written in. Two things about it matter here. It runs the arithmetic on a graphics card, and it does automatic differentiation: hand it a calculation and it works out the derivative of the answer with respect to any input, which is what lets a pose be optimised directly. xvr and nanoDRR are PyTorch. This page is not, which is why it has to estimate its gradients by nudging each parameter and re-rendering.

What is a cross-correlation?

A single number saying how similar two images are, from −1 to 1. Subtract each image's mean brightness, multiply the two images together pixel by pixel, add up the result and divide by both standard deviations. That normalising is the useful part: it makes the score ignore overall brightness and contrast, so a simulated DRR can be compared with a real radiograph even though the two are made in completely different ways. 1 means identical, and the optimiser above climbs towards it.