Benchmarks

This page compares aperture-sum performance, including the effect of validate=False in the object API. Times and speedup ratios depend on the machine, input layout, and software versions. See Performance for API-selection and reuse guidance.

Aperture-sum comparisons

The repository benchmark script validates numerical agreement before it prints timings. This page runs the same script with --repeats 5 and shows a focused subset: float32 exact aperture sums for circles, ellipses, rectangles, and pills at 1, 100, and 10,000 positions.

The benchmark labels are:

Label Timed operation
aap Construct an aperture with validate=True, then sum the image.
aapv The same construction and sum with validate=False.
aapr Optimized Python kernel path with validate=False.
pg photutils.geometry overlap kernels where available.
pa Photutils aperture construction and photometry.
sep SEP aperture sums where available.

Compare aap/aapr with aapv/aapr to measure the effect of validation. Their quotient is aap_time / aapv_time: values 1.8x and 1.2x, for example, mean the object workflow is 1.5x faster with validation disabled. These are construction-plus-sum timings; they do not measure constructor overhead alone.

For scale, earlier local CircAp constructor measurements showed roughly 2× faster construction for a single center: 1.66 µs → 0.75 µs, a saving of about 0.9 µs. A 10,000-center C-order case improved by only 1.3×. Construction was already measured in microseconds, so the full-workflow ratios below need not approach 2×. These illustrative timings are machine- and input-dependent; see construction timing context.

Both object paths receive the same prebuilt C-order (N, 2) coordinates and image, construct a fresh aperture, and call apsum_exact(return_npix=False). Both still copy coordinates into owned float64 storage. validate=False skips applicable checks in construction and the subsequent operation; use it only when inputs are already valid. Numerical agreement is checked before timing. A ratio close to 1x can reflect timing noise or little validation cost.

from io import StringIO
from pathlib import Path
import subprocess
import sys

import pandas as pd


for parent in [Path.cwd(), *Path.cwd().parents]:
    benchmark_script = parent / "benchmarks" / "benchmark_apertures.py"
    if benchmark_script.exists():
        break
else:
    raise FileNotFoundError("could not find benchmarks/benchmark_apertures.py")

benchmark_cmd = [
    sys.executable,
    str(benchmark_script),
    "--repeats",
    "5",
    "--format",
    "csv",
    "--tasks",
    "apsum",
    "--dtypes",
    "float32",
    "--shapes",
    "circle",
    "ellipse",
    "rectangle",
    "pill",
    "--counts",
    "1",
    "100",
    "10000",
]

benchmark_output = subprocess.run(
    benchmark_cmd,
    check=True,
    capture_output=True,
    text=True,
).stdout

benchmark_lines = [
    line for line in benchmark_output.splitlines()
    if line and not line.startswith("#")
]
benchmark_csv = "\n".join(
    line for line in benchmark_lines
    if not line.startswith("task,shape,dtype,n_apertures,fastest_library")
)
benchmark_raw = pd.read_csv(StringIO(benchmark_csv))
benchmark_raw = benchmark_raw[benchmark_raw["speedup_vs_library"].isna()].copy()
benchmark_raw["library"] = benchmark_raw["library"].replace(
    {
        "photutils.geometry": "pg",
        "photutils.Aperture": "pa",
    }
)

benchmark_summary = (
    benchmark_raw
    .pivot_table(
        index=["shape", "n_apertures"],
        columns="library",
        values="seconds",
        aggfunc="first",
    )
    .reset_index()
)

for column in ["aapr", "aap", "aapv", "pg", "pa", "sep"]:
    if column in benchmark_summary:
        benchmark_summary[f"{column}_us"] = benchmark_summary[column] * 1.0e6

for column in ["aap", "aapv", "pg", "pa", "sep"]:
    if column in benchmark_summary:
        benchmark_summary[f"{column}/aapr"] = (
            benchmark_summary[column] / benchmark_summary["aapr"]
        )

display_columns = [
    "shape",
    "n_apertures",
    "aapr_us",
    "aap/aapr",
    "aapv/aapr",
    "pg/aapr",
    "pa/aapr",
    "sep/aapr",
]
display_columns = [column for column in display_columns if column in benchmark_summary]
benchmark_table = benchmark_summary[display_columns].copy()

for column in benchmark_table.columns:
    if column.endswith("_us"):
        benchmark_table[column] = benchmark_table[column].map(lambda value: f"{value:.1f}")
    elif column.endswith("/aapr"):
        benchmark_table[column] = benchmark_table[column].map(
            lambda value: "--" if pd.isna(value) else f"{value:.1f}x"
        )

benchmark_table
library shape n_apertures aapr_us aap/aapr aapv/aapr pg/aapr pa/aapr sep/aapr
0 circle 1 21.5 2.3x 1.4x 1.6x 5.6x 1.4x
1 circle 100 150.5 1.2x 1.3x 8.9x 15.4x 2.0x
2 circle 10000 8205.8 1.1x 1.1x 14.6x 25.4x 3.0x
3 ellipse 1 41.2 1.6x 1.2x 2.5x 7.0x 1.5x
4 ellipse 100 376.7 1.3x 1.1x 12.4x 16.2x 2.4x
5 ellipse 10000 27980.5 1.0x 1.0x 15.8x 20.5x 2.8x
6 pill 1 78.5 2.4x 1.0x -- -- --
7 pill 100 2067.9 1.1x 1.1x -- -- --
8 pill 10000 195541.7 1.0x 1.0x -- -- --
9 rectangle 1 30.1 1.5x 1.7x -- 12.4x --
10 rectangle 100 178.2 1.1x 1.1x -- 103.3x --
11 rectangle 10000 8839.5 1.1x 1.0x -- 204.4x --

Times and ratios are hardware-dependent. In the table, values larger than 1x mean that backend was slower than aapr. The script interleaves timing order, validates results before reporting times, and drops one fastest and one slowest sample when at least three samples are available. For very large Photutils rows the script may reduce the effective Photutils repeat count to keep the benchmark from dominating docs build time; this is reported in the benchmark script notes.