Research papers, technical documents, catalogs, and older reports often preserve a graph even when the original data is gone. If you want to reuse a curve in Excel or recover XY coordinates from a PDF graph, reading every value by eye quickly becomes tedious.

Browser Kitty Graph Digitizer calibrates the axes of a graph image or PDF page, converts image positions into XY values, and exports CSV. It also supports colored-curve tracing, marker extraction, multiple series, Linear / log10 axes, project saving, and local processing of the selected files.

How graph digitizing converts image positions into numbers

A graph image normally does not contain metadata saying that a particular pixel means X=10 and Y=50. The first step is therefore to identify known tick marks on the X and Y axes.

For example, after registering the image positions of X=0, X=10, Y=0, and Y=100, the app can calculate a mapping between image coordinates and graph values. Clicked or automatically detected positions can then be converted into XY coordinates.

Values from Graph Digitizer are estimates calculated from information that remains in the image. A low-resolution graph cannot recreate digits or measurement precision that were lost when the original data became an image.

Load an image or PDF and crop to the graph when needed

Supported image inputs are PNG, JPEG, and WebP, using file selection, drag and drop, or clipboard paste. For PDFs, choose a page and optionally crop the graph region before committing it as the working image.

If the PDF page also contains body text, tables, multiple figures, notes, or legends, cropping to the graph makes calibration easier, simplifies color picking, and narrows the auto-trace area.

Opening a PDF does not immediately replace the current work. Confirm the page and crop region before switching the working image.

  • PNG
  • JPEG
  • WebP
  • PDF pages

Calibrate the axes with X1 / X2 / Y1 / Y2

To map image positions to graph values, place four calibration points: X1, X2, Y1, and Y2, using tick marks whose values are known. For example, X=0, X=10, Y=0, and Y=100 provide the reference needed to convert later positions into XY values.

Calibration points do not have to be the origin or the axis intersection. If X=20 and X=40 are clearer, use those instead. What matters is knowing exactly which number corresponds to each selected image position.

When possible, choose calibration ticks that are reasonably far apart. That reduces the relative impact of a one-pixel placement error.

  • X1
  • X2
  • Y1
  • Y2

Check Linear / log10 scales, reversed axes, and slight rotation

Graph axes are not always linear. If 1, 10, 100, and 1000 are equally spaced, the axis is logarithmic. Graph Digitizer lets X and Y use Linear or log10 independently, so semi-log and log-log plots are supported.

Treating a logarithmic axis as Linear changes the conversion formula and can produce large errors. Check whether ticks are equally spaced by value or equally spaced by powers of ten before digitizing.

Reversed axes are also supported by entering calibration values in the displayed direction. Straight XY axes can be calibrated even when a scanned graph is slightly rotated.

The app does not automatically correct perspective distortion, lens distortion, curved paper, or genuinely curved axes.

  • Linear
  • log10

Digitize small datasets manually and repair point placement precisely

For a graph with only a few points, manual digitizing is often the most reliable choice. Click the position you want to read and the app converts it into XY values for the current series.

Manual work is also useful for complex curves, similar series colors, strong grid lines, areas where auto-tracing struggles, or cases where only a few reference values are needed.

A selected point can be repaired with a loupe and one-pixel adjustments. On desktop, arrow keys move one pixel and Shift + arrow keys move ten. On mobile, place a temporary point, refine it under the loupe, and then confirm it.

Automatically trace a colored curve

For dense line plots or continuous curves, use Curve mode. Choose the curve color and the app searches along the X direction for matching pixels to generate candidate points.

You can tune the color, target region, optional start point, tolerance, continuity, and sampling interval. The search region can be the calibrated axis area or a custom rectangle.

Auto-traced points are not committed immediately. They appear in a preview so you can verify the detected curve, adjust settings, or remove false detections before applying the result to the series.

Extract visible plot markers as data points

In some line charts, the real observations are the centers of circles or squares rather than every point on the connecting line. Marker mode is intended for this case.

Curve mode follows a colored line along X, while Marker mode looks for locally thicker shapes such as visible circles or squares and returns their centers as data points.

If a legend, label, grid intersection, or another series is detected by mistake, select the false point in the preview and remove it with Delete / Backspace before applying the result.

Keep long gaps as separate segments and manage multiple series

When a curve contains a long missing section, Graph Digitizer keeps the separated pieces as different segments rather than silently interpolating through the gap. It avoids inventing values where no line is visible.

Graphs with red, blue, green, or other multiple series can store each series name, display color, and extracted points separately. Names such as Experiment A and Experiment B remain useful after CSV export.

Crossings, similar colors, overlapping lines, and grid lines close to a series color can cause tracking errors. Review and repair those areas manually against the source image.

Review results with source overlays and a redrawn graph

In graph digitizing, verifying the result is as important as extracting it. Graph Digitizer can overlay the extracted points on the source image so points that drift away from the curve or jump to another series are easier to spot.

It can also redraw a graph from the extracted XY data so you can compare the overall curve shape, peak positions, slope, and missing sections.

The source overlay, redrawn graph, review list, and coordinate table share the same selected point, making it easier to inspect and repair suspicious values.

Save unfinished work as a .graphdigitizer.json project

If calibration and multi-series extraction cannot be finished in one session, save the current project as .graphdigitizer.json. The project retains the working image, calibration, series, segments, extracted points, tracing settings, and CSV settings so the work can be resumed later.

When the source is a PDF, the entire original PDF is not embedded in the project. The project stores the working-page image and other information needed to resume the digitizing session.

Preview and export CSV for reuse in Excel and other tools

The standard CSV uses series,segment,x,y so multiple series and separated curve segments remain distinguishable. When there is only one series and one segment, a simpler x,y format is also available.

Before saving, preview the actual CSV text and verify the series names, columns, values, and segment numbers. You can also copy the CSV text directly into another application.

The saved CSV can be reused in Excel, Python, R, and other tools. Replotting X and Y can reproduce the graph shape, but the values remain estimates from image positions rather than the original raw measurements.

CSV format options
FormatUse
series,segment,x,yWhen multiple series or segments need to remain distinguishable
x,ySimple data with one series and one segment

It does not recover embedded source data or provide general chart OCR

Graph Digitizer does not inspect a PDF to recover hidden source tables or original vector data. It renders the PDF page and converts positions on the visible graph into XY values using calibration.

The main v1.0.0 target is an ordinary XY plot where each X position corresponds to roughly one Y value. It does not provide dedicated automatic interpretation for bar charts, histograms, polar or ternary plots, 3D charts, or filled areas.

It is also not a general chart-OCR system that recognizes axis labels or category names such as January and February. A category axis can instead be calibrated numerically, for example January through December as X=1 through X=12.

  • Bar charts
  • Histograms
  • Polar / ternary plots
  • 3D charts
  • Filled-area interpretation
  • General chart OCR

When auto-tracing struggles, improve the source and mix in manual repair

Curve tracing works best when the target color and shape are clearly separated from the rest of the image. JPEG artifacts, low resolution, very thin lines, strong grids, crossings, similar colors, and weak contrast can all make tracing harder.

When possible, prefer the original PDF, PNG, or a high-resolution source over a screenshot that has been repeatedly saved as JPEG. Clearer lines and tick marks also make calibration and manual point placement easier.

Not every section needs to be automated. Trace long clean regions automatically and repair crossings or missing points manually. Treat auto-tracing as candidate generation that reduces manual work rather than as unquestioned final data.

Image, PDF, and project-file size limits

In v1.0.0, images are limited to 20 MiB and 16 million decoded pixels, PDFs to 80 MiB with working-page rasters capped at 16 million pixels, and .graphdigitizer.json project files to 96 MiB.

A large scan can exceed the decoded-pixel limit even when its compressed file size is below 20 MiB. Very large PDFs and high-resolution images can also consume substantial device memory.

Main v1.0.0 input limits
InputLimit
Images20 MiB / 16 million decoded pixels
PDF80 MiB / 16 million pixels for the working page
Project file96 MiB

Images, PDFs, and extracted coordinates are processed locally

Graph Digitizer processes selected images and PDFs in the browser. The generated HTML uses a Content Security Policy containing connect-src 'none', with no runtime CDN, analytics, telemetry, external AI API, or remote model. PDF.js, its worker, and required WASM are embedded.

Opening the hosted version on Browser Kitty or GitHub Pages requires an initial request for the app HTML. The app does not then upload the selected images, PDFs, project files, or extracted coordinates.

For a session with the network completely disconnected, a verified standalone HTML build can also be opened locally.

Respect the precision available in the image and use the data accordingly

For better digitizing accuracy, use a clear source image or PDF, place calibration points carefully, choose Linear / log10 correctly, and review auto-traced points against the source overlay. Do not trust more numerical digits than the image can actually support.

Typical uses include recovering reference values from a paper, converting an old report curve to CSV, digitizing a product performance chart, re-analyzing results that survive only as a PDF graph, or redrawing archived graphs in Excel.

When using data extracted from papers or other third-party material, separately check copyright, usage terms, and any citation requirements that apply to the source.

Step by step

  1. Load the graph image or PDF
  2. Crop to the graph region when needed
  3. Calibrate X1 / X2 / Y1 / Y2
  4. Confirm Linear / log10 for each axis
  5. Create the series you need
  6. Digitize with Manual, Curve, or Marker mode
  7. Review the result on the source-image overlay
  8. Repair review-needed points
  9. Check the graph redrawn from the extracted values
  10. Preview the CSV and save it
Try it in Browser Kitty

Graph Digitizer

Calibrate graph axes, extract and repair XY data from images or PDF pages, and export the reviewed values as CSV.

Open toolView tool details

Tips and limitations

  • Use the clearest available source image or PDF
  • Place calibration points precisely on known tick positions
  • Do not confuse Linear and log10 axes
  • Use well-separated calibration ticks when possible
  • Do not accept auto-traced results without review
  • Check point positions against the source-image overlay
  • Fine-tune important points with the loupe
  • Do not trust more digits than the image can support

Frequently asked questions

Can I extract numerical data from a graph image?

Yes. Calibrate the X and Y axes with known tick values, then convert points or curve positions in the image into XY values.

Can I convert a graph in a PDF to CSV?

Yes. Choose the PDF page and graph region, digitize the data, and save the extracted values as CSV.

Does it support logarithmic graphs?

Yes. X and Y can each be set to Linear or log10, supporting both semi-log and log-log plots.

Can it automatically trace a curve?

Yes. Choose the curve color and tune tolerance, continuity, sampling interval, and other settings. The detected points are previewed before they are applied to a series.

Can it extract only the visible markers from a line chart?

Yes. Marker mode detects visible circles, squares, and similar plot markers and returns their centers as data points.

Can I remove incorrectly detected points?

Yes. Remove false detections from the preview with Delete / Backspace, and edit or delete points after they have been added to a series.

Does it support bar charts?

v1.0.0 primarily targets ordinary XY plots and does not include dedicated automatic analysis for bar charts or histograms.

Can it OCR text or month names from the graph?

It does not provide general chart OCR. For a category axis such as January through December, calibrate numerical positions such as X=1 through X=12 and replace them with labels later if needed.

Can it recover exactly the same values as the original dataset?

No guarantee. The values are estimates calculated from image positions, and precision that was lost when the data became an image cannot be restored.

Can I save unfinished work?

Yes. Save a .graphdigitizer.json project containing calibration, series, points, segments, auto-trace settings, and other state so the work can be resumed later.

What CSV format is used?

The standard format is series,segment,x,y. For one series with one segment, a simpler x,y format is also available.

Are graph images or PDFs uploaded to a server?

No. Image/PDF loading, PDF rendering, calibration, auto-tracing, coordinate calculations, and CSV generation run in the browser. The hosted version only needs the initial request for the application HTML.