Linear Regression Calculator
Fit a line or curve to x, y data: equation, R², standard errors, intervals and residuals.
Try before you buy.
- Free preview: a watermarked scatter plot with its fitted curve and the names of each table’s first rows (up to 3), with the equation, the working and every figure hidden.
- Locked until you unlock it: download and copy.
- Unlock: Pro pass, ₹179 for 30 days, a one-time payment that never renews.
Ways to unlock shows how to get the full result.
Printing this result is locked in the free preview.
Result
This is the last result worked out. Fix the input above to update it.
Locked in the free preview. Opens the ways to unlock this result.
Locked in the free preview. Batch runs unlock with a pass.
Locked in the free preview. Query results unlock with a pass.
About the Linear Regression Calculator
Paste your (x, y) data and fit a straight line, a polynomial of degree 2 to 6, an exponential y = a·e^(bx), a logarithmic y = a + b·ln x or a power curve y = a·x^b by least squares. You get the equation, every coefficient with its standard error, t-test and confidence interval, R² and adjusted R², the residual standard error, the F test and the ANOVA table, and predictions at any x with a confidence interval for the mean response and a prediction interval for a new observation.
The scatter plot shows the fitted curve with its confidence and prediction bands, and a residual plot with a table of residuals, standardized residuals and leverages shows whether the model really fits. The fit uses a numerically stable QR decomposition instead of the textbook normal equations; it reproduces the certified results of the NIST Statistical Reference Datasets for linear regression, including the difficult degree-5 Wampler problems.
How to use it
- Paste two columns copied from a spreadsheet, or type one pair per line such as
3.2, 7.9. A first row of names labels the axes; rows with a missing value are left out. - Choose the model: a straight line, a polynomial (pick its degree), or an exponential, logarithmic or power curve.
- Optionally type x values to predict at, and change the confidence level of the intervals (95 % by default).
- Read the equation, R² and the coefficients, then check the residual plot: a good model leaves no pattern.
- Copy the results or download a CSV file or the chart — with a Pro pass or after unlocking this result; without one you see the free preview.
Examples
The ozone-monitor calibration data of the NIST Statistical Reference Datasets
y = 1.00212x − 0.262323; R² = 0.99999375; s = 0.884796
NIST’s certified values: b₀ = −0.262323073774029 (SE 0.232818234301152), b₁ = 1.00211681802045 (SE 0.000429796848199937).
The load-cell calibration data, polynomial of degree 2
y = −3.16082 × 10⁻¹⁵x² + 7.32059 × 10⁻⁷x + 0.000673566; R² = 0.9999999
(10, 8.04), (8, 6.95), (13, 7.58), (9, 8.81), (11, 8.33), (14, 9.96), (6, 7.24), (4, 4.26), (12, 10.84), (7, 4.82), (5, 5.68)
y = 0.500091x + 3.00009; R² = 0.666542, r = 0.816421
Anscombe built four very different data sets with this same line and R²: always plot your data.
(1, 2), (2, 4), (3, 5), (4, 4), (5, 5), predict at x = 3
y = 0.6x + 2.2, R² = 0.6; at x = 3: ŷ = 4, 95 % CI [2.7, 5.3], prediction interval [0.9, 7.1]
Common uses
- Finding the line of best fit and its equation for lab data, sales, costs or any two related measurements.
- Calibration curves: fitting instrument readings to reference values and predicting with honest intervals.
- Fitting growth or decay (exponential), diminishing returns (logarithmic) and scaling laws (power).
- Checking a regression from coursework or software, including the standard errors and the ANOVA table.
How the fit is computed
Least squares chooses the coefficients that make the sum of squared residuals Σ(y − ŷ)² as small as possible (NIST/SEMATECH e-Handbook). For a straight line the slope is b = Sxy/Sxx and the intercept a = ȳ − b·x̄. For every model the calculator solves the problem with Householder QR on x centred and scaled, then converts the coefficients and their covariance s²(XᵀX)⁻¹ back to powers of x. That avoids the loss of digits of the normal equations: on the NIST Statistical Reference Datasets it agrees with the certified values to 12 or more significant digits for Norris and Pontius and to at least 7 for the degree-5 Wampler sets.
The exponential and power models become straight lines on a log scale — ln y = ln a + bx and ln y = ln a + b·ln x — and are fitted there, as most software does; the logarithmic model is a straight line in ln x.
R², standard errors and the F test
- R² = 1 − SSE/SST is the share of the variation of y the model explains; adjusted R² = 1 − (1 − R²)(n − 1)/(n − k) penalizes extra coefficients, so it is the one to compare models of different degrees by.
- The residual standard error s = √(SSE/(n − k)) is the typical distance of a point from the curve, in the units of y.
- Each coefficient has a standard error from s²(XᵀX)⁻¹, a t-test of whether it is 0 with n − k degrees of freedom, and a confidence interval estimate ± t·SE.
- The F test compares the model with a flat line: F = (SSR/(k − 1))/(SSE/(n − k)).
Confidence and prediction intervals
At a value x₀ the confidence interval ŷ ± t·s·√h₀ says where the average y at x₀ lies, and the prediction interval ŷ ± t·s·√(1 + h₀) where one new observation will fall, where h₀ is the leverage of x₀ (e-Handbook: average response, single response). The prediction interval is always wider, and both widen away from the centre of the data. For the exponential and power models the intervals are worked out on the log scale and turned back, so they are not symmetric around the curve.
Checking the fit with residuals
R² alone cannot show that a model is right — Anscombe’s four data sets share the same line and R² yet look completely different. Look at the residual plot: random scatter around 0 means the model captures the pattern; a curve means a missing term (try a polynomial or a transformed model); a funnel means the spread changes with x. Standardized residuals beyond about ±2 (red) or ±3 point to outliers, and a leverage much larger than the average k/n marks a point that pulls the fit strongly.
Limitations
- Up to 100,000 points; polynomials up to degree 6. The residual table on the page lists the first 1,000 points, and the CSV all of them.
- Least squares assumes independent errors with constant spread; for data in time order or with errors that grow with x, the standard errors can be too small.
- High-degree polynomials follow the noise and behave wildly outside the data; prefer the lowest degree that fits.
- The exponential and power models minimize the squared errors of ln y, not of y; their R² and residuals are on that scale.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.
Frequently asked questions
What do I get without a pass?
Without a pass, Linear Regression Calculator shows a watermarked scatter plot with its fitted curve and the names of each table’s first rows (up to 3), with the equation, the working and every figure hidden. Until you unlock it, the result can’t be downloaded or copied. A Pro, Premium or Ultimate pass, a one-time payment that never renews, unlocks the full result. The pricing page lists the passes and their prices.
How do I find the equation of the line of best fit?
Paste your x and y values and keep the straight-line model: the equation appears at the top. By hand, the slope is b = Σ(x − x̄)(y − ȳ)/Σ(x − x̄)² and the intercept a = ȳ − b·x̄; for the points (1, 2), (2, 4), (3, 5), (4, 4), (5, 5) that gives y = 0.6x + 2.2.
What is a good R²?
It depends on the field. Calibration lines often reach 0.999 or more, while models of human behaviour may be useful at 0.3. R² only measures how closely the points follow the fitted curve; it says nothing about whether the model is the right shape, which the residual plot shows.
What is the difference between a confidence interval and a prediction interval?
The confidence interval is for the mean y at a given x — it shrinks as you add data. The prediction interval is for a single new observation, so it also includes the scatter of individual points and never shrinks below about ±t·s. In the hand example at x = 3 they are [2.7, 5.3] and [0.9, 7.1].
How do I fit an exponential curve such as y = a·e^(bx)?
Choose Exponential. The calculator fits a straight line to ln y against x, so b is the slope and a = e^(intercept); every y must be greater than 0. The growth rate per unit of x is e^b − 1, and the doubling time is ln 2 / b when b > 0.
Which polynomial degree should I choose?
The lowest one that leaves residuals without a pattern. Adjusted R² helps: it rises only when an extra term improves the fit by more than chance. A degree close to the number of points fits the data exactly but predicts poorly.