CSV & Excel Data Profiler (EDA Report)
One-click exploratory data analysis: every column’s type, gaps, statistics and quirks.
Try before you buy.
- Free preview: the table’s totals, the first warnings and column profiles (up to 10), and the first lines of each schema (up to 20).
- Locked until you unlock it: download and copy.
- Unlock: Pro pass, ₹179 for 30 days, a one-time payment that never renews.
Ways to unlock shows how to get the full result.
Printing this result is locked in the free preview.
Your data
A CSV, TSV or Excel file, or cells pasted from a spreadsheet. The first row should hold the column names.
Or paste the table (from Excel, Google Sheets or a CSV)
Options
When the data could be read either way, your country’s order is used.
Summary
Warnings
Columns
| Column | Type | Missing | Distinct | Range |
|---|
Report, data dictionary and schema
Locked in the free preview. Opens the ways to unlock this result.
Locked in the free preview. Batch runs unlock with a pass.
Locked in the free preview. Query results unlock with a pass.
About the CSV & Excel Data Profiler (EDA Report)
Open a CSV or Excel file and get an exploratory data analysis of it in one step: for every column its inferred type (numbers, dates, text, yes/no, e-mails, phone numbers, web addresses, UUIDs), missing values, distinct values, minimum, maximum, mean, median and standard deviation, the most common values, text lengths, value patterns and a small histogram.
The report also warns about what usually breaks an import or an analysis: duplicate rows, empty rows and columns, constant columns, columns that mix types, dates that could be read two ways, codes with leading zeros, numbers written with currency signs, likely ID columns and columns that may hold personal data. It writes a printable report of the whole profile (a web page to print or save as PDF), a data dictionary (Markdown or CSV) and the inferred schema as SQL CREATE TABLE, JSON Schema, a Frictionless Table Schema or a BigQuery JSON schema. The file is read and profiled in your browser; nothing is uploaded.
How to use it
- Open a CSV, TSV, Excel (XLSX, XLS), OpenDocument or Numbers file under Your data, or paste cells copied from a spreadsheet. Pick the sheet and check that the first row holds the column names.
- Check the options: whether “NA”, “null”, “-” and similar cells count as missing, and how dates such as 04/10/2026 are meant (day first or month first).
- Read the summary and the warnings, then each column’s card: its type, gaps, statistics, top values, patterns and histogram.
- Choose a schema — SQL for PostgreSQL, MySQL, SQLite or SQL Server, JSON Schema, Frictionless or BigQuery — and its table name.
- Download the printable report (HTML), the data dictionary as Markdown or CSV, or the schema, or copy the schema — with a Pro pass, or after unlocking this result; without one the page shows a free preview.
Examples
postcode: 02134, 40001, 01100, 28001
Text, with the note “values have a leading zero: keep this column as text”
Read as numbers, 02134 would become 2134.
amount: 10, 12, 15, TBD, 18, 20 …
Whole numbers, with “Mixed: 1 value is not whole numbers (“TBD”)”
date: 04/10/2026, 05/11/2026
Dates, with a warning that they fit both day/month and month/day
One date such as 13/10/2026 decides the order for the whole column.
Common uses
- Understanding a file someone sent before working with it: what each column holds and how complete it is.
- Checking data quality before an import, a report or a machine-learning experiment.
- Writing the CREATE TABLE statement or the BigQuery schema for loading a CSV into a database.
- Documenting a dataset with a data dictionary for a team, a client or a data catalogue.
- Finding duplicate rows, empty columns and inconsistent values at a glance.
How column types are decided
Every filled value is classified as a whole number, a decimal number (including values written with digit grouping, currency or percent signs), a date, a date and time, a time of day, an e-mail address, a web address, a UUID, a phone number or text. A column takes the type of at least 80 % of its values — whole numbers count as decimals and dates as dates and times when the two are mixed — and the rest are listed as mixed values with examples. A column whose values are all one of the pairs yes/no, true/false, y/n or t/f is yes/no.
Two rules keep data safe: whole numbers with a leading zero (postal codes, account numbers) stay text, and dates such as 04/10/2026 are read in the order the column itself proves (any day above 12 decides it) or, when nothing decides it, in the order you choose. Columns of more than 100,000 values are classified on an even sample of 100,000; every count and statistic uses all values.
What a column card shows
- Filled, missing and distinct values; a column whose values are all filled and different is a likely ID or key.
- Numbers: minimum, maximum, mean, median, quartiles (Hyndman–Fan type 7, as in Excel’s QUARTILE.INC), standard deviation (n − 1), sum, zeros, negatives and outliers beyond 1.5 × IQR.
- Dates: the earliest and latest, the span in days and how they are written.
- Top values with counts, text lengths and the three most common patterns, where letters become A or a and digits 9 (“AB-1234” is “AA-9999”).
- A histogram: number ranges about as wide as the Freedman–Diaconis rule suggests (2 × IQR × n^−1/3), dates by year, month or day, and the most common values of text columns.
Schemas and the data dictionary
- SQL
CREATE TABLEfor PostgreSQL, MySQL/MariaDB, SQLite or SQL Server: integer, decimal (with the precision and scale the values need, currency and percent signs aside), boolean, date, timestamp, time, UUID and text types — whole numbers longer than 18 digits, which a 64-bit integer cannot hold, as exact decimals;NOT NULLfor columns without missing values; a primary key for a likely ID column named like id, key or code; and a comment where a column needs converting first (dates not written YYYY-MM-DD, numbers with currency signs). - JSON Schema (draft 2020-12) for the rows as an array of objects, with formats for dates, date-times written the RFC 3339 way, e-mails, web addresses and UUIDs, and optionally the observed ranges and value lists as constraints.
- A Frictionless Table Schema with types, date patterns such as
%d/%m/%Y(or%yfor two-digit years), digit grouping (groupChar) and signs (bareNumber) of numbers,requiredanduniqueconstraints and the missing-value markers found. - A BigQuery JSON schema (schema format) with valid column names,
INT64,NUMERIC,FLOAT64,BOOL,DATE,DATETIMEandTIMESTAMPtypes andREQUIREDorNULLABLEmodes.
When a column name repeats, the SQL and JSON schemas number the repeats (Amount, Amount_2) and say so in a comment, because a table cannot have two columns of one name. The data dictionary lists every column with its type, missing values, distinct values, range, most common values and notes. Schemas are inferred from the values in the file: check them before using them, for example widen a text column that will hold longer values later.
Limitations
- Up to 500,000 rows are profiled (text files up to 100 MB, workbooks up to 50 MB, as far as your device’s memory allows); more rows are left out with a note.
- Types are inferred from the values, not from meaning: a column of years is a column of whole numbers, and a phone number written without spaces or a + is a number.
- Excel files are read as their shown values; formulas count as their last saved results.
- The printable report is an HTML file, not a PDF: open it in a browser and print it or save it as PDF from the print dialog.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.
Frequently asked questions
What do I get without a pass?
Without a pass, CSV & Excel Data Profiler (EDA Report) shows the table’s totals, the first warnings and column profiles (up to 10), and the first lines of each schema (up to 20). Until you unlock it, the result can’t be downloaded or copied. A Pro, Premium or Ultimate pass, a one-time payment that never renews, unlocks the full result. The pricing page lists the passes and their prices.
What is data profiling?
A quick summary of what a dataset contains and how clean it is: each column’s type, how many values are missing or distinct, the range and distribution of values, and problems such as duplicates or mixed types. It is usually the first step of exploratory data analysis.
How do I get a CREATE TABLE statement from a CSV?
Open the file, choose SQL CREATE TABLE under Schema and your database, and type the table name. The statement uses the types inferred for each column; check the comments, which say where a column needs converting before it can be loaded.
Why is my number column shown as text?
Either most of its values are not numbers, or its numbers have leading zeros (02134), which would be lost if stored as numbers. The card’s notes say which values did not fit.
Can I print the profile or save it as a PDF?
Yes. With a Pro pass, or after unlocking the result, download the printable report (HTML), open it in your browser and print it, or choose Save as PDF in the print dialog. It holds the summary, the warnings, an overview of the columns and a section for each column with its figures, most common values, patterns and histogram. It is a plain page: no script, nothing loaded from the internet.
Which values count as missing?
Empty cells, and — unless you untick the option — cells that say NA, N/A, n.a., null, None, NaN, nil, -, --, – or —, #N/A, (blank), undefined or missing, in any case.
Is my data uploaded?
No. The file is read and profiled in your browser and stays on your device.