Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

Fine-Tuning Dataset Builder (Spreadsheet to Chat JSONL)

Spreadsheet rows in; checked training and validation files out, built on your device.

Developer No upload Free, no sign-up

Your examples

Or paste cells, CSV or JSON Lines

Columns

Open a sheet first: the page guesses the columns from its headers.

Used where the system column is empty or a conversation has none. OpenAI suggests putting the instructions that worked best before fine-tuning in every example.

Checks

Open a sheet or a dataset, or press Load example, to see the checks.

Tokens and cost

OpenAI’s limit per example for its fine-tunable GPT-4.1 and GPT-4o models.

Counting runs on this device; the first count downloads the tokenizer once.

From your provider’s current price list.
Empty: Auto, the default rule of OpenAI’s cookbook.

Split and download

0 for no validation file; at most 50.
The same seed gives the same split.
dataset-train.jsonl and dataset-validation.jsonl

The example in this format

Nothing to show yet.

Next steps

About the Fine-Tuning Dataset Builder (Spreadsheet to Chat JSONL)

Turn a spreadsheet of examples into the file a fine-tuning job needs. Open a CSV, TSV, Excel or OpenDocument sheet (or paste cells copied from one), say which columns hold the system prompt, the user’s message and the ideal answer — or group rows into multi-turn conversations by an id column — and download chat-format JSONL, the layout OpenAI’s fine-tuning and Hugging Face TRL read, or prompt–completion, Gemini tuning, ShareGPT and Alpaca files for other trainers.

Before you pay for a training run, the checks list what would waste it: empty answers, turns in the wrong order, copies, the same question answered two ways, examples longer than the token limit you set (counted with the LLM Token Counter’s tokenizers), unbalanced labels and too few examples. Split off a validation file — stratified by a label if you like, with a seed so the split repeats — mask emails, phone numbers and other personal data first, and see the token totals with a cost estimate from your price. You can also open an existing JSONL or JSON dataset to check it or convert it to another format. Everything runs in your browser; nothing is uploaded.

How to use it

  1. Choose a CSV, TSV, Excel or ODS file, paste cells copied from a spreadsheet, or open a JSONL or JSON dataset you already have. Load example fills in a small support-desk sheet.
  2. Under Columns, check the column with the user’s message and the one with the answer (the page guesses them from the headers). Add a system prompt for every example, an extra input column, or a label column for the balance check. For chat logs with one message per row, choose Rows grouped into conversations and pick the conversation id, role and message columns.
  3. Read Checks. Examples with an error, such as an empty answer, are left out of the files while Leave out examples with errors is ticked; warnings and tips say what to fix in your sheet.
  4. To keep personal data out of the files, tick Mask personal data: emails, phone numbers, card and bank numbers and the names you list become tags such as [EMAIL] before anything is counted or saved.
  5. Under Tokens and cost, pick the tokenizer of the model you will train and the length limit, press Count tokens, and enter the price per million training tokens to estimate the bill.
  6. Choose the format and the validation share, then press Download training file and Download validation file.

Examples

A row becomes a chat example
Input
question: Do you ship to Canada?
answer: We do. Delivery to Canada takes 5 to 8 working days.
system prompt: You are the support assistant of a shoe shop.
Result
{"messages":[{"role":"system","content":"You are the support assistant of a shoe shop."},{"role":"user","content":"Do you ship to Canada?"},{"role":"assistant","content":"We do. Delivery to Canada takes 5 to 8 working days."}]}
Chat logs, one message per row
Input
ticket 17 | turn 1 | customer | My discount code fails
ticket 17 | turn 2 | agent | Codes are case-sensitive…
ticket 17 | turn 3 | customer | It works now, thanks
Result
One example per ticket, in turn order: customer → user, agent → assistant. A warning notes that the last message is the customer’s, which teaches the model nothing.
The same example as ShareGPT and Alpaca
Input
user: Do you ship to Canada?
assistant: We do.
Result
ShareGPT: {"conversations":[{"from":"human","value":"Do you ship to Canada?"},{"from":"gpt","value":"We do."}]}
Alpaca: {"instruction":"Do you ship to Canada?","input":"","output":"We do."}
A training-cost estimate
Input
400 training examples of 300 tokens, 3 epochs, $25 per million training tokens
Result
120,000 tokens × 3 epochs = 360,000 billed tokens → $9.00

Enter your provider’s current training price: the page has no built-in prices.

Common uses

  • Turning a sheet of questions and approved answers from a support team into a first fine-tuning set.
  • Converting chat logs exported one message per row into multi-turn training examples.
  • Checking a dataset for empty answers, copies and unbalanced labels before paying for a training run.
  • Converting a dataset between chat JSONL, prompt–completion, Gemini, ShareGPT and Alpaca layouts for another trainer.
  • Masking customer emails and phone numbers before the data goes to a training service.

The file formats

  • Chat messages: {"messages": [{"role": "system" | "user" | "assistant", "content": "…"}]}, one example per line. It is OpenAI’s fine-tuning format (supervised fine-tuning) and Hugging Face TRL’s conversational format (dataset formats); Axolotl’s chat_template type (conversation datasets) and Unsloth (datasets guide) read it too. Train only on the last answer adds OpenAI’s "weight": 0 to the earlier answers of a conversation (fine-tuning best practices).
  • Prompt and completion: {"prompt": […], "completion": [{"role": "assistant", "content": "…"}]}. TRL’s SFTTrainer then computes the loss on the completion only, by default (SFT trainer).
  • Gemini tuning: {"systemInstruction": {"parts": [{"text": "…"}]}, "contents": [{"role": "user" | "model", "parts": [{"text": "…"}]}]}, for supervised tuning of Gemini models on Google Cloud (prepare the data).
  • ShareGPT: {"conversations": [{"from": "system" | "human" | "gpt", "value": "…"}]}, which Axolotl maps with field_messages and Unsloth converts with standardize_sharegpt.
  • Alpaca: {"instruction", "input", "output"}, plus "system", like Stanford Alpaca’s alpaca_data.json, which is one JSON array (Stanford Alpaca). It holds one turn; multi-turn conversations are left out unless you write the earlier turns as "history", which only some trainers read.

What the checks look for

  • Errors (left out of the files by default): no user message or no answer, a system message after the conversation has started, a conversation that starts with the assistant, rows with an unknown role or no conversation id.
  • Warnings: two user messages or two answers in a row; a conversation that ends with the user (models learn only from answers); exact copies; examples over the token limit, which trainers cut at the end — OpenAI: “Examples longer than the default are truncated … which removes tokens from the end of the training example”; labels where the largest has at least 3 times as many examples as the smallest; one answer given in 30% or more of the examples (at least 5). OpenAI’s guide warns that if 60% of the answers say “I cannot answer this” when only 5% should, “you will likely get an overabundance of refusals”.
  • Tips: the same question with different answers (a model can only be as consistent as its data), and the size of the set. OpenAI’s minimum is 10 examples, and it reports improvements from 50–100; Google suggests starting with 100 for Gemini (about Gemini tuning), and Unsloth at least 100 rows.

Length limits to choose from: 65,536 tokens per example for OpenAI’s fine-tunable GPT-4.1 and GPT-4o models, 131,072 for Gemini tuning, and 1,024, the default max_length of TRL’s SFTTrainer.

Tokens and the cost estimate

Tokens per example = the tokens of its messages, counted with the tokenizer you choose, + 4 per message + 3 for the roles and separators — the way OpenAI’s cookbook counts chat messages (counting tokens). Chat templates of open models add a similar few tokens. Claude and Gemini tokenizers are not published, so those counts are labelled estimates.

Billed training tokens = the tokens of the training file × the epochs (Google: “the number of tokens in your training dataset by the number of epochs”); an example over the limit counts only up to the limit, because the rest is cut off. Auto epochs follow OpenAI’s data-preparation cookbook: 3, raised for small sets until examples × epochs reaches 100 (at most 25), and lowered to 25,000 ÷ examples for very large sets (data preparation and analysis). The cost is billed tokens × your price per million ÷ 1,000,000.

Limitations

  • Text conversations only: tool calls, tool results, images and audio are not supported, and imported records that contain them are left out (the checks list their line numbers).
  • Token counts use the tokenizer you choose plus a fixed allowance for roles; the exact count depends on the model’s chat template, so leave a margin under a hard limit.
  • The checks find mechanical problems. They cannot tell whether an answer is right or written the way you want your model to answer: read a sample of the examples yourself.
  • Masking finds emails, phone numbers, card and bank numbers, identity numbers, IP addresses and the words you list; it cannot find every name or street address in free text.
  • Up to 200,000 rows are read from a sheet; formulas are read as the values the file last saved.

Privacy

Your sheet, the examples and the files you download are built in your browser and never uploaded. Counting tokens downloads only the tokenizer you choose, once.

Frequently asked questions

Which format should I choose?

The one your trainer reads. Choose Chat messages for OpenAI’s fine-tuning, Hugging Face TRL, Axolotl or Unsloth; Prompt and completion for TRL when the loss should cover the final answer only; Gemini tuning for Google Cloud; and ShareGPT or Alpaca for trainers and notebooks written for those layouts.

How many examples do I need?

OpenAI’s minimum is 10, and it reports improvements from 50 to 100 well-made examples; Google suggests starting with 100 for Gemini, and Unsloth recommends at least 100 rows, preferably over 1,000. Quality matters more than size: OpenAI notes that “a smaller amount of high-quality data is generally more effective than a larger amount of low-quality data”.

Why keep a validation file?

The trainer reports the loss on examples it does not learn from, which shows when the model starts memorising its training set instead of learning the task. The split is random but seeded, so the same seed gives the same files, and Stratify by label keeps each label’s share in both files.

Can I still fine-tune with OpenAI?

OpenAI says it is winding down its fine-tuning platform: it no longer takes new users, and existing users can create training jobs for the coming months (fine-tuning best practices). The chat-messages file is also the standard input of Hugging Face TRL, Axolotl and Unsloth, so the same file can train an open model.

Why are two rows with the same question flagged?

If two examples ask the same thing and answer differently, the model learns that both answers are right. OpenAI’s guide points out that a model is limited by how much the people who wrote the data agree. Keep the answer you want, or make the questions different.

Is my data uploaded?

No. The sheet is read, checked, masked and converted on your device, and the files are saved from your browser. The only download is the tokenizer, the first time you count tokens with it.

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.