Docs

Tips

The same 97 short notes CLAIR shows on its own launch screen, collected here by feature. Each one teaches a method, not a button.

Importing files

What CLAIR does the moment a file lands, and what it will not silently guess.

  • Start by dragging a spreadsheet onto the window. CLAIR reads it here, on this computer.
  • Your spreadsheets stay on this machine. CLAIR reads them here and never sends their contents anywhere.
  • A cover sheet or a notes tab in a workbook is not data. CLAIR shows the measured rows and columns for every sheet so you can tell which ones are worth importing.
  • Each sheet you import from a workbook becomes its own dataset in CLAIR, all kept together under the same file so you can find them later.
  • A subtotal or grand total row sitting among your data gets counted as its own record. CLAIR looks for rows that read like a total and tells you where they are, without removing them for you.
  • A row where every cell is blank carries no information, wherever it sits in your file. CLAIR removes those automatically and tells you how many it took out.
  • A row carrying more values than its header can quietly lose one. CLAIR warns you and keeps those columns as text, so a damaged number never passes as a clean one.
  • A row with fewer values than your header expects simply gets its last columns left blank, and CLAIR tells you which rows came up short so you can check the source file.
  • A file that separates its values with a semicolon or a vertical line instead of a comma used to open as one giant column. CLAIR now works out which symbol your file actually uses before reading it.
  • A company name or notice printed above the real header used to be read as if it were the header itself. CLAIR looks for that kind of banner and skips past it to the real one.
  • When a header spans two rows, like a category name sitting above several column names, CLAIR merges them into one clear column name instead of reading the category as a stray row of data.

Data Health

The automatic scan every file gets on arrival, before you ask CLAIR anything.

  • You do not have to know which analysis to run. Import a file and CLAIR profiles every column before you ask it anything.
  • CLAIR scores every file for health the moment it lands, with no model and no questions asked of you.
  • Some rows are wrong only in combination. An age of 8 alongside a doctorate is two ordinary values that cannot both be true, and CLAIR looks for exactly that.
  • A clean file gets no findings. The background scan is built so an ordinary dataset never manufactures problems just to have something to report.
  • When the odd rows in a file all come from one source, one site or one form, the oddness is in how it was collected rather than in the people.

Columns

  • A wide survey file can carry hundreds of columns. The search box at the top of Columns finds one by name.
  • Marking which columns are outcomes and which are identifiers takes a minute and changes what the rest of CLAIR will let you do to them.
  • Columns has its own Data Explorer tab: browse and rearrange the raw data without leaving the page where you are editing column classifications.

Clean

  • Clean lists what is wrong with a file before it changes anything: numbers stored as text, dates in several formats, the same answer spelled four ways.
  • Yes, YES, Y and 1 are one answer typed four ways, and they get counted as four. Clean finds those and merges them in one step.
  • Clean applies only the fixes it is certain about on its own. Anything that needs a judgement is shown to you first.
  • A scale that runs from Never to Always has an order, and merging two of its steps changes what it measures. Clean refuses that merge unless you insist.
  • Filling gaps with the column average makes every one of those rows look typical. If the gaps are not spread evenly across your groups, it quietly pulls the groups together.
  • A date that arrived as plain text will not sort, group or forecast. Convert a Column in Clean turns it back, and shows you both readings when a format like 03/04/2026 is ambiguous.
  • Clean can combine two datasets from the same project, either by matching a shared column between them or by stacking their rows together.
  • If one dataset stores its variables in rows and the other stores them in columns, merging them directly would garble the result. CLAIR spots the mismatch and turns one to match the other first.

Chat

  • Ask a question in plain English on the Chat page. When the answer came from a query CLAIR wrote and ran, you can open the query and the rows behind it.
  • Chat marks an answer verified only when it came from a query actually run against your table. Anything else is left unmarked.
  • If one row of your file is one measurement rather than one person, an average across rows is not an average across people. CLAIR says so in the answer when it spots that.
  • When a survey question let people pick more than one answer, counting the cells counts combinations rather than choices. CLAIR flags the answers where that matters.
  • A year spread across four columns is one thing stored as four. Totals across them are not totals, and CLAIR will tell you when a question runs into that.
  • An average over rows where some people appear more often than others leans towards those people. CLAIR looks for repeated identifiers before it answers.

Factor Finder & Segments

  • Pick an outcome in Factor Finder and CLAIR ranks every other column by how closely it moves with it.
  • Two columns that should move together and do not are a sign of a data problem, not a finding. Year in school and age are the classic pair.
  • A relationship far stronger than you expected is usually two columns measuring the same thing twice. Check the top of the ranking first.
  • Factor Finder describes your data, never your model, and it never says one thing caused another. It keeps the two rankings on separate lists for that reason.
  • A pattern that holds across everyone can weaken, vanish or reverse inside each group. Factor Finder can rank the same outcome group by group so you can see whether it does.
  • If a relationship only holds for one group, the overall number is an average of two different stories and describes neither.
  • Segments finds groups of rows that resemble each other, without being told what to look for, then tells you what makes each group different.
  • How much a column moves an outcome while the others are held still is a different question from which column moves with it most. Regression answers the first one.
  • Columns that only mark a missing response are hidden from Factor Finder's ranking by default, so one cannot pose as a real driver. CLAIR shows how many it hid and lets you bring them back.
  • A strong result in Factor Finder is a relationship, not proof it is real. A link on the results page takes you straight to the matching test in Statistics.

Statistics

  • Statistics builds the summary table most papers open with: counts, averages, spread and how much of each column is missing.
  • When a number was worked out from fewer rows than your file holds, CLAIR prints that smaller count next to it. A count that keeps shrinking is a column worth opening.
  • Measuring the same people twice is not the same as measuring two groups once. CLAIR checks which one you have before it compares them.
  • If you would have needed 300 people to see the difference you care about and you have 60, a result of "no difference" has not told you there is none.
  • Before you collect anything, CLAIR can tell you how many observations you would need to detect a difference of the size you expect.
  • Comparing a group of 400 with a group of 9 is not wrong, but the group of 9 decides how sure you can be. CLAIR reports the count behind each side.
  • Before comparing two groups with a t-test, CLAIR checks whether their values look like a bell curve. When they do not, it suggests a different test that does not need one.

Missing data

  • Seeing what is missing in every column at once is usually a faster read of a new file than looking at the values.
  • Missing values do not only cost you rows. If a column is empty for a fifth of your file, every number built on it was built on the other four fifths.
  • Missing at random costs you certainty. Missing for a reason costs you the answer. CLAIR checks whether the gaps in one column line up with the gaps in another.
  • If the people who skipped a question are not like the people who answered it, the answers you have are not the whole group.

Study design

  • Students sit in classrooms and patients sit in clinics. When rows are grouped like that, the ordinary comparison treats them as independent and reports more certainty than you have.
  • CLAIR reads the shape of your table to work out how the data was collected, and it will say it cannot tell rather than guess.
  • If one site, class or clinic contributed most of your rows, your result mostly describes that one. CLAIR checks whether your rows cluster that way.

Duplicate detection

  • A person is often only pinned down by several columns together, such as first name and last name and date of birth. CLAIR can check that the combination is unique even when no single column is.
  • The same person entered twice inflates your count, and every number built on that count is wrong by the same amount.
  • Duplicate rows are rarely spread evenly. If one site or one form double-entered, that site now weighs more than the others in every result.

Predict & Deep Analysis

  • Predict trains a model on your data and then shows you which columns it leaned on, so you can read what it found rather than only what it scored.
  • A model handed the answer among its inputs scores beautifully and means nothing. CLAIR refuses to train when an outcome column is in the input list.
  • A single predicted number hides how sure the model is. CLAIR gives a range instead, and refuses to give one at all when it has nothing to calibrate against.
  • A model learns the data it was given, including the rows that were recorded carelessly. CLAIR can flag rows whose recorded outcome disagrees with everything else about them.
  • If the rows flagged as possibly mis-recorded all come from one group, that group was recorded differently, which is a collection problem rather than a modelling one.
  • A Deep Analysis search that reads paused has genuinely stopped, so what you see is what is happening. Click Resume and it continues from where it left off instead of starting the search over.
  • When training says too few rows are left, one mostly-empty column is usually the cause, not the file. CLAIR names the sparsest ones it set aside so you can judge whether they were worth the rows.

Forecasting

  • With a date column in your file, CLAIR can project it forward and show you how each candidate method scored on the history it already has.
  • CLAIR scores every forecast against the plain guess that next season looks like last season. If nothing beats that guess, there is no pattern there to project.

Files & refresh

  • CLAIR remembers where a file came from and can tell you when the original has changed on disk since you imported it.
  • A result worked out from last month's copy of a spreadsheet is a result about last month. CLAIR checks whether the source has moved on since.
  • Upload a follow-up file and CLAIR will tell you, column by column, what moved since the first one and what did not.
  • A second wave of data is not automatically the same population as the first. CLAIR compares them before you trust a model built on the first.
  • If your new wave differs from the old one in who answered rather than in what they answered, comparing the two measures your recruiting, not your effect.

Signals

  • Signal scans your columns in the background for spikes, trends, outliers and columns that barely vary, then ranks what it found.
  • A column where nearly every row holds the same value cannot explain anything. Signal points those out so you do not spend an afternoon on one.
  • A run of missing values that starts on one date usually means something changed in how the data was collected, not in what was being measured.

De-identification

  • CLAIR flags which columns look identifying, both by their names and by what the values look like, and it never changes one on its own.
  • A de-identification report lists every column, including the ones nothing was done to. Leaving them out would read as nothing to see here, which is not a claim CLAIR can make.
  • An average of three people is an identification, not a statistic. CLAIR can hold back a group result when the group is too small to report.

Change history

  • Every change CLAIR makes to a file is recorded along with what it looked like before, so you can read back what happened weeks later.
  • A reviewer who asks how you got a number wants the steps, not the number. CLAIR keeps the steps and writes them into the report with it.

Templates

  • Save a set of questions as a template and run the whole set against a new file in one go.
  • Asking two files the same questions is the only fair way to compare their answers. A template is what keeps the question set from drifting between them.

Experiments

  • Experiments suggests which combination to try next based on what your runs so far have shown.
  • Running your treatments in the order you thought of them mixes the treatment up with the day. A planned run sheet shuffles that order for you.
  • Repeating a few runs at identical settings is how you learn how much your measurement wobbles. Without that you cannot tell a real effect from noise.
  • When something you cannot control changes between batches, put the whole comparison inside each batch. Then the batch cannot be the explanation.

Charts & dashboards

  • Choose the columns you want to see and Charts picks a sensible chart for them, which you can then change.
  • A chart drawn from a column with gaps is drawn from fewer rows than your file holds, so CLAIR prints that count on the chart.
  • A bar chart of averages says nothing about spread. Two groups with the same average can look nothing alike, so plot the spread before you believe the bars.
  • Arrange a dashboard in the order you want it read before exporting. The PDF follows that order, one panel per page, so the layout carries the argument.

Exports & documentation

  • CLAIR can write a codebook for a dataset: every column, what it holds, and either the range it covers or the values it takes.
  • A column description CLAIR wrote and nobody has checked is marked unreviewed in the codebook, so a reader can tell which is which.
  • A warning Chat adds to an answer, like a small group or dropped rows, is not left behind when you export it. The PDF, PowerPoint, HTML and JSON versions all carry the same warning.
  • You do not have to go through Chat to get the Full Dataset Report. The same export button is on Dashboards and Predict too.

Using CLAIR

  • If a word on screen is unfamiliar, the small info icon beside it explains it in one sentence.