---
name: numberdb-table
description: Make or update a table of numbers for NumberDB (numberdb.org) with the `numberdb` Python package — choosing the quantity, the convention, the precision and the rigour level, writing a generator, and publishing it. Use when asked to add, extend, correct or verify a NumberDB table, or when working from a "table wanted" issue.
---

# Making a NumberDB table

NumberDB answers one question: *here is a value — has anyone seen it before?*

The values are not only real numbers. A table may hold integers, rationals,
reals, complex numbers, p-adic numbers, or polynomials over Z or Q — and the
database is meant to be able to hold kinds nobody has added yet, which are
shown and cited even where they cannot be searched by their digits. Of the 107
tables today, 65 are real, 16 integer, 12 polynomial, 6 p-adic, 4 complex, 3
rational, and one holds hyperreals.

A table earns its place by making that answer possible and trustworthy. So the
work is not "compute some values"; it is "compute values somebody else can
check, indexed so they can be found, and labelled with how well they are
known".

Install: `pip install numberdb`, or `sage -pip install numberdb` inside
SageMath. In Sage use `import numberdb.sage as numberdb`, which returns Sage
objects; that works on SageMath and on a modular
[passagemath](https://github.com/passagemath/passagemath) install alike
(`pip install passagemath-symbolics numberdb`, which has everything but
p-adics). Without any Sage, plain `import numberdb` still works and returns
Python values.

## 1. Look at the database before writing anything

Search it, for two different reasons, and neither is optional.

**Is this table already here, under another name?** The question is about the
*family*, never about the values. A number appearing in several tables is the
point of the database, not a fault in it: pi is in T7 as itself, in T59 as a
value of the elliptic integral, and in T92 as a Sobolev constant, and somebody
who arrives holding 3.14159 is better served by three answers than by one.
Never drop a proposal because its numbers are already stored somewhere -- the
different context *is* the product.

What is worth avoiding is the same family twice by accident. Two tables of the
same objects are sometimes right -- the Hermite polynomials are held twice, in
the physicists' and the probabilists' conventions -- but that is a decision,
not an accident. Values that are a *derivable* multiple of stored ones are a
different question from values that are already **findable**: a reader with a
number tries a half, a two, a pi, and stops. If getting from their number to
the stored one means guessing a rational nobody would guess, the family is not
covered, whatever the identity says. The same question refuses the opposite
family: values that are a handful of small integers -- class numbers are
mostly 1, 2, 3 -- match everything and tell nobody anything, and putting them
in a table makes search by number worse for everyone. Those belong in the
comment on the entry they explain, which is where a reader meets them. `/drafts` lists what is being made right now and is invisible from
outside, so a table can be half-built and unfindable while you start it again.

**What does the database already hold that this table should point at?** A
reference to a table here is worth more than a link to Wikipedia for the same
thing: a reader following it lands on the numbers. Search titles, tags and
definitions before reaching for an external link. The Fibonacci polynomials
should point at the Chebyshev polynomials of the second kind, which are T99 --
not at an encyclopedia article about them.

    numberdb.search_text('Chebyshev')          # titles and tags
    numberdb.table('T99')                      # what it holds

Three things about those two calls, each of which has cost somebody an hour.
`search_text` returns an object with a `.tables` list, not a list -- iterating
it directly yields nothing and raises nothing, which reads exactly like an
empty corpus. The search **stems**, so "regulator" also matches "regular", and
it looks past titles into definitions and comments; read what came back rather
than counting it. And there is no call that lists the corpus, so to see
everything you walk the T-numbers.

`numberdb.table(...)` answers with the whole document, capitalised keys and
all -- `Title`, `Definition`, `Numbers` -- and `Numbers` is every entry the
table holds. Printing it is megabytes. Take the keys you want.

**Never guess a name the database resolves -- an address or a tag.**
`HREF{...}` takes the slug, and the slug is not the title with underscores:
mathematics is dropped from it, it is truncated, and a clash appends a number.
A table called "Power sum symmetric polynomials $p_k$" is at
`Power_sum_symmetric_polynomials`, and a link written to
`Power_sum_polynomials` from the shorter name in one's head points at nothing.
Tags are the same: the corpus says `zero`, not `zeros`, and a tag invented
from the plural in your head simply fails to group anything.

Read both off tables that already have them. `numberdb.table('T119')` returns
the document and **carries no address**; the slug is in a search result,
`numberdb.search_text('power sum').tables[0].url`, or in the address bar. Let
`manage.py audit_table` confirm afterwards rather than instead.

## 2. Decide what the table says before computing anything

Write these down first. Every one of them has gone wrong in this corpus, and
none was caught by a test.

- **Which quantity, exactly.** Not "the AGM" but "the AGM, taking at each step
  the square root nearer to the new arithmetic mean". Over the p-adics the
  other choice gives a different limit, and the table that failed to say so
  could not be reproduced from its own definition.
- **Which normalisation, branch, and indexing.** `elliptic_k` takes the
  parameter *m*, not the modulus *k*; the two differ from the second entry
  onwards. "The nth zero" needs a statement of where counting starts.
- **What the parameters are, and what constraint they actually satisfy.** One
  table said `k = 1 mod p` while 702 of its 856 entries were not. State the
  constraint the entries meet, and if the values extend beyond the obvious
  domain, say by what extension.
- **What each value will be**: an exact integer, rational or polynomial, or an
  approximation. Exact values are stronger and shorter — return them exactly
  rather than as a hundred digits of an integer.

**Hold the numbers that turn up.** The question a table has to answer about
its range is not how many entries that makes, nor how simply they can be
described, but whether anybody will arrive holding one of them.

$\Gamma(1/3)$ turns up: the multiplication formulas put it in people's hands.
$\operatorname{erf}(1.96)$ turns up, because that is a z-score.
$\operatorname{erf}(17/18)$ does not, and a bound on the denominator that
admits it is counting rather than choosing. $t^2+2$ turns up constantly and
almost never as the Weil polynomial of an abelian variety, so a table holding
every such polynomial in a small box answers searches it cannot inform.
$(x-3)^2$ turns up as a square and never as a Hecke polynomial worth
recording -- one draft held 89 entries of which 78 were $(x-a)^2$, because a
rational eigenvalue makes the characteristic polynomial a power of a linear
one whatever the field.

When a grid is what you need, let the arguments choose it rather than the
count. The six error-function tables and the four Airy ones were built on
every $a/b$ in lowest terms with $b\leq18$, a bound picked because it made
about a thousand entries; they now hold every argument of two decimal places
in their range, which is the same thousand entries and every one of them a
number a computation hands back.

No general test settles this. Which parameters are the interesting ones is a
fact about the family, and the table is the only thing that knows it: a
denominator bound, a size target, a count that matches the older tables --
none of them is a reason for any particular number to be present. Say which
range you chose and why in the completeness note. That sentence is the
argument, and it is what a reader checking whether their own number belongs
here actually reads.

The cost of getting it wrong is not storage. A table full of numbers nobody
looks for still answers when somebody looks for something else, and every hit
in the corpus is worth a little less for it.

**A title is what a search reaches.** The text index has four weights: the
title and `Keywords` first, then `Tags`, then the `Definition`, then the
table's `Comments`. **An entry's own comment is in none of them** -- a
constant that appears only as a row, with its name in that row's comment,
cannot be found by its name at all. So the title decides how finely to divide
a subject. "Feigenbaum constants" answers somebody typing "Feigenbaum";
"Constants of the regular continued fraction" answers nobody looking for
Khinchin, Levy or Lochs, and those were three tables pretending to be one.
See the three rules below. The test: *what would somebody type who is holding one of
these numbers and wants to know what it is?* If the title does not contain
that word, put it in `Keywords`, which is the same weight -- that is also how
an accent is handled, since the index has no `unaccent`: the title is
"Lévy's constant" and `Keywords` carries "Levy constant".

**Where several named objects share a subject but not a name, make several
tables.** Objects, not only constants: this corpus splits one table per
function and has done so consistently -- T20 and T21 are the zeros of the
Bessel functions of the first and second kind, T22 and T23 their extrema,
T55 to T58 the same for Airy Ai and Bi, and T25, T26, T59 are the three
complete elliptic integrals. Not one of them is "Values of the Bessel
functions". A table holding $\operatorname{Ei}$, $\operatorname{li}$,
$\operatorname{Si}$, $\operatorname{Ci}$, $\operatorname{Shi}$ and
$\operatorname{Chi}$ together is six tables wearing one title, and two of the
six are not findable by name in it.

**Where the members share an index, one table however few.** Khinchin's means
$K_p$ are one object at a sequence of orders, and belong together. A shared
*argument* is not that: seven functions evaluated at the same rational $x$ are
still seven functions, and $x$ is the index of each of them separately. The
question to ask is whether the parameter enumerates values of one thing or
names of several.

**Where two entries are one object in two conventions, one table.**
$E_1(x)=-\operatorname{Ei}(-x)$, $\beta$ and $e^\beta$ for Lévy's constant,
$r$ and $c$ for the logistic map: those are normalisations, and a reader
holding either should land in the same place.

**The test for a convention is that it is invertible.** Landing in the same
place is only possible if each form determines the other. A *specialisation*
does not: $C_n(q,1)$ and $C_n(q,q^{-1})$ are the Carlitz and MacMahon
$q$-Catalan numbers, each a different sequence of polynomials with its own
literature, and neither one gives $C_n(q,t)$ back. Two tables holding the same
numbers in two units are one table; a table and the thing you get by setting
one of its variables to 1 are two.

**Invertible is necessary and not sufficient.** Three things still argue for
two tables when each form does determine the other:

* **The parameters differ.** The Krawtchouk polynomials of the Hamming scheme
  are indexed by an integer alphabet size $q$; in the Askey normalisation by a
  rational $p$. Each is a rescaling of the other where they meet, and they
  meet only on a subset, so neither table contains the other.
* **One form is far more compact, so it can be carried further.** Measured
  over the four Ehrhart pairs, as mean characters per stored value: the
  Birkhoff polytopes write 381 for the Ehrhart polynomial against 134 for the
  $h^*$-polynomial, the hypersimplices 352 against 151, the root polytopes 223
  against 140 -- and the permutohedra the other way round, 180 against 300.
  Which form is small depends on the family, so one table holding both reaches
  a different range in each, and a range that is ragged for this reason is a
  sign of two tables rather than of one unfinished.
* **The conversion costs more than the storage.** A closed form both ways is
  the case for one table; a transform you would have to compute to answer a
  reader holding the other form is not.

**The question to ask about a parameter: does it name *what the number is of*,
or *which quantity is taken of it*?** The first indexes a family and belongs in
one table -- one Hausdorff dimension for each of forty named sets, one Ehrhart
polynomial for each of seventy-eight root systems, one entropy for each of
twenty-two distributions. The second is a second table wearing the same title.

**The tell is structural, and the audit now reports it on a draft.** A
parameter whose values are a handful of *names* -- `form: ehrhart | h-star`,
`quantity: psi | H`, `form: generating | signed` -- and which every other
parameter repeats under, is one table written twice, sharing a value column
with itself. The column gives it away too: a table of one quantity heads its
column with that quantity's symbol, `$\gamma_K$` or `$\phi(G,x)$` or
`$I(T,x)$`, and a table of two falls back on the word `value`, because no
symbol is true of every row. If you find yourself writing `value` there, ask
what the table holds.

Three answers to that question are good ones, and the audit says so rather than
insisting:

* **Parts of one object.** The $abc$-triples store $a$, $b$ and $c$ because a
  triple is the thing; an elliptic curve is stored as $N$, $c_4$, $c_6$
  because $c_4$ and $c_6$ identify the curve and $N$ says at a glance which
  curve it is. A reader holding one part wants the others beside it.
* **One number in two conventions**, as just above.
* **A parameter that is an argument**, not a name: $\nu = 0, 1, 2$ is one
  function at three orders.

Everything else -- the derivative of a function and its logarithmic
derivative, a polynomial and a different polynomial of the same object, a
constant and the family it is a member of -- is two tables.

**Splitting does not lose the connection; it puts it where it belongs.** The
argument for one table is usually that the objects belong together, and this
corpus has four ways to say that, each saying something different:

* **`Similar tables`** states the *relation*, in words: "the same expansion's
  limits, taken over the denominators of the convergents rather than over the
  partial quotients". That sentence is the valuable part, and one table cannot
  hold it, because inside one table there is nothing to relate.
* **`HREF{slug}`** in the prose links at first mention, and
  `HREF{slug#entry}` links the one value a sentence means.
* **`equals`** says the two entries *are the same number*, which is stronger
  than related and which search reads: it answers with the original and names
  the other beside it. Related is not equal, and saying equal where only
  related is true is worse than saying nothing.
* **`Tags`** say the two are about the same subject, which is how a reader
  browses sideways rather than by name.

$\operatorname{Ei}$, $\operatorname{li}$, $\operatorname{Si}$,
$\operatorname{Ci}$, $\operatorname{Shi}$ and $\operatorname{Chi}$ are
similar and are not the same. Six tables, each defining its own function,
tagged alike, linked to one another, and each findable by its own name.

**Do not discard what the source certified.** T174 read a certified spectral
data set whose eigenvalues alternate in sign -- the file even carries a
`sign_certified` column -- and stored $|\lambda_n|$. The sign is not
recoverable from the table: a reader cannot see an alternation in a column of
positive numbers, and a search for the number they hold does not match its
negative. Where the literature quotes a magnitude, as it does for the
Gauss-Kuzmin-Wirsing constant $|\lambda_2|$, that convention is about one
value; say so in that entry's comment and store the rest as they are.


**Decide exactness from the definition, never from the digits.** A run of
zeros is evidence and not proof: `3.000000000000000000` may be 3, or
3 + 10^-40 rounded, and no number of zeros separates them. The logistic map's
$r=3$ is exact because the fixed point loses stability where $|f'|=1$, which
is an argument. Where there is such an argument, write the integer or the
rational -- a decimal means plus or minus one unit in its last place however
long it is, so `3.000...` says a number known to *be* 3 is known to a hundred
places, and it understates exactly the rows a reader came for. Where there is
no such argument the zeros mean the opposite thing, a rounding presented as
sixty significant places, and the fix is fewer digits or a ball. Surds cancel
more often than one expects: $r=1+2\sqrt2$ gives $c=-r(r-2)/4=-7/4$ exactly.

## 3. How much to include

**A table is a reference, not a dump.** NumberDB exists so that somebody who
has met a value can find out what it is. That is the test for every entry: is
this a value somebody could plausibly encounter and want to identify? Nobody
meets the 500th Chebyshev polynomial and wonders what it is. A hundred of a
family is a reference; a thousand is a listing of something nobody was looking
for, and it makes the search results worse for everybody by burying the values
that are common.

So: **include what is common, and stop.** Then check what it costs.

See `docs/design/corpus-shape.md` for what the corpus does. Very roughly, and
only as a starting point to think from: 500–1000 entries for cheap
approximations, one or two integer parameters, 100 significant digits.

### Which parameter values, when the parameter is continuous

A family indexed by an integer chooses itself: the $n$th zero, the packing of
$n$ circles. A family indexed by a real number does not, and the choice is
the whole of the table's usefulness — because **search is by value, not by
parameter**. Nobody looks up $\xi(1.37)$; somebody has $0.3070\ldots$ and
searches for it. An entry earns its place only if a reader's parameter lands
on ours exactly.

**Use rationals of small height. Not rationals of short decimal length.**

A rational's size is its denominator in lowest terms. $1/2$, $2/3$, $3/4$
have heights 2, 3, 4; $7/100$ and $103/100$ have height 100. A decimal grid
is tidy in base ten and mathematically large-height, which is exactly why it
reads as arbitrary: $51/50$ is not a number anybody writes down for its own
sake. There is no natural reason for 10 here. It is a fact about human
notation, not about the function.

So, in order:

1. **Distinguished points first.** Closed forms, named values, singularities,
   the arguments other problems land on: $m=0,\tfrac12,1$; $q\to0$; the
   Blasius case $\beta=0$; the Falkner–Skan separation value. Prefer *exact*
   entries — an exact value outranks any number of decimals, and raises the
   table's rigour instead of spending it.

2. **Then rationals of bounded height.** All $p/q$ in the range with
   $q\leq N$ for a small $N$ — the Farey enumeration. It is canonical rather
   than chosen, it already contains $\tfrac12,\tfrac13,\tfrac23,\tfrac14$,
   and it is dense where the rationals are interesting rather than where base
   ten is. Say the bound in `complete-note`: "every $p/q$ with $q\leq8$ and
   $0\leq p/q\leq10$".

3. **Angles are small rational multiples of $\pi$.** $\pi/6$, $\pi/4$,
   $\pi/3$, $\pi/2$ — where the trigonometry is exact. Degree grids are the
   same mistake in another base: $5°$ is $\pi/36$, height 36, against
   height 3 to 6 for the values that mean something.

**No decimal grid, even when the function has no distinguished points.** That
case is real — $\operatorname{erf}$ has none beyond $\operatorname{erf}(0)=0$,
no closed form at any rational argument — and it is still not a reason for
tenths. Bounded height gives a table of comparable size whose every point is
a rational somebody might name.

**If a fine grid is genuinely right, it is because a named source publishes
that grid.** Cite it, and check it rather than assuming: the handbooks are
decimal because they were built for hand interpolation, which is a claim
about a book and needs to be true of the book you cite. "The corpus already
does this" is not a reason — it is how one table's choice became forty.

**Density is not completeness.** `complete-note` says which points were
chosen and why, not the arithmetic of the step size. "Every decimal grid
point $q=j/1000$" describes a step; it does not say why those are the
values a reader will have.

**Two things move that number down, and they are the usual case.**

*Expensive digits.* Few numbers known to great precision is as legitimate as
many known to a hundred digits; both at once is not.

*Values that grow.* A polynomial of degree n has about n/2 terms with
coefficients of O(n) digits, so it costs O(n²) characters and a table running
to n costs O(n³). Measured on Chebyshev polynomials of the first kind:

    n = 50      570 characters      table 0..50     11 KB
    n = 100    1892                 table 0..100    69 KB
    n = 200    6828                 table 0..200   472 KB   over the soft limit
    n = 500   39674                 table 0..500  6639 KB   over the hard limit

**The binding constraint is usually readability, not size.** An entry is
something a person looks at. The Fibonacci polynomials were first published to
n = 150 because that fitted comfortably inside every limit, and the range came
back down to 100 for a better reason: F_150 is 2248 characters and F_100 is
1107, and somewhere between those an entry stops being something anybody reads.
Ask what the largest entry looks like on a page before asking whether the table
fits.

**A comment on every entry is part of the entries block.** Notes are stored
with the values they belong to, so a line of prose on each of a thousand
entries is counted with them: on one table the block was 166 KB with a comment
on every row and 152 KB without. Worth having when it says something the row
cannot -- the field a discriminant names, a caveat about one value -- and
worth leaving out when it repeats what a formula in the document already says
for every row at once.

**Then aim at half the soft block limit -- about 160 KB -- not at the limit.** A
table that only just fits cannot be extended by the next person without
breaching it, and the limit is there to be a margin rather than a target.

For a family of polynomials indexed by degree, **n up to about 100** is the
house range, and it lands where it should:

    chebyshev_T   0..100     69 KB
    hermite       0..100    144 KB
    legendre_P    0..100    164 KB     rational coefficients cost more

Stop earlier when the coefficients are rational or the family grows faster, and
**measure the largest entry before choosing the range** rather than assuming.
Beyond that, the question is not what fits but what anybody is looking up.

The server enforces three limits (`numberdb_app/limits.py`):

| | recommended | soft | hard |
|---|---|---|---|
| entries | 1000 | 1200 | 50,000 |
| digits | 100 | 500 | 10,000 |
| entries block | — | 320 KB | 4 MB |

A soft limit may be passed by an author who explains why, recorded in the table
as `Size exception`. A hard limit is not a judgement and cannot be passed. The
digit limits do not apply to exact tables. The block limit is the one that
binds for polynomials, and it is a limit on the *whole table*: it admits many
small entries or a few large ones, and refuses both at once.

## 4. Where each thing goes

A table has sections and they are not interchangeable. The definitions of the
two newest tables in the database each grew to hold a definition, two
conventions, a caveat about indexing and a pointer to a companion table, and
had to be taken apart again.

| section | holds | not |
|---|---|---|
| **Title** | what the thing is called, LaTeX allowed | a description |
| **Definition** | one or two sentences saying what the object *is* | caveats, history, relations |
| **Parameters** | each parameter's type and the constraint the entries actually satisfy | aspiration |
| **Comments** | conventions, caveats, what the values mean at a special point, alternative indexing, notation used elsewhere in the table | formulas |
| **Formulas** | identities, closed forms, generating functions, recurrences | prose about them |
| **Similar tables** | tables *in this database* it relates to, with the relation named | external links |
| **Links** | sources outside: Wikipedia, LMFDB, OEIS, MathWorld | anything the database holds itself |
| **References** | papers and books, cited from the prose with `CITE{}` | uncited decoration |
| **Programs** | the standard incantation for a reader who wants one more value | the generator |
| **Data properties** | `type`, `rigour`, how the digits were obtained, and `repeats` if this table repeats another's | anything else |
| **`rigour details`** | how the digits were obtained, at whatever length that needs: blank lines make paragraphs, `- ` makes a list, backticks mark code, and a long note folds after its first paragraph | a second copy of the definition |
| **References** | `bib`, and the identifiers beside it: `arxiv`, `doi`, `zbl`, `mr` -- each renders as a link | an identifier written into the bib text, where it is not one |

Two rules that follow from the table above and are worth stating alone:

**A parameter says what the family is indexed by, not how much of it you
computed.** These are two different facts and they have two different homes. A
table of `zeta_K(s)` at negative odd `s` has a parameter `s` that is *a
negative odd integer* -- writing `$s\in\{-1,-3,-5\}$` describes this
afternoon's run, and the next person to extend the table has to edit the
definition of the family to add a row. The range that is actually here belongs
under `Data properties`: `complete: no`, with `complete-note` saying which part
is finished. It renders inside "complete: no (...)", so write a clause that
finishes that sentence -- "it holds every real fundamental discriminant with
$D\leq 1000$, at $s=-1,-3,-5$" -- and it is the
sentence a reader wants: not that the table is incomplete, which is true of
almost all of them, but what it *does* cover.

**A parameter's key is an address, not a word.** It identifies every entry
under it, so it is frozen when the table is published and a citation resolves
on it. Two things follow. Choose a specific one while you still can --
twenty-two tables here call a parameter `expression` for three unrelated
things, and `quantity`, `form` or `normalisation` each say which. And never
name a parameter by its key in prose: "the quantity named by the parameter
`expression`" tells a reader nothing, because the key is how the document
addresses a column, not a word anybody outside it knows. Name the quantities,
or use the parameter's `display` symbol.

**A value's display is a label, not a formula.** It is printed on every row
that value indexes, so a conversion rule put there is repeated once per row:
one table carried `$c=-r(r-2)/4$` on all twenty-four of its $c$ rows where
`$c$` was wanted, and the rule belongs in `Formulas`. The same for a closed
form the entry's comment already gives. The value column has a label too,
`Display properties: number-header`, and it is a claim about every value under
it: T174 stored $\lambda_n$ and went on heading the column $|\lambda_n|$,
which says the table holds something it does not. When what a table stores
changes, three places say what it holds -- the title, the definition and that
header -- and the header is the one nobody rereads.

**`param-latex` replaces the label of the parameter group it sits in**, and
that is the whole of what it does. It is for a parameter whose values are
words -- `normalisation: relative` shown as `$\operatorname{vol}_{\mathrm{rel}}(B_5)$`
-- and it is wrong on a parameter whose value is a number, because there the
value *is* the label. Two tables named each polynomial with it and the column
that should have read `2` read `$L_{\Delta(2,4)}(t)$`, which is the column
header restated one row at a time.

**Entries are shown in the order the document writes them**, so the order the
generator enumerates in is what a reader sees. For an index running over the
negatives that is not the order of $\mathbb Z$: a table led with $a=-50$ and
put $a=2$ -- the row anybody arrives wanting -- halfway down. Enumerate by
$|a|$, positive first, so the small cases are at the top and $a$ and $-a$ are
adjacent, which is what a reader compares.

**A value belongs in a row, not in a comment.** If the table holds a quantity
in two normalisations, both are entries. A number written into an entry
comment answers no search, carries no type, and cannot be cited -- one table
gave each window's $c$ as `$c=-r(r-2)/4=-1.75000000000000000000000$` inside
the comment, where none of that is true of it.

**If the table repeats another's values, say so.** `repeats:
HREF{Other_table}` under `Data properties` means: where these two tables hold
the same number, that one states it first. Search then answers with the
original and names this table beside it, instead of saying the same thing
twice. Only where this table repeats that one's computation -- the same
quantity of the same object, transcribed. Two tables whose values agree
because of a theorem are both worth being told about, and that is the
commoner and more interesting case: Hermite's constant in dimension 8 equals
the Hermite number of $E_8$, and both tables should answer.

**Link the first mention of a thing to what explains it.** A table if the
corpus holds one; otherwise a reference the table already declares --
`CITE{Wiki}` beside the first "Dedekind zeta function" costs nothing and the
article is in `Links` already. The rule is the same in both cases: first
mention, once per section.

**Name the thing, do not point at it.** "The first factor is in the table of
Bernoulli numbers and the second, for $D=5$ and $D=8$, in the table of
generalized Bernoulli numbers" makes a reader count backwards through a
formula with four factors in it, and be wrong. Write "$B_{2m}$ is in ... and
$B_{2m,\chi_D}$ ... in ...". The same goes for "the former", "the latter",
"as above", and "the second one": the symbol is shorter than the phrase
pointing at it, and it stays right when the formula is edited.

**Link a table the first time it is named, and only the first time.** The
convention is Wikipedia's, and it is worth following in both directions. A
family the corpus holds should be a link, not a plain phrase -- somebody
reading about Bernoulli numbers here is one click from the table of them -- and
linking the same name four times in one comment is noise. First mention in a
section, `HREF{Bernoulli_numbers}[the Bernoulli numbers]`; afterwards, plain
text.

**Link the number, not the table, when the sentence means one number.**
`HREF{Feigenbaum_constants}` points at a table of six constants where the
sentence said $\delta$, and `HREF{Golden_ratio}` at three where it said
$\varphi$. A link may name an entry: `HREF{Feigenbaum_constants#delta}[$\delta$]`
and `HREF{Golden_ratio#phi}[$\varphi$]` land on the row. The address after the
`#` is the entry's identity -- its parameter values, comma separated, as a
citation writes them -- so a row of a three-parameter table is
`HREF{slug#Z,1,density}`, and a row of this same table is `HREF{#delta}`.
Read the keys out of the table rather than inventing them; they are the keys
of `Numbers`, not the labels the page prints. Link the whole table when the
sentence means the whole family, which is the commoner case.

**A comment states a fact about the mathematics.** How strong a search hit
would be, how distinctive a value is, what a reader ought to conclude: those
are remarks about the website, and a table full of them reads as apology.
Where a value looks surprising the fact is the explanation -- "the value is
$x$ because $j(\zeta_3)=0$" -- and the sentence after it, explaining that a
search for $x$ therefore means little, is the one to leave out.

**Define notation where you use it.** A comment saying the values are
`U_n(x,-1)` is useless until something says what `U` is. If a symbol appears in
a formula or a comment, its definition belongs in the same table.

**A reference is for citing, not for listing.** `CITE{Koshy}` in a comment
earns its place; a bibliography nobody points at is furniture.

Every table has at least Title, Definition, Tags (two is typical), Links and
Data properties. Prefer an existing tag; a tag list is a way through the
corpus and a tag with one table on it is not. But the list is not finished and
was never meant to be: propose a new one when **three or more tables would
carry it**, counting the others in the same batch, and say which they are. Six
tables on percolation, the Ising model, critical exponents, lattice entropy and
self-avoiding walks were filed under "physics" and "combinatorics" for want of
this, which tells a reader almost nothing about any of them. Use `$...$`
for mathematics, `CITE{key}` for a reference or link, and
`HREF{slug}[caption]` for a table here. Link outward only to sources that will
still exist: Wikipedia, LMFDB, OEIS, MathWorld, mpmath, or a paper.

**A formula states a relation, not who checked it.** "Checked on every
entry", "both were computed and agree" are facts about your run, and in a
`Formulas` section they read as an apology. The field for them exists:
`rigour details`. Put what you verified there, once, and let the formula be a
formula.

**Write in sentences, not in dashes.** A `--` in a definition or a comment is
a parenthesis the reader has to hold open, and the renderer makes no
typographic substitution, so it reaches the page as two hyphens: in a field
full of minus signs that reads as mathematics. Say it with a comma, a colon
or a full stop. "converges uniformly for continuous $f$, which is Bernstein's
proof of the Weierstrass approximation theorem"; "each entry comment says
which argument makes it exact. In $c$ the surd cancels: $r=1+\sqrt6$ gives
$-r(r-2)/4=-(6-1)/4$."

The em dashes a reader does see, in the parameter list and in `Similar
tables`, are the site's own: it sets one between a parameter and its title,
and between a table and the relation glossing it. Those are punctuation the
page supplies, not something an author writes. A hyphen inside a word or a
key is a third thing again and is right as it is: period-doubling, `onset-c`.

**Write sentences, not notes to yourself.** Everything a reader sees is read
as English. "every rule with $n\leq 30$ is here, nodes and weights, both
halves" is four fragments stapled together, and a reader cannot tell whether
both halves of every rule are listed or both halves of something else. The
register is an encyclopedia's: precise, unhurried, and using the table's own
word for a thing -- if the parameter is titled "number of nodes", write nodes,
not points. Telegraphic is not concise; it is unfinished.

**`Programs` and `generate.py` answer different questions.** `Programs` is the
standard incantation in Sage, PARI or mpmath for a reader who wants one more
value. `generate.py` is the program that reproduces and extends *this* table,
attached to it. A table wants both where both apply.

## 5. Types: what a value is, how it is written, what to return

`type` in Data properties says what the table holds. Seven are searchable by
their digits:

| type | holds | written as |
|---|---|---|
| `Z` | integers | `3`, `-1729` |
| `Q` | rationals | `-3/2` |
| `R` | reals | see below |
| `C` | complex | `0.309... + i * 0.951...` |
| `Qp` | p-adics | `2^4 * 111736... + O(2^167)`, or `Q2:1.110` |
| `Z[]`, `Q[]` | polynomials | `x^2 - x - 1` |

A type outside that set is allowed but is *shown and cited rather than found*:
it must also carry a `type name` (T41's four hyperreals are `*R`, "hyperreals").
A misspelling is a typo; a new symbol with a name beside it is somebody
deciding something.

**A real is stored as an interval that contains it**, in one of four forms, and
the first is the one the corpus is written in:

    3.14                 the interval [3.13, 3.15] -- the last digit may be
                         off by one. `12e2` means [1100, 1300].
    [2, 2.3728596]       endpoints, exactly
    3.14 +/- 2e-2        centre and radius, exactly: [3.12, 3.16]
    1p31415              p-notation: 0.31415e1 with the last digit uncertain

A string with no `.` and no `e` is an **exact integer**, not an approximation.
This convention is the whole reason a hundred digits means something: `3.14`
*is* an interval, so a table never has to say separately how far to trust it.

**What `value()` should return**, and what each return means:

| return | recorded as | rigour it can support |
|---|---|---|
| `int`, `ZZ(n)` | exact integer | `exact` |
| `Fraction`, `QQ(a)/b` | exact rational | `exact` |
| `RealBallField(prec)(x)` (arb) | interval, from the ball | `proven` |
| `ComplexBallField(prec)(x)` | complex interval | `proven` |
| `RealIntervalField(prec)` element of nonzero width | interval | `proven` |
| a Sage polynomial | polynomial | `exact` |
| a `Qp` element | p-adic, with its own `O(p^n)` | `proven` |
| a string | taken verbatim | not `proven` |
| a `float` | **refused** — it does not say how precise it is | — |

Prefer **balls** (`RealBallField`, `ComplexBallField`) for anything
transcendental: arb carries the error through every step, so the digits written
are the digits the result supports. `RealIntervalField` is fine when the whole
computation is interval arithmetic (MPFI), and is a trap when it merely wraps a
fixed-precision result — see below.

An entry may also say it is deliberately less precise than the table's
`digits`, by returning `{'number': x, 'digits': 8}`.

## 6. Write a generator

**A generator fills a table; it does not create one.** A new table is made on
the site, by a person, deliberately -- it takes a permanent T-number, a title
in every listing, and a parameter order that can never change because citations
resolve on it. Creating tables through the API is board-only for that reason.
So: agree the title, definition, parameters and type with whoever owns the
database, let them create it, then point a generator at it.

**Division in a generator is where exactness goes.** `sage -python` has no
preparser, so `factorial(30)` is a Python `int` and `factorial(n) / k` is
*float* division -- exact up to 2^53 and quietly wrong after it. A Bessel
polynomial built from its closed form this way had the right coefficients up
to n = 15 and wrong ones from n = 16, differing in the last two digits, which
is exactly the kind of wrong that gets published. Two ways out: build from a
recurrence that only multiplies and adds, or write every division between Sage
rationals, `QQ(a) / QQ(b)`. Note that `c in ZZ` is true of a float that
happens to be integral, so that check will not catch it.

**Name the rings you use, rather than importing `sage.all`.** A modular
passagemath has no `sage.all`, and the generator below runs unchanged on both
it and a full SageMath -- verified against the live table it fills. Importing
`numberdb.sage` first is what makes the direct imports work: a ring module
imported before anything has initialised Sage raises "cannot import name QQ",
and that module does the initialising.

What else the named imports do not bring: root finding (`sage.numerical.optimize`), `RealBall.str`, and `QQ('1.25')` -- a rational is built from a decimal *string* by hand, `QQ(125)/10**2`, or it raises. Bisection and a rational built from its digits are usually the shortest way past all three.

```python
import sys
import numberdb.sage as numberdb          # import this before the rings
from sage.rings.rational_field import QQ
from sage.rings.complex_arb import ComplexBallField

WORKING_GUARD = 64          # bits beyond what the digits need, measured

class CompleteEllipticK(numberdb.Generator):
    table = 'T25'           # the table must already exist
    parameters = ('m',)
    type = 'R'              # Z, Q, R, C, Qp, Z[], Q[]
    digits = 100
    rigour = 'proven'

    def enumerate(self, denominator=50):
        for b in range(1, denominator + 1):
            for a in range(0, denominator + 1):
                m = QQ(a) / QQ(b)
                if m >= 1 or m.denominator() != b:
                    continue
                yield {'m': str(m)}

    def value(self, params, digits):
        field = ComplexBallField(numberdb.bits(digits, losing=WORKING_GUARD))
        value = field(QQ(params['m'])).elliptic_k()
        assert value.imag().contains_zero()
        return value.real()

if __name__ == '__main__':
    generator = CompleteEllipticK()
    if '--publish' in sys.argv:
        print(generator.publish(message='recomputed in ball arithmetic'))
    else:
        report = generator.verify()
        print(report)
        sys.exit(0 if report.ok else 1)
```

- `verify()` recomputes and compares against the stored table. Needs no key,
  writes nothing. `verify(sample=None)` checks every entry.
- `preview()` computes and compares and sends nothing — the right thing to run
  unattended.
- `publish()` writes. Needs `NUMBERDB_API_KEY`.
- `numberdb.bits(digits, losing=n)` converts decimal digits to bits and adds a
  guard. State the guard as a constant with the measurement behind it: how many
  digits the worst entry actually retained. Do not tune it until the run stops
  complaining.

**Open the file with the commands to run it.** The generator is attached to the
table and downloaded from it, by somebody who has neither this repository nor a
way to guess:

```
Run it with SageMath:

    $ sage -pip install numberdb          # once
    $ sage -python generate.py            # check the table against this code
    $ sage -python generate.py --publish  # send it, with NUMBERDB_API_KEY set
```

Then say what the file does and, where it matters, what was decided and why: a
working precision that was measured, a convention that had to be chosen, an
error bound and where it comes from.

## 7. Rigour: say how well the digits are known

One value per table, in `rigour`. The first five are ordered, weakest last.

| level | means |
|---|---|
| `exact` | an integer, rational or polynomial. No precision to choose. |
| `proven` | interval or ball arithmetic throughout, so the digits follow from the width of the result. |
| `assumed-bound` | fixed precision with an error bound you assert *and justify*. |
| `heuristic (agreement-checked)` | computed at two or more precisions, keeping the digits they agree on. |
| `heuristic` | one computation and a guard chosen by judgement. |
| `measured` | not computed at all — an experimental value. Not on the scale. |

**`proven` is enforced, in one direction**: a value carrying no error of its
own is refused. That refusal exists because of the commonest mistake in this
field:

```python
RealIntervalField(prec)(some_float_function(x))   # width zero: claims exactness
```

Wrapping a fixed-precision result in an interval field does not make it an
enclosure. Check with `arb`/ball arithmetic (`ComplexBallField`,
`RealBallField`) rather than `RealIntervalField` around a float — arb
implements a great deal (`elliptic_k`, `barnes_g`, `zeta`, `zetaderiv`, `agm`,
`airy_ai`, `bessel_J`, Hurwitz `zeta(s, a)`, `sin_integral`, …).

Ball arithmetic also disposes of the argument-rounding problem: `CBF(QQ(1)/3)`
is a ball *containing* 1/3, so an argument you cannot represent is inside the
bound rather than outside the claim.

When you cannot bound it, say so — a weaker level honestly stated is a
contribution, a false `proven` is not:

```python
numberdb.agreeing(lambda working: compute(params, working), at=(150, 200))
numberdb.assume_accurate(value, ulps=2, because='PARI ellL1 at 38 digits; ...')
```

`agreeing` takes decimal digits, not bits, and does not escalate on its own:
the file attached to a table is meant to say how the numbers were made.
`assume_accurate` requires `because` — checked against documentation, most
libraries state no accuracy at all.

## 8. What the refusals mean

The package stops rather than guesses. Each of these has caught a real error:

- **"N digits were asked for and this value carries M"** — the working
  precision was too low, or a field was built in digits where Sage counts bits
  (a factor of ~3.3). Raise the guard; do not lower `digits` silently.
- **"rigour is 'proven', and this value carries no error of its own"** — see
  above. Compute in ball arithmetic, or state the weaker level.
- **"the table holds X and this run produced Y"** — a real disagreement.
  Investigate before passing `correcting=True`; a table once held 200 digits of
  which the last three were wrong, and it was found exactly here.
- **"the table holds N digits and this run produced M"** — you are about to
  shorten stored values. `lowering=True` only if the stored precision was never
  justified.

## 9. Publishing, and what happens next

**A table is proposed, filled, reviewed and then public**, and those are four
separate acts:

1. **Propose it as a draft** -- `POST /api/tables` with `X-Draft: yes`, or made
   on the site. It takes its permanent T-number at once, and keeps it: a
   generator is written against that number while the table is still being set
   up. A draft is invisible, in no listing, answers no search, and may have no
   numbers in it yet, because the prose is written first.
2. **Fill it** with a generator pointed at that T-number.
3. **Offer it for review** when it is finished -- the button on `/drafts`, or
   `X-Draft: ready` at creation for a run that proposes and fills in one go. A
   draft in progress asks for nobody's attention; offering it says the work is
   done. Nothing can work that out from outside, which is why it is a
   statement rather than a rule, and it can be withdrawn.
4. **Somebody reviews it.** An offered draft appears in the review queue as
   "waiting to be published", and confirming it publishes it -- the two are
   the same act, since what they have in common is that somebody competent
   looked.
5. **It is public**, and its values answer search by number.

Do not expect to do step 4. It is the point at which a person takes
responsibility for a table existing.

- Publishing needs the owner's API key. Do not ask for it, and do not put a key
  in a file you commit.
- **Say what you are.** If an assistant is running the publish, set

      export NUMBERDB_ASSISTED_BY=claude-opus-5     # or codex-cli, or ...

  and the revision records it, beside the generator and the Sage and package
  versions, where readers and reviewers already see it. Do not write the name
  into the generator instead: the file outlives the run, and a name hard-coded
  there keeps claiming one tool's work after another tool edits and republishes
  it.

  The author of a submission is the person whose key published it -- authorship
  is accountability, and a model can neither answer for a wrong value nor agree
  to the licence. What an assistant did is a disclosed method, not a
  co-authorship. Disclose when it made a decision a reader would otherwise
  attribute to a person: chose the convention, the range or the
  parameterisation, wrote the definition, wrote the generator. Not for
  formatting or renaming.
- A "table wanted" issue is answered **in the issue**, not in the table. Say
  there that the table exists, what it covers and what it does not; that is
  what the person who asked is waiting to read. The table carries no trace of
  the request: it is encyclopedic, and who asked for it is not a fact about
  the mathematics. Cite the issue only in the rare case that it contains
  something a reader of the mathematics needs -- a derivation, a convention,
  a source -- and then cite that content, not the request. Discussion of a
  table has its own home besides the issue: every table has a discussion
  page.
- Values are held out of search by number until a board member reviews them.
  That is deliberate: a reader looking at a table can see an entry is
  unreviewed, and somebody typing digits into a search box cannot.
- Correcting values that are already public is a human decision. Show the
  measured discrepancy — in units of the last place — and ask.

## 10. Check the work

- `verify(sample=None)` after publishing.
- `manage.py sweep_arb` (server-side) recomputes stored values from
  independent definitions and reports anything the site's own parsers would
  read as a different number.
- **Check new values against something independent.** Not the code that
  produced them: the Fibonacci polynomials were checked against Sage's own
  `fibonacci()` at x = 1, and against identities that tie the two tables
  together -- L_n = F_(n-1) + F_(n+1), F_2n = F_n L_n, gcd(F_m, F_n) =
  F_gcd(m,n). A family with known identities gives you a free test suite.
- **Verify a claim before writing it into a table, including one somebody
  suggested.** A suggestion is a hypothesis. "These are orthogonal polynomials,
  tag them so" sounds obviously right and is false: Favard's condition fails,
  and what holds instead is an *indefinite* pairing on the imaginary axis. The
  table now says that, which is worth more than either the wrong tag or
  silence.
- **A measurement needs a control that returns a known answer.** The first
  attempt at that orthogonality check used Simpson's rule on a weight with an
  endpoint singularity and reported -0.023 for a pairing that is exactly zero.
  Worse, the control -- the Chebyshev family, whose answers are known --
  silently returned zero for everything because of a coercion error, so it
  confirmed nothing while looking like confirmation. Run the control first and
  check it gives the answer you already know.
- **An enclosure that contains everything agrees with everything.** The same
  rule, in the form it takes in ball arithmetic, where it is easy to miss
  because the check reports success. `psi(-1/2)` was compared against its
  closed form and reported `overlaps = True` while its value was `nan` -- arb
  takes a power as `exp(y log x)`, so a negative base with a negative exponent
  is not a number, and a nan ball overlaps every interval there is. A ball
  merely *wide* does the same quietly: one of radius 1.4e18 contains zero, and
  every other answer. Before believing a comparison, check that both sides are
  finite and that the radius is small enough to mean anything.
- **Check the claim the digits make, not the value at their midpoint.** What a
  written number claims is the interval its last place denotes, so the thing
  to verify is that a zero lies *inside* it -- the function taking opposite
  signs at the two ends. Feeding the midpoint back in and asking for something
  near zero measures conditioning instead: near the poles of `psi^(n)` the
  derivative is enormous, and a point 10^-100 from a true zero evaluates to
  order one. The same 607 values failed that test twenty times in sixty-three
  and passed the bracket test 1050 times out of 1050.
- If your recomputation disagrees with a stored value, suspect your
  recomputation first. In this corpus every disagreement after the first two
  was the checker's fault: a Teichmüller limit that collapses at p = 2, an
  exponential series that does not converge where the Artin–Hasse exponential
  is defined, a dodecahedron's inradius out by a factor of √5.
