An EPUB is a zip archive. Rename book.epub to book.zip, unzip it, and you
get a folder of XHTML files, a stylesheet, and a handful of small XML documents
that tell a reading device how to assemble them into a book.
That is worth knowing for one practical reason: when a reader says your chapters do not appear in the Go To menu, or a store rejects your upload, the fault is almost always in one of those small XML files rather than in your writing.
The whole archive
Here is the complete file list from an EPUB this site produced a moment ago — a short book with a title page and one chapter:
mimetype
META-INF/container.xml
META-INF/com.apple.ibooks.display-options.xml
EPUB/content.opf
EPUB/toc.ncx
EPUB/nav.xhtml
EPUB/text/title_page.xhtml
EPUB/text/ch001.xhtml
EPUB/styles/stylesheet1.css
Nine files. A four-hundred-page novel has the same nine plus one XHTML file per chapter. There is nothing else in an ebook.
mimetype
Twenty bytes, containing the string application/epub+zip and no newline.
It has two rules that no other file in the archive has: it must be the first entry in the zip, and it must be stored uncompressed. This is so that a program can identify an EPUB by reading the first few dozen bytes without unzipping anything.
It is also the single most common way a hand-made EPUB fails validation. Zipping a folder with the operating system's right-click Compress almost never produces this layout, which is why "I zipped it myself and it will not open" is such a familiar complaint.
META-INF/container.xml
The pointer. It exists to answer one question: where is the package document?
<rootfiles>
<rootfile full-path="EPUB/content.opf"
media-type="application/oebps-package+xml" />
</rootfiles>
That indirection is why the rest of the archive can be laid out however the
producer likes. EPUB/content.opf is a convention, not a requirement — the only
fixed paths in the whole format are mimetype and META-INF/container.xml.
content.opf — the package document
The book's identity and its contents, in three sections.
Metadata is what the store reads. This is where your title and author actually live:
<dc:identifier id="epub-id-1">urn:uuid:439bf7c8-…</dc:identifier>
<dc:title id="epub-title-1">The Glass Hours</dc:title>
<dc:language>en-US</dc:language>
<dc:creator id="epub-creator-1">Marta Reyes</dc:creator>
If you have ever uploaded an ebook and found it listed as "Untitled", this is
the reason. The filename is not the title; dc:title is. A converter that was
never told the book's name will happily produce a valid EPUB called Untitled.
The manifest lists every file in the archive. A file that exists but is not manifested is invisible to the reading system; a file that is manifested but does not exist is a validation error. This is the source of the most common EPUBCheck complaint of all.
The spine is the reading order — the sequence of pages you get by turning one page at a time from the beginning.
<spine toc="ncx">
<itemref idref="title_page_xhtml" linear="yes" />
<itemref idref="nav" />
<itemref idref="ch001_xhtml" />
</spine>
The spine is order, not structure. It knows chapter one comes after the title page; it does not know that chapter one is a chapter.
nav.xhtml — the navigation document
This is the file that makes a Go To menu work, and the one most worth understanding.
<nav epub:type="toc" role="doc-toc" id="toc">
<ol class="toc">
<li><a href="text/ch001.xhtml#chapter-one">Chapter One</a></li>
</ol>
</nav>
It is a real page of XHTML containing an ordered list of links. Every entry a reader sees in the chapter menu is a line in that list. If your ebook opens fine but the chapter menu is empty or shows one entry called "Book", the navigation document is why.
And here is the thing that catches people: nav entries come from headings, not from appearance. A chapter title that you set in Word by selecting the line and making it 18pt, bold and centred is, as far as any converter can tell, a paragraph that happens to look large. It gets no nav entry. A chapter title styled as Heading 1 gets one. The two look identical on your screen and produce different ebooks.
The same file usually carries a second, hidden landmarks nav, which is how a
device knows which page is the cover and where the body text begins.
toc.ncx
The navigation document as EPUB 2 defined it, kept for old devices. EPUB 3
replaced it with nav.xhtml, but plenty of hardware in circulation predates
that, so most producers ship both. It is redundant on anything made in the last
decade and harmless everywhere.
The XHTML files, and the stylesheet
Your actual book: one file per chapter, by convention. They are XHTML, which is
stricter than the HTML on the web — every tag must be closed and every attribute
quoted. An unclosed <i> that a browser would silently forgive is a hard error
here.
The stylesheet is a suggestion rather than an instruction. The reader chose their font size, their typeface and their margins, and those choices win. This is not a limitation to work around; it is what a reflowable ebook is.
What to do with this
Two things, if something is wrong.
Rename the file to .zip, unzip it, and open nav.xhtml in a text editor. If
your chapters are not in that list, they will not be in the reader's menu, and
no amount of adjusting the visual layout will change that — the fix is upstream,
in how the headings were marked in the source document.
And run EPUBCheck before uploading anywhere. It is the reference validator, and every store runs some version of the same checks; it is much less painful to read its output yourself than to have an upload rejected without explanation.