Synced from
geml-spec/gemlate6829e7— edit it there.
Write a GEML parser in your language
English | 中文
The highest-impact thing you can do for GEML: implement it from the spec in another language. Two independent parsers that agree are the proof the spec is unambiguous — and what makes GEML a standard, not one repo.
It's a weekend project, and you can self-certify: reproduce a set of JSON conformance cases, then parse the spec's own .geml file cleanly. Building one? Open an implementation issue — we'll help and link it. No need to finish it all at once.
Conformant means five things
Your parser turns GEML source into a document model (blocks and inline nodes). It's a conforming parser (§8.2) when:
- It reproduces every case in the conformance suite (below).
- It parses the dogfood spec
GEML-spec.gemlwith zeroerrordiagnostics — that exercises fences, attributes, references, tables, charts, and metadata. - References resolve (§8): every
#idis unique, and every[[#id]],[[doc.geml#id]],[text](#id),[^id], a table's or chart'ssrc=/data=, and anembed'ssrc=points at something real. - It normalizes its input exactly as §0.5 says: UTF-8, strip one leading BOM, line endings → LF,
U+0000→U+FFFD. Four lines of code, and skipping them is the most common way a second implementation silently disagrees with the reference on real-world files. - Every diagnostic carries the code and severity from Appendix A. The message text is yours to word (or translate); the code is the contract, and it's what makes your error paths testable against ours.
The suite pins the ambiguous parts (inline emphasis, list nesting); the dogfood covers the rest.
Two things you also owe an untrusted document, per §9: bound your recursion depth (block, list, and inline nesting — emit the *-nesting-too-deep error and keep going, never blow the stack), and neutralize non-http/https/mailto/tel URL schemes when you build the model, not at the rendering sink.
The conformance suite
Plain JSON — copy it in and run it with your own harness. In geml-parser/test/conformance/:
| File | Covers |
|---|---|
inline.json | emphasis / strong / strikethrough, the rule of three, escapes, nesting |
precedence.json | atom vs. emphasis order: code, math, links, images, footnotes, breaks |
lists.json | ordered/unordered, start, indentation nesting, tight vs. loose, task markers |
Each case is { name, geml, want }:
{ "name": "em inside strong", "geml": "**a *b* c**", "want": "strong(\"a \" em(\"b\") \" c\")" }want is a projection of the parsed model — a compact string. Project your model the same way and assert it equals want.
Projection format
text "abc" (JSON-quoted)
emphasis em( … )
strong strong( … )
strikethrough s( … )
code span code("…")
inline math math("…")
hard break br
image img("src")
link link("target" children…) target = href | #anchor | doc#anchor
auto-reference ref("target")
footnote ref fn("id")
paragraph children, space-separated
heading level N hN( children… )
list ul[…] | ol[…] suffix "*" if loose, "@N" if ordered start ≠ 1
list item li(…) | li[ ](https://github.com/geml-spec/geml/blob/e6829e7c41c38c308ae9de93993878c2c4359eac/docs/…) | li[x](https://github.com/geml-spec/geml/blob/e6829e7c41c38c308ae9de93993878c2c4359eac/docs/…) nested lists appended insideChildren join with one space. _project.mjs is the reference projection — the tiebreaker if anything is unclear. impl2.mjs is a full parser + projection written only from the spec (a few hundred lines) — a worked example of what you're building.
Build order
Each step maps to a spec section and what tests it. Do them incrementally.
- Normalize the input (§0.5) — decode UTF-8, strip one leading BOM, collapse CRLF/CR to LF, replace
U+0000. Do this first and everything downstream gets simpler; every step preserves the line count, so you can still index the original bytes by line. →preliminaries - Fences + block scanner (§2–§3) — a
=-run opens a block, an equal-length run closes it, a longer fence nests; ATX headings, lists, paragraphs,%%lines. → dogfood - Attribute object
{#id .class key=val}(§4) — value typing; a bare word with no=is a boolean flag. → dogfood meta+{{key}}interpolation (§3–§4) — substitute in flow source text, skipping the verbatim atoms (code spans, inline math) and escaped{{key}}. →interp.json- Inline (§5) — emphasis/strong/strike (rule of three), code, math, links, auto-refs, footnotes, images, breaks, escapes. The hard part; lean on the fixtures. →
inline.json,precedence.json - Lists (§2.1) — ordering,
start, nesting, tight/loose,[ ]/[x]. →lists.json - References + checks (§8) — collect ids, resolve refs, error on duplicates and dangling. → dogfood
- Tables and views (§6, §6.1) — pipe grid and
format=csv/tsvparse to one model, which holds facts; aviewderives from one withcompute=,summary=,where=,order=,limit=,select=,by=/aggregate=. → dogfood - Diagrams & charts (§7) — diagram bodies are never interpreted;
geml-chart data=#idcharts a table by reference. → dogfood
Step 0 plus 1–5 give a useful parser. 6 is what makes GEML GEML. 7–8 are the payoff.
Self-certify
for file in [inline.json, precedence.json, lists.json, interp.json]:
for case in load(file):
assert project(parse(case.geml)) == case.want
doc = parse(read("spec/in_geml_format/GEML-spec.geml"))
assert no "error" diagnostic in doc.diagnostics
# §0.5 — the same document, four ways, must give the same model
base = "# T\n\n- a\n- b\n"
assert parse(base) == parse("" + base) == parse(base.replace("\n", "\r\n"))
# Appendix A — every code you emit is in the catalogue, at its declared severity
for d in parse(read("spec/in_geml_format/GEML-spec.geml")).diagnostics + your_error_fixtures():
assert d.code in APPENDIX_A and d.severity == APPENDIX_A[d.code]Suite green + dogfood clean + §0.5 + Appendix A = an independent, conformant GEML parser. Open an issue or PR (CONTRIBUTING.md) and we'll add it to the README.
Reference
- Spec:
GEML-spec.md(§0–§9 + Appendices A/B) +GEML-history-spec.md. GEML-spec.geml— the spec in GEML; your end-to-end test.geml-parser/— the reference implementation (a guide; the spec is the definition).