Research & Unicode Data Methodology
This methodology defines how UnicodeSymbols.com turns runtime character data and primary standards into checked public reference pages.
What this page covers
Properties are calculated
Names, categories, code points, and bytes come from the stored text and runtime data.
Collections are editorial
Subject groups help people browse; they are not official Unicode properties.
Claims keep their boundary
The review separates encoded facts, font behavior, and cultural interpretation.
Where character properties come from
The application derives formal names, general categories, bidirectional classes, combining classes, East Asian width values, and UTF-8 encodings from the Unicode database supplied by its Python runtime. Code-point notation and byte output are calculated from the stored text rather than copied from a third-party symbol page.
What the collections mean
The twelve subject collections are editorial browsing aids, not normative Unicode properties. A character can appear in more than one collection when it performs a genuine role in each subject. Collection totals describe this site's selected inventory; they are not claims about the size of the Unicode Standard.
How an editorial claim is reviewed
- Identify the exact code point or ordered sequence.
- Separate standard-defined properties from font behavior and cultural interpretation.
- Check the relevant Unicode report, encoding specification, or web standard.
- State uncertainty where a source does not support one universal meaning.
- Test internal links, metadata, and rendered examples before release.
Browser-tool boundaries
The inspector, normalization checker, and converter run their transformations in JavaScript in the open tab. The inspector downloads a cacheable dictionary of catalog names; it does not send the pasted specimen to obtain a result. Search is different: its query is sent to the site's bounded search endpoint, and the resulting search page carries a noindex instruction.
Corrections and review dates
A reported issue should include a URL, an untouched copy of the encoded value, the intended result, the observed result, and relevant software details. Review dates change after substantive verification, not merely to create an appearance of freshness. Read the separate editorial policy for publication safeguards, consult the source register, or use the correction channel.