Methodology
How DBEmoji counts and verifies emoji
Every number on this site is derived from Unicode Consortium source files by a script in this repository, not entered by hand. This page states the source, the counting rules, and where DBEmoji deliberately differs from Unicode’s own figures.
- Unicode Emoji
- 16.0
- Entries
- 4,070
- RGI emoji
- 3,781
- Verified
- 2026-09-18
Where the data comes from
The dataset is built from emoji-test.txt for Unicode Emoji 16.0 (published 2024-08-14, 23:51:54 GMT) and from UnicodeData.txt for official character names. Both files are vendored into the repository so a build is reproducible and reviewable; refreshing them is a single command, and the dataset is then re-derived and re-validated.
Codepoints, Unicode names, CLDR short names, Unicode groups and subgroups, and version numbers are copied verbatim from those files. Nothing in that list is edited.
What counts as one emoji
DBEmoji ships 4,070 entries. Unicode Emoji 16.0 defines 3,790 sequences Recommended for General Interchange (RGI). The two numbers differ, and here is exactly why.
- Skin-tone variants are counted separately. 1,840 entries carry a Fitzpatrick modifier. Unicode also counts these individually, so they are not the source of the difference — but they are why the total is far larger than the ~1,900 distinct pictures most people picture when they hear “emoji”.
- ZWJ sequences are counted separately. 1,615 entries are zero-width-joiner sequences (👨👩👧, 🧑🚒). Each is its own entry because each has its own codepoint sequence and its own page.
- Components are included. The 9 skin-tone and hair-style modifiers have entries, flagged as components. They are not meant to be sent on their own.
- Flags are counted as emoji. 289 entries sit in the flags collection, including regional tag sequences.
- 280 entries are outside Unicode Emoji 16.0. These are text symbols, non-RGI sequences, and characters proposed for a future release. They are the whole of the difference between our total and Unicode’s.
If you want the figure that matches Unicode, use 3,781 RGI emoji plus 9 components. If you want the number of pages on this site, use 4,070.
Entries outside the Unicode emoji set
280 entries are present in this database but not in Unicode Emoji 16.0. They fall into three groups: text symbols that are not emoji (chess pieces, geometric shapes, power symbols), sequences that are syntactically valid but not recommended (skin-tone variants of multi-person emoji, most subdivision flags), and characters proposed for a future Unicode release.
Every one of them is labelled “Not in the Unicode emoji set” on its page, because on mainstream platforms they render as plain text or an empty box rather than a colour emoji. Their pages remain reachable and their data is preserved, but they carry noindex and are excluded from the sitemap: a search result that leads to a character the reader cannot actually use is a dead end.
Unicode categories vs DBEmoji collections
Unicode defines 10 emoji groups. DBEmoji browses through 15, because Unicode’s “People & Body” group alone holds more than 2,000 sequences and is not practical to browse as one list. The extra 5 are DBEmoji collections, not part of the standard, and every category page says which it is.
Unicode groups
DBEmoji collections
- Face — from People & Body
- Activity — from People & Body
- Sport — from People & Body
- Role — from People & Body
- Family & Love — from People & Body
Collections re-group emoji; they never re-label them. An emoji’s Unicode group and subgroup are always shown on its page, unchanged.
Tags, meanings and tone
Tags, meanings, common uses and tone are DBEmoji editorial interpretation, labelled as such wherever they appear. They are not part of the Unicode Standard and should not be cited as if they were.
Theme tags are validated against each emoji’s own Unicode data: a tag is kept only when Unicode’s subgroup licenses it, or when a whole word supporting it appears in the CLDR short name or the emoji’s shortcodes. The dataset this site inherited was seeded from a keyword list that matched substrings, which produced results like 😂 “face with tears of joy” being tagged drink — “tea” is a substring of “tear”. All 22 governed theme tags are now re-checked on every build, and a validation script fails the build if an unsupported one reappears.
Where a mood is recorded, it overrides contradicting mood tags. An emoji is never described as both cheerful and sad on the same page, and use cases are drawn from Unicode’s subgroup rather than from the tag list, so a laughing emoji is never suggested for a condolence message.
Where we have no confident basis for a section, the section is omitted rather than filled. See the editorial policy.
Keeping up with Unicode
When the Unicode Consortium publishes a new Emoji release, the vendored source files are refreshed, the dataset is re-derived, and the validator checks every entry against the new release before anything ships. The version and verification date shown across the site come from that build, so they cannot drift from what is actually in the database.
Current: Unicode Emoji 16.0 · dataset 2026-09-18 · verified 2026-09-18
