DEVELOPER TOOL

i18nExtractionTool

Quickly process HTML files to extract text content and generate standardized internationalization JSON resource files, streamlining multilingual website development.

index.html
<div class="header">
  <h1>Welcome</h1>
  <p>Sample text</p>
  <button>Start</button>
</div>
en.json
{
  "header.title": "Welcome",
  "header.description": "Sample text",
  "header.button": "Start"
}
KEY FEATURES

Core Capabilities

Two powerful flows: extract copy from HTML into JSON, and restore original HTML from keyed templates using JSON

HTML Parsing

Powerful HTML parsing capabilities that accurately identify and extract text content from various tags while automatically filtering scripts and style code

Standard JSON Output

Automatically generates i18n-standard JSON resource files with key-value mappings and nested structures for direct integration

HTML Restore

Provide translation JSON and t(...) keyed HTML to rebuild original HTML and recover copy without altering your templates

SIMPLE PROCESS

Three Simple Steps

We've simplified the HTML internationalization process to complete the entire workflow in just minutes

1

Upload HTML Files

Select or drag and drop your HTML files for processing. Supports multiple files for batch processing.

2

Configure Extraction Options

Set extraction rules, prefixes, and exclusions to customize the output structure for specific requirements.

3

Export JSON Files

Generate and download JSON format internationalization resource files for immediate integration into your project.

EXAMPLES

See It In Action

Experience how our tool transforms HTML content into structured internationalization resources

Input: HTML Filesource.html
<div class="header">
  <h1>Welcome to our service</h1>
  <p>This is a sample text to demonstrate i18n extraction.</p>
  <button title="Click to continue">Get Started</button>
</div>
Output: JSON Filetranslations.json
{
  "header.title": "Welcome to our service",
  "header.description": "This is a sample text to demonstrate i18n extraction.",
  "header.button.text": "Get Started",
  "header.button.title": "Click to continue"
}

Integration Example

After extraction, you can easily integrate the JSON files with popular i18n libraries:

// React + i18next (example)
import i18n from 'i18next'
import { useTranslation } from 'react-i18next'
import translations from './translations.json'

i18n.init({ resources: { en: { translation: translations } } })

function Header() {
  const { t } = useTranslation()
  return (
    <div className="header">
      <h1>{t('header.title')}</h1>
      <p>{t('header.description')}</p>
      <button title={t('header.button.title')}>{t('header.button.text')}</button>
    </div>
  )
}

Ready to Internationalize?

Start extracting internationalization resources from your HTML files today and streamline your multilingual development

Start Using Now

Parse HTML for i18n Strings Without Leaving Your Machine

Internationalizing a web application starts with extraction: finding every user-facing string in your markup and organizing it into a translation file. Done by hand this is tedious and error-prone; done on a server you must upload your source. This parser does it locally - paste your HTML (or a template from Vue, React, Svelte, or plain markup) and it extracts every candidate string, deduplicates it, and outputs clean JSON (or CSV) ready to hand to translators or a CMS.

The extraction is pattern-aware: it skips script and style blocks, strips attributes that are not user-visible, flags strings that look like URLs or placeholders, and preserves interpolation markers (, {count}) so translators see exactly what gets filled at runtime. The result is a structured, sorted list where each string has a stable key, context (surrounding text), and occurrence count.

Why local parsing matters for i18n workflows

Translation files are the backbone of your localization pipeline: they go to machine translation, to human translators, into your CI. Keeping the extraction step offline means your source markup - which often contains product names, unreleased features, and proprietary copy - is never sent to a third party. Re-running the parser after each sprint gives you a diff of newly added strings, which is exactly what a continuous localization workflow needs.

Common Questions

Does this tool work with framework templates?
Yes. It understands plain HTML plus common template syntax (Vue, React/JSX, Svelte) and treats interpolation markers as placeholders rather than literal text.
What format does the output come in?
Clean JSON by default, with CSV available for spreadsheet-based translation workflows. Every string gets a stable key, context, and occurrence count.
Is my code uploaded anywhere?
No. Parsing runs entirely in your browser, so your markup never leaves your machine.
How are duplicates handled?
Identical strings are deduplicated and carry an occurrence count, so translators translate each unique string once and the tool maps it back to every usage.
Can I use this for a new string diff each release?
Yes. Run the parser before and after a change and compare outputs; new or removed keys are the delta your translators need.