ikawaha / kuromoji.js

JavaScript implementation of Japanese morphological analyzer

Geek Repo:Geek Repo

Github PK Tool:Github PK Tool

kuromoji.js

JavaScript implementation of Japanese morphological analyzer. This is a pure JavaScript porting of Kuromoji.

You can see how kuromoji.js works in demo site.

Directory

Directory tree is as follows:

demo/        -- Demo
dist/        -- Released objects
  browser/   -- JavaScript file for browser (Browserified)
  dict/      -- Dictionaries for tokenizer (gzipped)
  node/      -- JavaScript file for Node.js
example/     -- Examples to use in Node.js
jsdoc/       -- API Documentation
scripts/     -- Node.js scripts
src/         -- JavaScript source
test/        -- Unit test

Usage

You can tokenize sentences by only 5 lines of code. If you need working examples, see also files under demo or example directory.

Node.js

Install by npm package manager:

npm install kuromoji

Load this library as follows:

var kuromoji = require("kuromoji");

You can prepare tokenizer like this:

kuromoji.builder({ dicPath: "path/to/dictionary/dir/" }).build(function (tokenizer) {
    // tokenizer is ready
    var path = tokenizer.tokenize("すもももももももものうち");
    console.log(path);
});

Browser

Only you need dist/browser/kuromoji.js, and dist/dict/*.dat.gz files

Install by Bower package manager:

bower install kuromoji

Or you can use kuromoji.js file and dictionary files from GitHub repository.

In your HTML:

<script src="url/to/kuromoji.js"></script>

In your JavaScript:

kuromoji.builder({ dicPath: "/url/to/dictionary/dir/" }).build(function (tokenizer) {
    // tokenizer is ready
    var path = tokenizer.tokenize("すもももももももものうち");
    console.log(path);
});

API

The function tokenize() returns JSON array like this:

[ {
    word_id: 509800,          // 辞書内での単語ID
    word_type: 'KNOWN',       // 単語タイプ(辞書に登録されている単語ならKNOWN, 未知語ならUNKNOWN)
    word_position: 1,         // 単語の開始位置
    surface_form: '黒文字',    // 表層形
    pos: '名詞',               // 品詞
    pos_detail_1: '一般',      // 品詞細分類1
    pos_detail_2: '*',        // 品詞細分類2
    pos_detail_3: '*',        // 品詞細分類3
    conjugated_type: '*',     // 活用型
    conjugated_form: '*',     // 活用形
    basic_form: '黒文字',      // 基本形
    reading: 'クロモジ',       // 読み
    pronunciation: 'クロモジ'  // 発音
  } ]

(This is defined by src/util/IpadicFormatter.js)

See also JSDoc under jsdoc directory.

About

JavaScript implementation of Japanese morphological analyzer


Languages

Language:JavaScript 93.6%Language:Shell 4.5%Language:CSS 1.9%