Query Details

Detecting Non Latin Characters With KQL

Query

let foreignCharPattern = @"[\p{Han}\p{Hiragana}\p{Katakana}\p{Hangul}\p{Arabic}\p{Hebrew}\p{Devanagari}\p{Bengali}\p{Tamil}\p{Thai}]";

About this query

Detecting Non-Latin Characters with KQL

Query Information

I recently discovered how easy it is to detect characters from non-Latin writing systems using Unicode script categories:

A simple regular expression can identify scripts such as Chinese, Japanese, Korean, Arabic, Hebrew, and several Indic scripts. A small but useful insight for text analysis and data quality checks.

Explanation

This KQL query is designed to identify email subjects that mention file-sharing platforms like Teams or SharePoint and contain non-Latin characters. Here's a simple breakdown:

  1. Pattern Definition: The query defines a pattern (foreignCharPattern) using a regular expression that matches characters from various non-Latin scripts, such as Chinese, Japanese, Korean, Arabic, Hebrew, and several Indic scripts.

  2. Platform Keywords: It specifies a list of file-sharing platforms (SharingPlatform) to look for in the email subjects.

  3. Data Filtering:

    • It searches the EmailEvents table for emails received in the last 30 days.
    • It filters these emails to find subjects that mention any of the specified platforms.
    • It further filters to keep only those subjects that contain at least one character from the non-Latin scripts defined in the pattern.
  4. Character Extraction: The query uses extract_all to list all non-Latin characters found in each subject, making it easy to see which scripts are used and how many such characters are present.

  5. Purpose: This query helps identify potentially suspicious emails, such as phishing attempts, which might use mixed scripts to deceive recipients. It can also highlight emails from international or multilingual sources.

Overall, the query is a tool for detecting and analyzing emails with mixed script content, which could be unusual or suspicious in organizations that typically use only Latin scripts.