Announcement

Collapse
No announcement yet.
X
  • Filter
  • Time
  • Show
Clear All
new posts

  • Display of contents of string variable containing text in Hindi using Devanāgarī characters

    Hi all,

    I am reading a file from the web into Stata that contains Hindi text in Devanāgarī script. Some Devanāgarī characters display correctly in Stata graphs, but not in my Results window. This differs from Arabic, Chinese, Japanese, and Cyrillic scripts, which render correctly in both graphs and the Results window.

    When I copy and paste the Results window contents into my editor (Emacs on Linux), all characters appear correctly, suggesting a rendering rather than an encoding problem. I suspect the issue is related to the font used in the Results window, but I have not been able to identify a font setting that fixes it.

    Any hints how to solve this would be greatly appreciated.

    Here is an example of how to read the file, in case anyone wants to try it out:

    Code:
    . copy https://hi.wiktionary.org/wiki/%E0%A4%B5%E0%A4%BF%E0%A4%95%E0%A5%8D%E0%A4%B7%E0%A4%A8%E0%A4%B0%E0%A5%80:%E0%A4%A6%E0%A5%87%E0%A4%B6%E0%A5%8B%E0%A4%82/%E0%A4%B0%E0%A4%BE%E0%A4%B7%E0%A5%8D%E0%A4%9F%E0%A5%8D%E0%A4%B0%E0%A5%8B%E0%A4%82_%E0%A4%95%E0%A5%87_%E0%A4%A8%E0%A4%BE%E0%A4%AE country_code_hi.txt
    
    . filefilter country_code_hi.txt x.txt, from(`"<tr "') to(\n) replace
    . import delimited using x.txt, stripquotes(yes) delimiters("\t") encoding(UTF-8) clear
    . gen iso2 = ustrtrim(ustrregexs(1)) if ustrregexm(v1,`"<td.+?>([A-X]+)</td><td.+?>.+?</td><td.+?>(.+?)</td>"')
    . gen name_hi = ustrtrim(ustrregexs(2)) if ustrregexm(v1,`"<td.+?>([A-X]+)</td><td.+?>.+?</td><td.+?>(.+?)</td>"')
    . keep if !mi(iso2,name_hi)
    . keep iso2 name_hi
    . list

  • #2
    I can replicate your problem on StataNow/MP 19.5 on macOS Tahoe 26.2. The problem is visible even between the Data Editor window and the Results windows -- the characters render correctly in the Data Editor, but not in the Results window. Initially I thought this is because of the font -- my Data Editor was using Helvetica while the Results window was using Monaco. But then I changed the Data Editor font to Monaco, and it still rendered correctly there.

    Screenshots attached.
    Click image for larger version

Name:	Screenshot 2026-01-08 at 2.31.06 AM.png
Views:	1
Size:	216.7 KB
ID:	1784140

    Click image for larger version

Name:	Screenshot 2026-01-08 at 2.31.36 AM.png
Views:	1
Size:	163.8 KB
ID:	1784141

    Comment


    • #3
      Stata's backend Unicode functions are fully capable of handling any languages. This is a display issue in the result window when encountering complex scripts, Hebrew, Arabic, Hindi, etc.

      Try the following undocumented setting:

      set usecharalignment off

      Code:
      copy https://hi.wiktionary.org/wiki/%E0%A4%B5%E0%A4%BF%E0%A4%95%E0%A5%8D%E0%A4%B7%E0%A4%A8%E0%A4%B0%E0%A5%80:%E0%A4%A6%E0%A5%87%E0%A4%B6%E0%A5%8B%E0%A4%82/%E0%A4%B0%E0%A4%BE%E0%A4%B7%E0%A5%8D%E0%A4%9F%E0%A5%8D%E0%A4%B0%E0%A5%8B%E0%A4%82_%E0%A4%95%E0%A5%87_%E0%A4%A8%E0%A4%BE%E0%A4%AE country_code_hi.txt  
      
      filefilter country_code_hi.txt x.txt, from(`"<tr "') to(\n) replace
      import delimited using x.txt, stripquotes(yes) delimiters("\t") encoding(UTF-8) clear
      gen iso2 = ustrtrim(ustrregexs(1)) if ustrregexm(v1,`"<td.+?>([A-X]+)</td><td.+?>.+?</td><td.+?>(.+?)</td>"')
      gen name_hi = ustrtrim(ustrregexs(2)) if ustrregexm(v1,`"<td.+?>([A-X]+)</td><td.+?>.+?</td><td.+?>(.+?)</td>"')
      keep if !mi(iso2,name_hi)
      keep iso2 name_hi  
      
      set usecharalignment off
      list
      Longer explanation, Stata's result window is deeply rooted in the Unix terminal, which emphasizes on the column-based character by character alignment. Hence by default, the result window is displaying one character aligned to a column boundary at a time. The behavior naturally is not up to handle complex script language, which the shape or positioning of a glyph depending on its relation to other glyphs. Examples of complex script languages are Hebrew, Arabic, most of the South Asian languages (Bengali, Hindi, Nepali, etc.).

      Note, set usecharalignment off is a crude attempt to address some issues of displaying complex scripts. It does not really solve any issues, especially displaying the text in a table context. For proper handling complex scripts, the result window needs to be rewritten to handle complex script layout, and very likely its current behavior of displaying will have to change.

      For more information about the complex script layout, see https://en.wikipedia.org/wiki/Comple...ai%20alphabet.
      Last edited by Hua Peng (StataCorp); 07 Jan 2026, 17:45.

      Comment


      • #4
        Thank you Hemanshu Kumar for replication of the problem and Hua Peng for pointing me to -set usecharalignment off- which is good enough for my purpose. Also many thanks for the introduction to "complex script layout". That's interesting.

        Comment

        Working...
        X