Skip to content

(🐞) Column marker in error message off for certain string errors #102312

Description

@KotlinIsland
"\N{a}"
👉 python test.py
  File "/test.py", line 1
    "\N{a}"
           ^
SyntaxError: (unicode error) 'unicodeescape' codec can't decode bytes in position 0-4: unknown Unicode character name

Here the caret is pointing to blackspace after the code

Expected

👉 python test.py
  File "/test.py", line 1
    "\N{a}"
        ^
SyntaxError: (unicode error) 'unicodeescape' codec can't decode bytes in position 0-4: unknown Unicode character name

Activity

  1. arhadthedev commented on Feb 28, 2023

    @arhadthedev
    Member

    Reproduced in the CLI too:

    >>> "\N{a}"
      File "<stdin>", line 1
        "\N{a}"
               ^
    SyntaxError: (unicode error) 'unicodeescape' codec can't decode bytes in position 0-4: unknown Unicode character name

    By the way, the interpreter hangs for around five seconds on first execution of \N{a} with no delay on further attempts.

  2. KotlinIsland commented on Feb 28, 2023

    @KotlinIsland
    ContributorAuthor

    @arhadthedev Am I missing something? I Can't see any caret in your repro.

  3. arhadthedev commented on Feb 28, 2023

    @arhadthedev
    Member

    @KotlinIsland Thank you, updated my initially undercopied output.

  4. arhadthedev commented on Feb 28, 2023

    @arhadthedev
    Member

    @benjaminp, @pablogsal (as parser experts)

    update: @gvanrossum, @lysnikolaou (as new parser experts)

  5. terryjreedy commented on Feb 28, 2023

    @terryjreedy
    Member

    This is likely a duplicate of #102310.
    Testing with Python 3.12.0a5+ (heads/main:0f89acf6cc, Feb 27 2023, 21:33:28), message is "incomplete input" and closing quote is marked. I interpret this as having been fixed already.
    EDIT: the different message is because I initially ran in IDLE, which uses codeop, which uses compile slightly differently.
    Get same messages in REPL compiled 2 hours ago and in installed 3.12.0a5 and installed 3.10.10.

  6. gvanrossum commented on Feb 28, 2023

    @gvanrossum
    Member

    By the way, the interpreter hangs for around five seconds on first execution of \N{a} with no delay on further attempts.

    Could that just be importing the (humongous) unicodedata module to resolve the name?

    @terryjreedy What is it a duplicate of? Your link points to this same issue.

  7. gvanrossum commented on Feb 28, 2023

    @gvanrossum
    Member

    PS. If Benjamin is still listed as parser expert that's likely out of date.

  8. terryjreedy commented on Feb 28, 2023

    @terryjreedy
    Member

    #102310, about same problem with b"Ā". Sorry about bad link (now fixed).

  9. KotlinIsland commented on Feb 28, 2023

    @KotlinIsland
    ContributorAuthor

    I don't think it's the same issue as #102310, in that issue the column marker is offset by a non ascii character, in this issue the marker appears after the string. While this issue does make an appearance in #102310, I do feel that they are distinct.

    👉 $c:temp = "'\N{a}'"
    👉 py temp
      File "C:\temp", line 1
        '\N{a}'
               ^
    👉 $c:temp = "b'Ā'"
    👉 py temp
      File "C:\temp", line 1
        b'Ā'
             ^
    👉 $c:temp = "b'ĀĀĀĀĀĀĀĀ'"
    👉 py temp
      File "C:\temp", line 1
        b'ĀĀĀĀĀĀĀĀ'
                           ^
    SyntaxError: bytes can only contain ASCII literal characters
    SyntaxError: (unicode error) 'unicodeescape' codec can't decode bytes in position 0-4: unknown Unicode character name
    
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    interpreter-core(Objects, Python, Grammar, and Parser dirs)type-bugAn unexpected behavior, bug, or error

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions