Skip to content

ElementTree parser limitation of input string size #83895

Description

@AnanthVijalapuram
BPO 39714
Nosy @scoder

Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.

Show more details

GitHub fields:

assignee = None
closed_at = None
created_at = <Date 2020-02-21.17:59:58.910>
labels = ['expert-XML', 'type-bug', '3.9', '3.10']
title = 'ElementTree parser limitation of input string size'
updated_at = <Date 2020-09-08.04:55:41.878>
user = 'https://bugs.python.org/AnanthVijalapuram'

bugs.python.org fields:

activity = <Date 2020-09-08.04:55:41.878>
actor = 'scoder'
assignee = 'none'
closed = False
closed_date = None
closer = None
components = ['XML']
creation = <Date 2020-02-21.17:59:58.910>
creator = 'Ananth Vijalapuram'
dependencies = []
files = []
hgrepos = []
issue_num = 39714
keywords = []
message_count = 2.0
messages = ['362418', '376545']
nosy_count = 2.0
nosy_names = ['scoder', 'Ananth Vijalapuram']
pr_nums = []
priority = 'normal'
resolution = None
stage = None
status = 'open'
superseder = None
type = 'behavior'
url = 'https://bugs.python.org/issue39714'
versions = ['Python 3.9', 'Python 3.10']

Linked PRs

Activity

  1. AnanthVijalapuram commented on Feb 21, 2020

    AnanthVijalapurammannequin
    MannequinAuthor

    I am trying to parse a very large XML file. Here is the output:

    python3.7.4 crif_parser.py
    Retrieved 3593891712 characters <- this is printed from my script
    Traceback (most recent call last):
      File "crif_parser.py", line 9, in <module>
        tree = ET.fromstring(data)
      File "python3/3.7.4/lib/python3.7/xml/etree/ElementTree.py", line 1315, in XML
        parser.feed(text)
    OverflowError: size does not fit in an int
  2. scoder commented on Sep 8, 2020

    @scoder
    Contributor

    I'd suggest feeding the data into the parser in chunks, or letting it read from a file-like object, or something like that.

    Also, you probably want to do incremental processing on the data (see the XMLPullParser and iterparse), because reading 3.5GB of XML data into an in-memory tree can easily result in 10x the memory usage. You may have 40GB of RAM on your machine, but even then, I'd still recommend processing the data in incrementally.

  3. added and removed on Sep 8, 2020
  4. changed the title [-]ElementTree limitation[/-] [+]ElementTree parser limitation of input string size[/+] on Sep 8, 2020
  5. changed the title [-]ElementTree limitation[/-] [+]ElementTree parser limitation of input string size[/+] on Sep 8, 2020
  6. transferred this issue fromon Apr 10, 2022
  7. vadimkantorov commented on Dec 14, 2023

    @vadimkantorov

    This error is hit when reading wiktionary dumps :(

  8. scoder commented on Dec 15, 2023

    @scoder
    Contributor

    I personally don't think we should do something about this. It's very easy to work around the failure by using the much more efficient parse() function instead of first reading so much data into a string in memory and then parsing from that. That just uselessly wastes both memory and time. Whether it's a good idea to parse the whole document into memory at all is then yet another question. Most people should be better off using iterparse() or the XMLPullParser instead.

    If nothing else, the failure at least makes users aware that they are not doing something recommended.

  9. vadimkantorov commented on Dec 15, 2023

    @vadimkantorov

    For modern big machines with dozens gigabytes of memory, it might be not so senseless (and be a good trade-off for its simplicity). If it's an int32 limitation, lifting it to int64 might be good

  10. serhiy-storchaka commented on Aug 31, 2026

    @serhiy-storchaka
    Member

    PR #156746 removes the limitation.

    The C implementation raised OverflowError because Expat takes the length as an int. xml.parsers.expat has fed Expat in chunks since bpo-17089 (2013), so the pure Python implementation of ElementTree already accepted input larger than 2 GiB; only the accelerator refused it:

    >>> data = b'<r>' + b'x' * ((1 << 31) + (1 << 27)) + b'</r>'
    >>> len(ET.fromstring(data).text)
    2281701376

    @scoder, your advice stays right: parse() or iterparse() is better than reading gigabytes into a string first. But xml.parsers.expat accepts such input, so this was not a deliberate limitation, and the accelerated module should not be able to do less than the pure Python one. When I fixed pyexpat in bpo-17089 I could not test this half of it: machines had 2 GiB of RAM. The new test needs 6 GiB and is skipped unless -M is given.

    Chunking is not slower for ordinary documents either: 1 GiB with a million elements is parsed in 2.37 s in chunks of 1 MiB and in 2.60 s in a single call.

  11. added 2 commits that reference this issue on Sep 1, 2026
  12. added 2 commits that reference this issue on Sep 1, 2026
  13. scoder commented on Sep 1, 2026

    @scoder
    Contributor

    Change looks good and helpful to me. Thanks for implementing it.

  14. added a commit that references this issue on Sep 12, 2026
  15. added a commit that references this issue on Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions