Repository navigation
ElementTree parser limitation of input string size #83895
Description
Activity
AnanthVijalapuram commented
on Feb 21, 2020 AnanthVijalapurammannequinMannequinAuthorMore actionsI am trying to parse a very large XML file. Here is the output:
python3.7.4 crif_parser.py Retrieved 3593891712 characters <- this is printed from my script Traceback (most recent call last): File "crif_parser.py", line 9, in <module> tree = ET.fromstring(data) File "python3/3.7.4/lib/python3.7/xml/etree/ElementTree.py", line 1315, in XML parser.feed(text) OverflowError: size does not fit in an int
- added3.7 (EOL)end of lifeend of lifetype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on Feb 21, 2020 I'd suggest feeding the data into the parser in chunks, or letting it read from a file-like object, or something like that.
Also, you probably want to do incremental processing on the data (see the XMLPullParser and iterparse), because reading 3.5GB of XML data into an in-memory tree can easily result in 10x the memory usage. You may have 40GB of RAM on your machine, but even then, I'd still recommend processing the data in incrementally.
- added3.9 (EOL)end of lifeend of life3.10 (EOL)end of lifeend of lifeand removed3.7 (EOL)end of lifeend of life
on Sep 8, 2020 - changed the title
[-]ElementTree limitation[/-][+]ElementTree parser limitation of input string size[/+]on Sep 8, 2020 - changed the title
[-]ElementTree limitation[/-][+]ElementTree parser limitation of input string size[/+]on Sep 8, 2020 This error is hit when reading wiktionary dumps :(
I personally don't think we should do something about this. It's very easy to work around the failure by using the much more efficient
parse()function instead of first reading so much data into a string in memory and then parsing from that. That just uselessly wastes both memory and time. Whether it's a good idea to parse the whole document into memory at all is then yet another question. Most people should be better off usingiterparse()or theXMLPullParserinstead.If nothing else, the failure at least makes users aware that they are not doing something recommended.
Reacted by Vadim KantorovFor modern big machines with dozens gigabytes of memory, it might be not so senseless (and be a good trade-off for its simplicity). If it's an int32 limitation, lifting it to int64 might be good
PR #156746 removes the limitation.
The C implementation raised
OverflowErrorbecause Expat takes the length as anint.xml.parsers.expathas fed Expat in chunks since bpo-17089 (2013), so the pure Python implementation of ElementTree already accepted input larger than 2 GiB; only the accelerator refused it:>>> data = b'<r>' + b'x' * ((1 << 31) + (1 << 27)) + b'</r>' >>> len(ET.fromstring(data).text) 2281701376
@scoder, your advice stays right:
parse()oriterparse()is better than reading gigabytes into a string first. Butxml.parsers.expataccepts such input, so this was not a deliberate limitation, and the accelerated module should not be able to do less than the pure Python one. When I fixedpyexpatin bpo-17089 I could not test this half of it: machines had 2 GiB of RAM. The new test needs 6 GiB and is skipped unless-Mis given.Chunking is not slower for ordinary documents either: 1 GiB with a million elements is parsed in 2.37 s in chunks of 1 MiB and in 2.60 s in a single call.
Change looks good and helpful to me. Thanks for implementing it.
- added a commit that references this issue
on Oct 9, 2026
Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.
Show more details
GitHub fields:
bugs.python.org fields:
Linked PRs