Feature or enhancement
Proposal:
The conversion between Python values and CSV fields is hard-coded in both directions. The reader converts unquoted fields with float(), and only in the QUOTE_NONNUMERIC and QUOTE_STRINGS modes. The writer converts every non-string value with str().
This is the common cause of several open issues:
I propose two parameters, mirroring parse_float in json:
csv.reader(f, converter=None) -- called as converter(index, field) instead of float().
csv.writer(f, formatter=None) -- called as formatter(index, value) instead of str(). It must return a string.
index is the 0-based position of the field in the record. Both default to None, which keeps the current behavior. The hooks only replace the existing calls -- what is not passed to float() or str() now is not passed to them either. Quoting is still decided by the original value.
The index goes first, like in enumerate(). This also makes a wrong one-argument callable fail at once: converter=int raises TypeError on the first field instead of taking the index as the base.
The index makes the hooks per-column, which is what the dtype and converters parameters of pandas.read_csv() are used for:
>>> types = [str, int, Decimal, Fraction]
>>> list(csv.reader(['spam,42,1.10,1/2'], quoting=csv.QUOTE_NONNUMERIC,
... converter=lambda i, field: types[i](field)))
[['spam', 42, Decimal('1.10'), Fraction(1, 2)]]
>>> def money(index, value):
... return format(value, '.2f') if index == 2 else str(value)
>>> csv.writer(sys.stdout, formatter=money).writerow(['a', 1, 0.0, 3.14159])
a,1,0.00,3.14159
gh-85002 no longer needs a parameter of its own -- a strict writer is a formatter which refuses everything except numbers.
I have a working prototype (about 90 lines in Modules/_csv.c).
Open question: should these be parameters of the reader and the writer, or attributes of the dialect? A dialect is a portable description of the file syntax -- it is registered under a global name, sniffed, and copied -- so keeping callables out of it seems better.
Linked PRs
Feature or enhancement
Proposal:
The conversion between Python values and CSV fields is hard-coded in both directions. The reader converts unquoted fields with
float(), and only in theQUOTE_NONNUMERICandQUOTE_STRINGSmodes. The writer converts every non-string value withstr().This is the common cause of several open issues:
boolis written unquoted asTrue, which cannot be read back.csvdoes not round-trip forcomplexnumbers #98485 -- the same forcomplex;FractionandIntEnumare affected too, andDecimalsilently round-trips throughfloat.QUOTE_NONNUMERICmode.I propose two parameters, mirroring
parse_floatinjson:csv.reader(f, converter=None)-- called asconverter(index, field)instead offloat().csv.writer(f, formatter=None)-- called asformatter(index, value)instead ofstr(). It must return a string.index is the 0-based position of the field in the record. Both default to
None, which keeps the current behavior. The hooks only replace the existing calls -- what is not passed tofloat()orstr()now is not passed to them either. Quoting is still decided by the original value.The index goes first, like in
enumerate(). This also makes a wrong one-argument callable fail at once:converter=intraisesTypeErroron the first field instead of taking the index as the base.The index makes the hooks per-column, which is what the
dtypeandconvertersparameters ofpandas.read_csv()are used for:gh-85002 no longer needs a parameter of its own -- a strict writer is a formatter which refuses everything except numbers.
I have a working prototype (about 90 lines in
Modules/_csv.c).Open question: should these be parameters of the reader and the writer, or attributes of the dialect? A dialect is a portable description of the file syntax -- it is registered under a global name, sniffed, and copied -- so keeping callables out of it seems better.
Linked PRs