[prev in list] [next in list] [prev in thread] [next in thread]
List: hadoop-dev
Subject: [jira] [Resolved] (HADOOP-17414) Magic committer files don't have the count of bytes written collect
From: "Steve Loughran (Jira)" <jira () apache ! org>
Date: 2021-01-26 19:48:01
Message-ID: JIRA.13344592.1607351315000.517581.1611690481224 () Atlassian ! JIRA
[Download RAW message or body]
[ https://issues.apache.org/jira/browse/HADOOP-17414?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel \
]
Steve Loughran resolved HADOOP-17414.
-------------------------------------
Fix Version/s: 3.3.1
Resolution: Fixed
> Magic committer files don't have the count of bytes written collected by spark
> ------------------------------------------------------------------------------
>
> Key: HADOOP-17414
> URL: https://issues.apache.org/jira/browse/HADOOP-17414
> Project: Hadoop Common
> Issue Type: Sub-task
> Components: fs/s3
> Affects Versions: 3.2.0
> Reporter: Steve Loughran
> Assignee: Steve Loughran
> Priority: Major
> Labels: pull-request-available
> Fix For: 3.3.1
>
> Time Spent: 12h 20m
> Remaining Estimate: 0h
>
> The spark statistics tracking doesn't correctly assess the size of the uploaded \
> files as it only calls getFileStatus on the zero byte objects -not the \
> yet-to-manifest files. Which, given they don't exist yet, isn't \
> easy to do.
> Solution:
> * Add getXAttr and listXAttr API calls to S3AFileSystem
> * Return all S3 object headers as XAttr attributes prefixed "header." That's custom \
> and standard (e.g header.Content-Length). The setXAttr call isn't implemented, so \
> for correctness the FS doesn't declare its support for the API in \
> hasPathCapability(). The magic commit file write sets the custom header
> set the length of the data final data in the header
> x-hadoop-s3a-magic-data-length in the marker file.
> A matching patch in Spark will look for the XAttr
> "header.x-hadoop-s3a-magic-data-length" when the file
> being probed for output data is zero byte long.
> As a result, the job tracking statistics will report the
> bytes written but yet to be manifest.
--
This message was sent by Atlassian Jira
(v8.3.4#803005)
---------------------------------------------------------------------
To unsubscribe, e-mail: common-dev-unsubscribe@hadoop.apache.org
For additional commands, e-mail: common-dev-help@hadoop.apache.org
[prev in list] [next in list] [prev in thread] [next in thread]
Configure |
About |
News |
Add a list |
Sponsored by KoreLogic