hadoop

Author	SHA1	Message	Date
Steve Loughran	81d90fd65b	HADOOP-18073. S3A: Upgrade AWS SDK to V2 (#5995 ) This patch migrates the S3A connector to use the V2 AWS SDK. This is a significant change at the source code level. Any applications using the internal extension/override points in the filesystem connector are likely to break. This includes but is not limited to: - Code invoking methods on the S3AFileSystem class which used classes from the V1 SDK. - The ability to define the factory for the `AmazonS3` client, and to retrieve it from the S3AFileSystem. There is a new factory API and a special interface S3AInternals to access a limited set of internal classes and operations. - Delegation token and auditing extensions. - Classes trying to integrate with the AWS SDK. All standard V1 credential providers listed in the option fs.s3a.aws.credentials.provider will be automatically remapped to their V2 equivalent. Other V1 Credential Providers are supported, but only if the V1 SDK is added back to the classpath. The SDK Signing plugin has changed; all v1 signers are incompatible. There is no support for the S3 "v2" signing algorithm. Finally, the aws-sdk-bundle JAR has been replaced by the shaded V2 equivalent, "bundle.jar", which is now exported by the hadoop-aws module. Consult the document aws_sdk_upgrade for the full details. Contributed by Ahmar Suhail + some bits by Steve Loughran	2023-09-11 14:30:25 +01:00
Viraj Jasani	0e3aafe6c0	HADOOP-18399. S3A Prefetch - SingleFilePerBlockCache to use LocalDirAllocator (#5054 ) Contributed by Viraj Jasani	2023-04-18 16:37:48 +01:00
Steve Loughran	7c3d94a032	HADOOP-18637. S3A to support upload of files greater than 2 GB using DiskBlocks (#5543 ) Contributed By: HarshitGupta and Steve Loughran	2023-04-12 05:17:45 +05:30
Viraj Jasani	5b1657278c	HADOOP-18377. hadoop-aws build to add a -prefetch profile to run all tests with prefetching (#4914 ) Contributed by Viraj Jasani	2022-09-20 10:26:13 +01:00
Steve Loughran	e199da3fae	HADOOP-17833. Improve Magic Committer performance (#3289 ) Speed up the magic committer with key changes being * Writes under __magic always retain directory markers * File creation under __magic skips all overwrite checks, including the LIST call intended to stop files being created over dirs. * mkdirs under __magic probes the path for existence but does not look any further. Extra parallelism in task and job commit directory scanning Use of createFile and openFile with parameters which all for HEAD checks to be skipped. The committer can write the summary _SUCCESS file to the path `fs.s3a.committer.summary.report.directory`, which can be in a different file system/bucket if desired, using the job id as the filename. Also: HADOOP-15460. S3A FS to add `fs.s3a.create.performance` Application code can set the createFile() option fs.s3a.create.performance to true to disable the same safety checks when writing under magic directories. Use with care. The createFile option prefix `fs.s3a.create.header.` can be used to add custom headers to S3 objects when created. Contributed by Steve Loughran.	2022-06-17 19:11:35 +01:00
Viraj Jasani	66b72406bd	HADOOP-18131. Upgrade maven enforcer plugin and relevant dependencies (#4000 ) Reviewed-by: Akira Ajisaka <aajisaka@apache.org> Reviewed-by: Wei-Chiu Chuang <weichiu@apache.org> Signed-off-by: Takanobu Asanuma <tasanuma@apache.org>	2022-03-08 17:27:04 +09:00
Steve Loughran	14ba19af06	HADOOP-17409. Remove s3guard from S3A module (#3534 ) Completely removes S3Guard support from the S3A codebase. If the connector is configured to use any metastore other than the null and local stores (i.e. DynamoDB is selected) the s3a client will raise an exception and refuse to initialize. This is to ensure that there is no mix of S3Guard enabled and disabled deployments with the same configuration but different hadoop releases -it must be turned off completely. The "hadoop s3guard" command has been retained -but the supported subcommands have been reduced to those which are not purely S3Guard related: "bucket-info" and "uploads". This is major change in terms of the number of files changed; before cherry picking subsequent s3a patches into older releases, this patch will probably need backporting first. Goodbye S3Guard, your work is done. Time to die. Contributed by Steve Loughran.	2022-01-17 18:08:57 +00:00
Viraj Jasani	516f36c6f1	HADOOP-17967. Keep restrict-imports-enforcer-rule for Guava VisibleForTesting in hadoop-main pom (#3555 )	2021-10-21 16:54:25 +09:00
Viraj Jasani	79e5a7f3e3	HADOOP-17962. Replace Guava VisibleForTesting by Hadoop's own annotation in hadoop-tools modules (#3540 )	2021-10-14 17:43:32 +09:00
Viraj Jasani	4ef27a596f	HADOOP-17753. Keep restrict-imports-enforcer-rule for Guava Lists in top level hadoop-main pom (#3087 )	2021-06-11 12:15:52 +09:00
Viraj Jasani	f4b24c68e7	HADOOP-17743. Replace Guava Lists usage by Hadoop's own Lists in hadoop-common, hadoop-tools and cloud-storage projects (#3072 )	2021-06-07 13:24:09 +09:00
Viraj Jasani	986d0a4f1d	HADOOP-17732. Keep restrict-imports-enforcer-rule for Guava Sets in hadoop-main pom (#3049 ) Signed-off-by: Takanobu Asanuma <tasanuma@apache.org>	2021-05-26 17:14:31 +09:00
Viraj Jasani	e4062ad027	HADOOP-17115. Replace Guava Sets usage by Hadoop's own Sets in hadoop-common and hadoop-tools (#2985 ) Signed-off-by: Sean Busbey <busbey@apache.org>	2021-05-20 10:47:04 -05:00
Akira Ajisaka	23b343aed1	HADOOP-16870. Use spotbugs-maven-plugin instead of findbugs-maven-plugin (#2753 ) Removed findbugs from the hadoop build images and added spotbugs instead. Upgraded SpotBugs to 4.2.2 and spotbugs-maven-plugin to 4.2.0. Reviewed-by: Masatake Iwasaki <iwasakims@apache.org>	2021-03-11 10:56:07 +09:00
Akira Ajisaka	9a298d180d	Revert "HADOOP-16870. Use spotbugs-maven-plugin instead of findbugs-maven-plugin (#2454 )" This reverts commit `4cf3531583`.	2021-02-19 11:09:10 +09:00
Akira Ajisaka	4cf3531583	HADOOP-16870. Use spotbugs-maven-plugin instead of findbugs-maven-plugin (#2454 ) Use spotbugs instead of findbugs. Removed findbugs from the hadoop build images, and added spotbugs in the images instead. Reviewed-by: Masatake Iwasaki <iwasakims@apache.org> Reviewed-by: Inigo Goiri <inigoiri@apache.org> Reviewed-by: Dinesh Chitlangia <dineshc@apache.org>	2021-02-17 10:38:20 +09:00
Steve Loughran	617af28e80	HADOOP-17271. S3A connector to support IOStatistics. (#2580 ) S3A connector to support the IOStatistics API of HADOOP-16830, This is a major rework of the S3A Statistics collection to * Embrace the IOStatistics APIs * Move from direct references of S3AInstrumention statistics collectors to interface/implementation classes in new packages. * Ubiquitous support of IOStatistics, including: S3AFileSystem, input and output streams, RemoteIterator instances provided in list calls. * Adoption of new statistic names from hadoop-common Regarding statistic collection, as well as all existing statistics, the connector now records min/max/mean durations of HTTP GET and HEAD requests, and those of LIST operations. Contributed by Steve Loughran.	2020-12-31 21:55:39 +00:00
Steve Loughran	5092ea62ec	HADOOP-13230. S3A to optionally retain directory markers. This adds an option to disable "empty directory" marker deletion, so avoid throttling and other scale problems. This feature is not backwards compatible. Consult the documentation and use with care. Contributed by Steve Loughran. Change-Id: I69a61e7584dc36e485d5e39ff25b1e3e559a1958	2020-08-15 12:51:08 +01:00
Akira Ajisaka	c40cbc57fa	HADOOP-17091. [JDK11] Fix Javadoc errors (#2098 )	2020-08-03 10:46:51 +09:00
Brahma Reddy Battula	8914cf9167	Preparing for 3.4.0 development	2020-03-29 23:24:25 +05:30
Sahil Takiar	f206b736f0	HADOOP-16346. Stabilize S3A OpenSSL support. Introduces `openssl` as an option for `fs.s3a.ssl.channel.mode`. The new option is documented and marked as experimental. For details on how to use this, consult the peformance document in the s3a documentation. This patch is the successor to HADOOP-16050 "S3A SSL connections should use OpenSSL" -which was reverted because of incompatibilities between the wildfly OpenSSL client and the AWS HTTPS servers (HADOOP-16347). With the Wildfly release moved up to 1.0.7.Final (HADOOP-16405) everything should now work. Related issues: * HADOOP-15669. ABFS: Improve HTTPS Performance * HADOOP-16050: S3A SSL connections should use OpenSSL * HADOOP-16371: Option to disable GCM for SSL connections when running on Java 8 * HADOOP-16405: Upgrade Wildfly Openssl version to 1.0.7.Final Contributed by Sahil Takiar Change-Id: I80a4bc5051519f186b7383b2c1cea140be42444e	2020-01-21 16:37:51 +00:00
Steve Loughran	f44abc3e11	HADOOP-16207 Improved S3A MR tests. Contributed by Steve Loughran. Replaces the committer-specific terasort and MR test jobs with parameterization of the (now single tests) and use of file:// over hdfs:// as the cluster FS. The parameterization ensures that only one of the specific committer tests run at a time -overloads of the test machines are less likely, and so the suites can be pulled back into the parallel phase. There's also more detailed validation of the stage outputs of the terasorting; if one test fails the rest are all skipped. This and the fact that job output is stored under target/yarn-${timestamp} means failures should be more debuggable. Change-Id: Iefa370ba73c6419496e6e69dd6673d00f37ff095	2019-10-04 14:12:31 +01:00
Steve Loughran	b15ef7dc3d	HADOOP-16384: S3A: Avoid inconsistencies between DDB and S3. Contributed by Steve Loughran Contains - HADOOP-16397. Hadoop S3Guard Prune command to support a -tombstone option. - HADOOP-16406. ITestDynamoDBMetadataStore.testProvisionTable times out intermittently This patch doesn't fix the underlying problem but it * changes some tests to clean up better * does a lot more in logging operations in against DDB, if enabled * adds an entry point to dump the state of the metastore and s3 tables (precursor to fsck) * adds a purge entry point to help clean up after a test run has got a store into a mess * s3guard prune command adds -tombstone option to only clear tombstones The outcome is that tests should pass consistently and if problems occur we have better diagnostics. Change-Id: I3eca3f5529d7f6fec398c0ff0472919f08f054eb	2019-07-12 13:02:25 +01:00
Steve Loughran	e02eb24e0a	HADOOP-15183. S3Guard store becomes inconsistent after partial failure of rename. Contributed by Steve Loughran. Change-Id: I825b0bc36be960475d2d259b1cdab45ae1bb78eb	2019-06-20 09:56:40 +01:00
Steve Loughran	309501c6fa	Revert "HADOOP-16050: s3a SSL connections should use OpenSSL" This reverts commit `b067f8acaa`. Change-Id: I584b050a56c0e6f70b11fa3f7db00d5ac46e7dd8	2019-06-05 13:54:55 +01:00
Akira Ajisaka	afd844059c	HADOOP-16331. Fix ASF License check in pom.xml Signed-off-by: Takanobu Asanuma <tasanuma@apache.org>	2019-05-29 17:25:13 +09:00
Steve Loughran	0c73dba3a6	HADOOP-16332. Remove S3A dependency on http core. Contributed by Steve Loughran. Change-Id: I53209c993a405fefdb5e1b692d5a56d027d3b845	2019-05-28 22:50:37 +01:00
Akira Ajisaka	9f933e6446	HADOOP-16323. https everywhere in Maven settings.	2019-05-27 15:24:59 +09:00
Ben Roling	a36274d699	HADOOP-16085. S3Guard: use object version or etags to protect against inconsistent read after replace/overwrite. Contributed by Ben Roling. S3Guard will now track the etag of uploaded files and, if an S3 bucket is versioned, the object version. You can then control how to react to a mismatch between the data in the DynamoDB table and that in the store: warn, fail, or, when using versions, return the original value. This adds two new columns to the table: etag and version. This is transparent to older S3A clients -but when such clients add/update data to the S3Guard table, they will not add these values. As a result, the etag/version checks will not work with files uploaded by older clients. For a consistent experience, upgrade all clients to use the latest hadoop version.	2019-05-19 22:29:54 +01:00
Sahil Takiar	b067f8acaa	HADOOP-16050: s3a SSL connections should use OpenSSL (cherry picked from commit aebf229c175dfa19fff3b31e9e67596f6c6124fa)	2019-05-16 08:57:54 -06:00
Steve Loughran	9f1c017f44	HADOOP-16058. S3A tests to include Terasort. Contributed by Steve Loughran. This includes - HADOOP-15890. Some S3A committer tests don't match ITest* pattern; don't run in maven - MAPREDUCE-7090. BigMapOutput example doesn't work with paths off cluster fs - MAPREDUCE-7091. Terasort on S3A to switch to new committers - MAPREDUCE-7092. MR examples to work better against cloud stores	2019-03-21 11:15:37 +00:00
Akira Ajisaka	1129288cf5	HADOOP-14178. Move Mockito up to version 2.23.4. Contributed by Akira Ajisaka and Masatake Iwasaki.	2019-01-29 18:29:56 -08:00
Steve Loughran	6d0bffe17e	HADOOP-14556. S3A to support Delegation Tokens. Contributed by Steve Loughran and Daryn Sharp.	2019-01-14 17:59:27 +00:00
Akira Ajisaka	7f78397036	Revert "HADOOP-14556. S3A to support Delegation Tokens." This reverts commit `d7152332b3`.	2019-01-08 14:51:30 +09:00
Steve Loughran	d7152332b3	HADOOP-14556. S3A to support Delegation Tokens. Contributed by Steve Loughran.	2019-01-07 13:18:03 +00:00
Steve Loughran	a668f8e6c6	HADOOP-16015. Add bouncycastle jars to hadoop-aws as test dependencies. Contributed by Steve Loughran.	2018-12-20 18:09:01 +00:00
Sunil G	58fa96b697	Changed version in trunk to 3.3.0-SNAPSHOT.	2018-10-02 22:41:41 +05:30
Steve Loughran	d7c0a08a1c	HADOOP-15426 Make S3guard client resilient to DDB throttle events and network failures (Contributed by Steve Loughran)	2018-09-12 21:04:49 -07:00
Sean Mackrory	b089a06793	HADOOP-14918. Remove the Local Dynamo DB test option. Contributed by Gabor Bota.	2018-06-20 16:45:08 -06:00
Chris Douglas	45d1b0fdcc	HADOOP-14696. parallel tests don't work for Windows. Contributed by Allen Wittenauer	2018-03-12 20:05:39 -07:00
Wangda Tan	60f9e60b3b	Preparing for 3.2.0 development Change-Id: I6d0e01f3d665d26573ef2b957add1cf0cddf7938	2018-02-11 11:17:38 +08:00
Steve Loughran	de8b6ca5ef	HADOOP-13786 Add S3A committer for zero-rename commits to S3 endpoints. Contributed by Steve Loughran and Ryan Blue.	2017-11-22 15:28:12 +00:00
Akira Ajisaka	6903cf096e	HADOOP-13514. Upgrade maven surefire plugin to 2.20.1 Signed-off-by: Allen Wittenauer <aw@apache.org>	2017-11-19 12:39:37 -08:00
Aaron Fabbri	49467165a5	HADOOP-14738 Remove S3N and obsolete bits of S3A; rework docs. Contributed by Steve Loughran.	2017-09-14 14:10:48 -07:00
Andrew Wang	0d419c984f	Preparing for 3.1.0 development	2017-09-01 11:53:48 -07:00
Steve Loughran	621b43e254	HADOOP-13345 HS3Guard: Improved Consistency for S3A. Contributed by: Chris Nauroth, Aaron Fabbri, Mingliang Liu, Lei (Eddy) Xu, Sean Mackrory, Steve Loughran and others.	2017-09-01 14:13:41 +01:00
Steve Loughran	7fc324aabd	HADOOP-14126. Remove jackson, joda and other transient aws SDK dependencies from hadoop-aws. Contributed by Steve Loughran (cherry picked from commit ced547d5f0dbea571cbc472c5f55fe89d5900a6f)	2017-08-04 11:09:08 +01:00
Andrew Wang	af2773f609	Updating version for 3.0.0-beta1 development	2017-06-29 17:57:40 -07:00
Andrew Wang	16ad896d5c	Update maven version for 3.0.0-alpha4 development	2017-05-26 14:09:44 -07:00
Akira Ajisaka	0d5c8ed8e0	HADOOP-14401. maven-project-info-reports-plugin can be removed. Contributed by Andras Bokor.	2017-05-11 16:37:32 -05:00

1 2

79 Commits