hadoop

Author	SHA1	Message	Date
Pranav Saxena	4d968add52	HADOOP-19271. NPE in AbfsManagedApacheHttpConnection.toString() when not connected (#7040 ) Contributed by: Pranav Saxena	2024-09-16 12:21:20 -05:00
Steve Loughran	ea6e0f7cd5	HADOOP-19221. S3A: Unable to recover from failure of multipart block upload attempt (#6938 ) This is a major change which handles 400 error responses when uploading large files from memory heap/buffer (or staging committer) and the remote S3 store returns a 500 response from a upload of a block in a multipart upload. The SDK's own streaming code seems unable to fully replay the upload; at attempts to but then blocks and the S3 store returns a 400 response "Your socket connection to the server was not read from or written to within the timeout period. Idle connections will be closed. (Service: S3, Status Code: 400...)" There is an option to control whether or not the S3A client itself attempts to retry on a 50x error other than 503 throttling events (which are independently processed as before) Option: fs.s3a.retry.http.5xx.errors Default: true 500 errors are very rare from standard AWS S3, which has a five nines SLA. It may be more common against S3 Express which has lower guarantees. Third party stores have unknown guarantees, and the exception may indicate a bad server configuration. Consider setting fs.s3a.retry.http.5xx.errors to false when working with such stores. Signification Code changes: There is now a custom set of implementations of software.amazon.awssdk.http.ContentStreamProvidercontent in the class org.apache.hadoop.fs.s3a.impl.UploadContentProviders. These: * Restart on failures * Do not copy buffers/byte buffers into new private byte arrays, so avoid exacerbating memory problems.. There new IOStatistics for specific http error codes -these are collected even when all recovery is performed within the SDK. S3ABlockOutputStream has major changes, including handling of Thread.interrupt() on the main thread, which now triggers and briefly awaits cancellation of any ongoing uploads. If the writing thread is interrupted in close(), it is mapped to an InterruptedIOException. Applications like Hive and Spark must catch these after cancelling a worker thread. Contributed by Steve Loughran	2024-09-13 20:02:14 +01:00
Smith Cruise	c835adb3a8	HADOOP-19201 S3A. Support external-id in assume role (#6876 ) The option fs.s3a.assumed.role.external.id sets the external id for calls of AssumeRole to the STS service Contributed by Smith Cruise	2024-09-10 15:38:32 +01:00
K0K0V0K	c9e9bce361	YARN-11729. Broken 'AM Node Web UI' link on App details page (#7030 )	2024-09-09 16:33:40 +02:00
Saikat Roy	6881d12da4	HADOOP-19262: Upgrade wildfly-openssl:1.1.3.Final to 2.1.4.Final to support Java17+ (#7026 ) Contributed by Saikat Roy	2024-09-09 15:14:03 +01:00
Benjamin Teke	8c41fbcaf5	Revert "YARN-11709. NodeManager should be shut down or blacklisted when it ca…" (#7028 ) Some checks failed website / build (push) Has been cancelled Details This reverts commit `f00094203b`.	2024-09-07 08:48:38 +02:00
PJ Fanning	a00b1c06f3	HADOOP-19269. Upgrade maven-shade-plugin 3.6.0 (#7029 ) Contributed by PJ Fanning	2024-09-05 20:29:44 +01:00
Steve Loughran	57e62ae07f	Revert "YARN-11664. Remove HDFS Binaries/Jars Dependency From Yarn (#6631 )" This reverts commit `6c01490f14`.	2024-09-05 14:35:50 +01:00
Shintaro Onuma	1f302e83fd	HADOOP-18938. S3A: Fix endpoint region parsing for vpc endpoints. (#6466 ) Contributed by Shintaro Onuma	2024-09-05 14:14:04 +01:00
Syed Shameerur Rahman	6c01490f14	YARN-11664. Remove HDFS Binaries/Jars Dependency From Yarn (#6631 ) To support YARN deployments in clusters without HDFS some changes have been made in packaging * new hadoop-common class org.apache.hadoop.fs.HdfsCommonConstants * hdfs class org.apache.hadoop.hdfs.protocol.datatransfer.IOStreamPair moved from hdfs-client to hadoop-common * YARN handlers for DSQuotaExceededException replaced by use of superclass ClusterStorageCapacityExceededException. Contributed by Syed Shameerur Rahman	2024-09-04 13:26:42 +01:00
Cheng Pan	9486844610	HADOOP-16928. Make javadoc work on Java 17 (#6976 ) Contributed by Cheng Pan	2024-09-04 11:50:59 +01:00
Steve Loughran	3bbfb2be08	HADOOP-19257. S3A: ITestAssumeRole.testAssumeRoleBadInnerAuth failure (#7021 ) Remove the error string matched on so that no future message change from AWS will trigger a regression Contributed by Steve Loughran	2024-09-03 21:20:47 +01:00
zhengchenyu	1655acc5e2	HADOOP-19250. [Addendum] Fix test TestServiceInterruptHandling.testRegisterAndRaise. (#7008 ) Contributed by Chenyu Zheng	2024-08-30 12:05:13 +01:00
Steve Loughran	b404c8c8f8	HADOOP-19252. Upgrade hadoop-thirdparty to 1.3.0 (#7007 ) Update the version of hadoop-thirdparty to 1.3.0 across all shaded artifacts used. This synchronizes the shaded protobuf library with those of all other shaded artifacts (guava, avro) Contributed by Steve Loughran	2024-08-30 11:50:51 +01:00
litao	a962aa37e0	HDFS-17599. EC: Fix the mismatch between locations and indices for mover (#6980 )	2024-08-30 12:56:33 +08:00
Ayush Saxena	0837c84a9f	Revert "HADOOP-19231. Add JacksonUtil to manage Jackson classes (#6953 )" This reverts commit `fa9bb0d1ac`.	2024-08-29 14:42:03 +05:30
Cheng Pan	0aab1a2976	HADOOP-19248. Protobuf code generate and replace should happen together (#6975 ) Contributed by Cheng Pan	2024-08-28 20:18:46 +01:00
K0K0V0K	e4ee3d560b	YARN-10345 HsWebServices containerlogs does not honor ACLs for completed jobs (#7013 ) - following rest apis did not have access control - - /ws/v1/history/containerlogs/{containerid}/{filename} - - /ws/v1/history/containers/{containerid}/logs Change-Id: I434f6138966ab22583d356509e40b70d328d9e7c	2024-08-27 17:55:07 +02:00
Sung Dong Kim	89e38f08ae	HDFS-17573. Allow turn on both FSImage parallelization and compression (#6929 ). Contributed by Sung Dong Kim. Signed-off-by: He Xiaoqiao <hexiaoqiao@apache.org>	2024-08-25 17:51:14 +08:00
Kevin Cai	5745a7dd75	HDFS-16084. Fix getJNIEnv crash due to incorrect state set to tls var (#6969 ). Contributed by Kevin Cai. Signed-off-by: He Xiaoqiao <hexiaoqiao@apache.org>	2024-08-25 16:47:05 +08:00
slfan1989	6be04633b5	YARN-11711. Clean Up ServiceScheduler Code. (#6977 ) Contributed by Shilun Fan. Reviewed-by: Steve Loughran <stevel@apache.org> Signed-off-by: Shilun Fan <slfan1989@apache.org>	2024-08-22 17:49:42 +08:00
Heagan A	f6c45e0bcf	HDFS-17546. Follow-up backport from branch3.3 (#6908 ) HDFS-17546. Follow-up backport from branch3.3	2024-08-21 13:11:32 -07:00
Carl Levasseur	68fcd7234c	HADOOP-18542. Keep MSI tenant ID and client ID optional (#4262 ) Contributed by Carl Levasseur	2024-08-21 14:15:28 +01:00
Tsz-Wo Nicholas Sze	012ae9d1aa	HDFS-17606. Do not require implementing CustomizedCallbackHandler. (#7005 )	2024-08-20 16:32:53 -07:00
Anuj Modi	b15ed27cfb	HADOOP-19187: [ABFS][FNSOverBlob] AbfsClient Refactoring to Support Multiple Implementation of Clients. (#6879 ) Refactor AbfsClient into DFS and Blob Client. Contributed by Anuj Modi	2024-08-20 18:07:07 +01:00
dhavalshah9131	33c9ecb652	HADOOP-19249. KMSClientProvider raises NPE with unauthed user (#6984 ) KMSClientProvider raises a NullPointerException when an unauthorised user tries to perform the key operation Contributed by Dhaval Shah	2024-08-20 14:03:05 +01:00
Steve Loughran	2fd7cf53fa	HADOOP-19253. Google GCS compilation fails due to VectorIO changes (#7002 ) Fixes a compilation failure caused by HADOOP-19098 Restore original sortRanges() method signature, FileRange[] sortRanges(List<? extends FileRange>) This ensures that google GCS connector will compile again. It has also been marked as Stable so it is left alone The version returning List<? extends FileRange> has been renamed sortRangeList() Contributed by Steve Loughran	2024-08-19 19:58:15 +01:00
Stephen O'Donnell	df08e0de41	HDFS-17605. Reduce memory overhead of TestBPOfferService (#6996 )	2024-08-19 11:35:11 +01:00
zhengchenyu	e5b76dc99f	HADOOP-19180. EC: Fix calculation errors caused by special index order (#6813 ). Contributed by zhengchenyu. Reviewed-by: He Xiaoqiao <hexiaoqiao@apache.org> Signed-off-by: Shuyan Zhang <zhangshuyan@apache.org>	2024-08-19 12:40:45 +08:00
PJ Fanning	59dba6e1bd	HADOOP-19134. Use StringBuilder instead of StringBuffer. (#6692 ). Contributed by PJ Fanning	2024-08-18 21:29:12 +05:30
slfan1989	b5f88990b7	HADOOP-19136. Upgrade commons-io to 2.16.1. (#6704 ) Contributed by Shilun Fan.	2024-08-16 19:42:26 +01:00
zhengchenyu	bf804cb64b	HADOOP-19250. Fix test TestServiceInterruptHandling.testRegisterAndRaise (#6987 ) Contributed by Chenyu Zheng	2024-08-16 17:16:28 +01:00
Ferenc Erdelyi	f00094203b	YARN-11709. NodeManager should be shut down or blacklisted when it cacannot run program /var/lib/yarn-ce/bin/container-executor (#6960 )	2024-08-16 16:33:10 +02:00
Steve Loughran	5f93edfd70	HADOOP-19153. hadoop-common exports logback as a transitive dependency (#6999 ) - Critical: remove the obsolete exclusion list from hadoop-common. - Diligence: expand the hadoop-project exclusion list to exclude all ch.qos.logback artifacts Contributed by Steve Loughran	2024-08-16 13:54:59 +01:00
PJ Fanning	fa9bb0d1ac	HADOOP-19231. Add JacksonUtil to manage Jackson classes (#6953 ) New class org.apache.hadoop.util.JacksonUtil centralizes construction of Jackson ObjectMappers and JsonFactories. Contributed by PJ Fanning	2024-08-15 16:44:54 +01:00
Steve Loughran	55a576906d	HADOOP-19131. Assist reflection IO with WrappedOperations class (#6686 ) 1. The class WrappedIO has been extended with more filesystem operations - openFile() - PathCapabilities - StreamCapabilities - ByteBufferPositionedReadable All these static methods raise UncheckedIOExceptions rather than checked ones. 2. The adjacent class org.apache.hadoop.io.wrappedio.WrappedStatistics provides similar access to IOStatistics/IOStatisticsContext classes and operations. Allows callers to: * Get a serializable IOStatisticsSnapshot from an IOStatisticsSource or IOStatistics instance * Save an IOStatisticsSnapshot to file * Convert an IOStatisticsSnapshot to JSON * Given an object which may be an IOStatisticsSource, return an object whose toString() value is a dynamically generated, human readable summary. This is for logging. * Separate getters to the different sections of IOStatistics. * Mean values are returned as a Map.Pair<Long, Long> of (samples, sum) from which means may be calculated. There are examples of the dynamic bindings to these classes in: org.apache.hadoop.io.wrappedio.impl.DynamicWrappedIO org.apache.hadoop.io.wrappedio.impl.DynamicWrappedStatistics These use DynMethods and other classes in the package org.apache.hadoop.util.dynamic which are based on the Apache Parquet equivalents. This makes re-implementing these in that library and others which their own fork of the classes (example: Apache Iceberg) 3. The openFile() option "fs.option.openfile.read.policy" has added specific file format policies for the core filetypes * avro * columnar * csv * hbase * json * orc * parquet S3A chooses the appropriate sequential/random policy as a A policy `parquet, columnar, vector, random, adaptive` will use the parquet policy for any filesystem aware of it, falling back to the first entry in the list which the specific version of the filesystem recognizes 4. New Path capability fs.capability.virtual.block.locations Indicates that locations are generated client side and don't refer to real hosts. Contributed by Steve Loughran	2024-08-14 14:43:00 +01:00
Viraj Jasani	fa83c9a805	HADOOP-19072 S3A: Override fs.s3a.performance.flags for tests (ADDENDUM 2) (#6993 ) Second followup to #6543; all hadoop-aws integration tests complete correctly even when fs.s3a.performance.flags = * Contributed by Viraj Jasani	2024-08-14 10:57:44 +01:00
Viraj Jasani	74ff00705c	HADOOP-19072. S3A: Override fs.s3a.performance.flags for tests (ADDENDUM) (#6985 ) This is a followup to #6543 which ensures all test pass in configurations where fs.s3a.performance.flags is set to "*" or contains "mkdirs" Contributed by VJ Jasani	2024-08-12 14:16:44 +01:00
Viraj Jasani	321a6cc55e	HADOOP-19072. S3A: expand optimisations on stores with "fs.s3a.performance.flags" for mkdir (#6543 ) If the flag list in fs.s3a.performance.flags includes "mkdir" then the safety check of a walk up the tree to look for a parent directory, -done to verify a directory isn't being created under a file- are skipped. This saves the cost of multiple list operations. Contributed by Viraj Jasani	2024-08-08 17:48:51 +01:00
Masatake Iwasaki	2a50911734	HADOOP-17609. Make SM4 support optional for OpenSSL native code. (#3019 ) Reviewed-by: Steve Loughran <stevel@apache.org> Reviewed-by: Wei-Chiu Chuang <weichiu@apache.org>	2024-08-08 21:03:05 +09:00
Tsz-Wo Nicholas Sze	b189ef8197	HDFS-17575. SaslDataTransferClient should use SaslParticipant to create messages. (#6954 )	2024-08-05 10:42:12 -07:00
Cheng Pan	59d5e0bb2e	HADOOP-19244. Pullout arch-agnostic maven javadoc plugin configurations in hadoop-common (#6970 ) Contributed by Cheng Pan. Reviewed-by: Steve Loughran <stevel@apache.org> Signed-off-by: Shilun Fan <slfan1989@apache.org>	2024-08-05 15:30:36 +08:00
zhengchenyu	b08d492abd	HADOOP-19246. Update the yasm rpm download address (#6973 ) Reviewed-by: Shilun Fan <slfan1989@apache.org> Signed-off-by: Tao Li <tomscut@apache.org>	2024-08-05 09:57:16 +08:00
Steve Loughran	2cf4d638af	HADOOP-19245. S3ABlockOutputStream no longer sends progress events in close() (#6974 ) Contributed by Steve Loughran	2024-08-02 16:01:03 +01:00
PJ Fanning	c593c17255	HADOOP-19237. Upgrade to dnsjava 3.6.1 due to CVEs (#6961 ) Contributed by P J Fanning	2024-08-01 20:07:36 +01:00
Takanobu Asanuma	059e996c02	HDFS-17591. RBF: Router should follow X-FRAME-OPTIONS protection setting (#6963 )	2024-07-30 10:14:33 +09:00
Mukund Thakur	038636a1b5	HADOOP-19238. Fix create-release script for arm64 based MacOS (#6962 ) Contributed by Mukund Thakur	2024-07-29 19:45:14 +01:00
Steve Loughran	a5806a9e7b	HADOOP-19161. S3A: option "fs.s3a.performance.flags" to take list of performance flags (#6789 ) 1. Configuration adds new method `getEnumSet()` to get a set of enum values from a configuration string. <E extends Enum<E>> EnumSet<E> getEnumSet(String key, Class<E> enumClass, boolean ignoreUnknown) Whitespace is ignored, case is ignored and the value "" is mapped to "all values of the enum". If "ignoreUnknown" is true then when parsing, unknown values are ignored. This is recommended for forward compatiblity with later versions. 2. This support is implemented in org.apache.hadoop.fs.s3a.impl.ConfigurationHelper -it can be used elsewhere in the hadoop codebase. 3. A new private FlagSet class in hadoop common manages a set of enum flags. It implements StreamCapabilities and can be probed for a specific option being set (with a prefix) S3A adds an option fs.s3a.performance.flags which builds a FlagSet with enum type PerformanceFlagEnum which initially contains {Create, Delete, Mkdir, Open} * the existing fs.s3a.create.performance option sets the flag "Create". * tests which configure fs.s3a.create.performance MUST clear fs.s3a.performance.flags in test setup. Future performance flags are planned, with different levels of safety and/or backwards compatibility. Contributed by Steve Loughran	2024-07-29 11:33:51 +01:00
Raphael Azzolini	4525c7e35e	HADOOP-19197. S3A: Support AWS KMS Encryption Context (#6874 ) The new property fs.s3a.encryption.context allow users to specify the AWS KMS Encryption Context to be used in S3A. The value of the encryption context is a key/value string that will be Base64 encoded and set in the parameter ssekmsEncryptionContext from the S3 client. Contributed by Raphael Azzolini	2024-07-23 17:09:04 +01:00
Aswin M Prabhu	e2a0dca43b	HDFS-16690. Automatically format unformatted JNs with JournalNodeSyncer (#6925 ). Contributed by Aswin M Prabhu. Signed-off-by: He Xiaoqiao <hexiaoqiao@apache.org>	2024-07-23 20:55:57 +08:00

1 2 3 4 5 ...

27410 Commits