hadoop

Author	SHA1	Message	Date
Jia Fan	4f0f5a546c	HADOOP-19049. Fix StatisticsDataReferenceCleaner classloader leak (#6488 ) Contributed by Jia Fan	2024-02-03 14:48:52 +00:00
Viraj Jasani	7504b8505f	HADOOP-18980. S3A credential provider remapping: make extensible (#6406 ) Contributed by Viraj Jasani	2024-02-02 17:02:48 +00:00
Tsz-Wo Nicholas Sze	da34ecdb83	HADOOP-19035. CrcUtil/CrcComposer should not throw IOException for non-IO. (#6443 )	2024-01-25 10:35:32 -08:00
PJ Fanning	76691dfa14	HADOOP-18894: upgrade sshd-core due to CVEs (#6060 ) Contributed by PJ Fanning. Reviewed-by: He Xiaoqiao <hexiaoqiao@apache.org> Reviewed-by: Steve Loughran <stevel@cloudera.com> Signed-off-by: Shilun Fan <slfan1989@apache.org>	2024-01-21 08:13:25 +08:00
Xing Lin	453e264eb4	HADOOP-18981. Move oncrpc and portmap packages to hadoop-common (#6280 ) Move the org.apache.hadoop.{oncrpc, portmap} packages from the hadoop-nfs module to the hadoop-common module. This allows for use of the protocol beyond just NFS -including within HDFS itself. Contributed by Xing Lin	2024-01-11 14:06:15 +00:00
Lei Yang	661c784662	HDFS-17290: Adds disconnected client rpc backoff metrics (#6359 )	2024-01-04 20:24:10 -08:00
Anika Kelhanka	62cc673d00	[HADOOP-19010] - NullPointerException in Hadoop Credential Check CLI (#6351 )	2023-12-16 12:23:52 +05:30
Steve Loughran	e221231e81	HADOOP-18996. S3A to provide full support for S3 Express One Zone (#6308 ) This adds borad support for Amazon S3 Express One Zone to the S3A connector, particularly resilience of other parts of the codebase to LIST operations returning paths under which only in-progress uploads are taking place. hadoop-common and hadoop-mapreduce treewalking routines all cope with this; distcp is left alone. There are still some outstanding followup issues, and we expect more to surface with extended use. Contains HADOOP-18955. AWS SDK v2: add path capability probe "fs.s3a.capability.aws.v2 * lets us probe for AWS SDK version * bucket-info reports it Contains HADOOP-18961 S3A: add s3guard command "bucket" hadoop s3guard bucket -create -region us-west-2 -zone usw2-az2 \ s3a://stevel--usw2-az2--x-s3/ * requires -zone if bucket is zonal * rejects it if not * rejects zonal bucket suffixes if endpoint is not aws (safety feature) * imperfect, but a functional starting point. New path capability "fs.s3a.capability.zonal.storage" * Used in tests to determine whether pending uploads manifest paths * cli tests can probe for this * bucket-info reports it * some tests disable/change assertions as appropriate ---- Shell commands fail on S3Express buckets if pending uploads. New path capability in hadoop-common "fs.capability.directory.listing.inconsistent" 1. S3AFS returns true on a S3 Express bucket 2. FileUtil.maybeIgnoreMissingDirectory(fs, path, fnfe) decides whether to swallow the exception or not. 3. This is used in: Shell, FileInputFormat, LocatedFileStatusFetcher Fixes with tests * fs -ls -R * fs -du * fs -df * fs -find * S3AFS.getContentSummary() (maybe...should discuss) * mapred LocatedFileStatusFetcher * Globber, HADOOP-15478 already fixed that when dealing with S3 inconsistencies * FileInputFormat S3Express CreateSession request is permitted outside audit spans. S3 Bulk Delete calls request the store to return the list of deleted objects if RequestFactoryImpl is set to trace. log4j.logger.org.apache.hadoop.fs.s3a.impl.RequestFactoryImpl=TRACE Test Changes * ITestS3AMiscOperations removes all tests which require unencrypted buckets. AWS S3 defaults to SSE-S3 everywhere. * ITestBucketTool to test new tool without actually creating new buckets. * S3ATestUtils add methods to skip test suites/cases if store is/is not S3Express * Cutting down on "is this a S3Express bucket" logic to trailing --x-s3 string and not worrying about AZ naming logic. commented out relevant tests. * ITestTreewalkProblems validated against standard and S3Express stores Outstanding * Distcp: tests show it fails. Proposed: release notes. --- x-amz-checksum header not found when signing S3Express messages This modifies the custom signer in ITestCustomSigner to be a subclass of AwsS3V4Signer with a goal of preventing signing problems with S3 Express stores. ---- RemoteFileChanged renaming multipart file Maps 412 status code to RemoteFileChangedException Modifies huge file tests -Adds a check on etag match for stat vs list -ITestS3AHugeFilesByteBufferBlocks renames parent dirs, rather than files, to replicate distcp better. ---- S3Express custom Signing cannot handle bulk delete Copy custom signer into production JAR, so enable downstream testing Extend ITestCustomSigner to cover more filesystem operations - PUT - POST - COPY - LIST - Bulk delete through delete() and rename() - list + abort multipart uploads Suite is parameterized on bulk delete enabled/disabled. To use the new signer for a full test run: <property> <name>fs.s3a.custom.signers</name> <value>CustomSdkSigner:org.apache.hadoop.fs.s3a.auth.CustomSdkSigner</value> </property> <property> <name>fs.s3a.s3.signing-algorithm</name> <value>CustomSdkSigner</value> </property>	2023-12-01 14:16:33 +00:00
PJ Fanning	f609460bda	HADOOP-18957. Use StandardCharsets.UTF_8 (#6231 ). Contributed by PJ Fanning. Signed-off-by: Ayush Saxena <ayushsaxena@apache.org>	2023-11-20 23:44:48 +05:30
K0K0V0K	a32097a921	HADOOP-18954. Filter NaN values from JMX json interface. (#6229 ). Reviewed-by: Ferenc Erdelyi Signed-off-by: He Xiaoqiao <hexiaoqiao@apache.org>	2023-11-09 17:14:14 +08:00
Tom	f58945d7d1	HDFS-16791. Add getEnclosingRoot() API to filesystem interface and implementations (#6198 ) The enclosing root path is a common ancestor that should be used for temp and staging dirs as well as within encryption zones and other restricted directories. Contributed by Tom McCormick	2023-11-08 14:25:21 +00:00
huhaiyang	f85ac5b60d	HADOOP-18920. RPC Metrics : Optimize logic for log slow RPCs (#6146 )	2023-10-25 13:56:39 +08:00
huhaiyang	9d48af8d70	HADOOP-18868. Optimize the configuration and use of callqueue overflow trigger failover (#5998 )	2023-10-23 14:06:02 -07:00
Zita Dombi	4c04818d3d	HADOOP-18919. Zookeeper SSL/TLS support in HDFS ZKFC (#6194 )	2023-10-23 11:03:15 -07:00
Steve Loughran	9bc159f4ac	HADOOP-18487. Make protobuf 2.5 an optional runtime dependency. (#4996 ) Protobuf 2.5 JAR is no longer needed at runtime. The option common.protobuf.scope defines whether the protobuf 2.5.0 dependency is marked as provided or not. * New package org.apache.hadoop.ipc.internal for internal only protobuf classes ...with a ShadedProtobufHelper in there which has shaded protobuf refs only, so guaranteed not to need protobuf-2.5 on the CP * All uses of org.apache.hadoop.ipc.ProtobufHelper have been replaced by uses of org.apache.hadoop.ipc.internal.ShadedProtobufHelper * The scope of protobuf-2.5 is set by the option common.protobuf2.scope In this patch is it is still "compile" * There is explicit reference to it in modules where it may be needed. * The maven scope of the dependency can be set with the common.protobuf2.scope option. It can be set to "provided" in a build: -Dcommon.protobuf2.scope=provided * Add new ipc(callable) method to catch and convert shaded protobuf exceptions raised during invocation of the supplied lambda expression * This is adopted in the code where the migration is not traumatically over-complex. RouterAdminProtocolTranslatorPB is left alone for this reason. Contributed by Steve Loughran	2023-10-13 13:48:38 +01:00
Kevin Risden	5c22934d90	HADOOP-18922. Race condition in ZKDelegationTokenSecretManager creating znode (#6150 ). Contributed by Kevin Risden. Signed-off-by: He Xiaoqiao <hexiaoqiao@apache.org>	2023-10-12 23:21:26 +08:00
Viraj Jasani	27cb551821	HADOOP-18829. S3A prefetch LRU cache eviction metrics (#5893 ) Contributed by: Viraj Jasani	2023-09-21 14:31:44 +05:30
ConfX	23360b3f6b	HADOOP-18824. ZKDelegationTokenSecretManager causes ArithmeticException due to improper numRetries value checking (#6052 )	2023-09-14 15:53:31 -07:00
Szilard Nemeth	9342ecf6cc	HADOOP-18870. CURATOR-599 change broke functionality introduced in HADOOP-18139 and HADOOP-18709. Contributed by Ferenc Erdelyi	2023-09-06 21:32:36 -04:00
Steve Loughran	28c533a582	Revert "HADOOP-18860. Upgrade mockito version to 4.11.0 (#5977 )" This reverts commit `1046f9cf98`.	2023-08-31 14:54:53 +01:00
Anmol Asrani	1046f9cf98	HADOOP-18860. Upgrade mockito version to 4.11.0 (#5977 ) As well as the POM update, this patch moves to the (renamed) verify methods. Backporting mockito test changes may now require cherrypicking this patch, otherwise use the old method names. Contributed by Anmol Asrani	2023-08-29 12:12:27 +01:00
Chunyi Yang	42b4525f75	HDFS-17156. Client may receive old state ID which will lead to inconsistent reads. (#5951 ) Reviewed-by: Simbarashe Dzinamarira <sdzinamarira@linkedin.com> Signed-off-by: Takanobu Asanuma <tasanuma@apache.org>	2023-08-18 01:56:34 +09:00
Liangjun He	b6edcb9a84	HADOOP-18840. Add enQueue time to RpcMetrics (#5926 ). Contributed by Liangjun He. Reviewed-by: Shilun Fan <slfan1989@apache.org> Reviewed-by: Xing Lin <linxingnku@gmail.com> Signed-off-by: He Xiaoqiao <hexiaoqiao@apache.org>	2023-08-10 10:38:48 +08:00
WangYuanben	1e3e246934	HADOOP-18810. Document missing a lot of properties in core-default.xml. (#5912 ) Contributed by WangYuanben. Reviewed-by: Shilun Fan <slfan1989@apache.org> Signed-off-by: Shilun Fan <slfan1989@apache.org>	2023-08-08 07:37:26 +08:00
hfutatzhanghb	b95595158f	HADOOP-18801. Delete path directly when it can not be parsed in trash. (#5744 ). Contributed by farmmamba. Signed-off-by: Ayush Saxena <ayushsaxena@apache.org> Signed-off-by: He Xiaoqiao <hexiaoqiao@apache.org>	2023-07-16 12:20:46 +08:00
Viraj Jasani	e7d74f3d59	HADOOP-18291. S3A prefetch - Implement thread-safe LRU cache for SingleFilePerBlockCache (#5754 ) Contributed by Viraj Jasani	2023-07-14 10:21:01 +01:00
Xing Lin	427366b73b	HDFS-17042 Add rpcCallSuccesses and OverallRpcProcessingTime to RpcMetrics for Namenode (#5730 )	2023-06-15 13:59:58 -07:00
Steve Loughran	7a45ef4164	MAPREDUCE-7435. Manifest Committer OOM on abfs (#5519 ) This modifies the manifest committer so that the list of files to rename is passed between stages as a file of writeable entries on the local filesystem. The map of directories to create is still passed in memory; this map is built across all tasks, so even if many tasks created files, if they all write into the same set of directories the memory needed is O(directories) with the task count not a factor. The _SUCCESS file reports on heap size through gauges. This should give a warning if there are problems. Contributed by Steve Loughran	2023-06-09 17:00:59 +01:00
Ayush Saxena	1d0c9ab433	Revert "HADOOP-18207. Introduce hadoop-logging module (#5503 )" This reverts commit `03a499821c`.	2023-06-05 09:34:40 +05:30
Szilard Nemeth	e0a339223a	HADOOP-18709. Add curator based ZooKeeper communication support over SSL/TLS into the common library. Contributed by Ferenc Erdelyi	2023-06-04 14:40:41 -04:00
Viraj Jasani	03a499821c	HADOOP-18207. Introduce hadoop-logging module (#5503 ) Reviewed-by: Duo Zhang <zhangduo@apache.org>	2023-06-02 18:07:34 -07:00
Steve Loughran	160b9fc3c9	HADOOP-18755. openFile builder new optLong() methods break hbase-filesystem (#5704 ) This is a followup to HADOOP-18724. Open file fails with NumberFormatException for S3AFileSystem Contributed by Steve Loughran	2023-06-01 14:31:08 +01:00
Patrick GRANDJEAN	4627242c44	HADOOP-18652. Path.suffix raises NullPointerException (#5653 ). Contributed by Patrick Grandjean. Reviewed-by: Wei-Chiu Chuang <weichiu@apache.org> Signed-off-by: Ayush Saxena <ayushsaxena@apache.org>	2023-05-19 05:16:55 +05:30
WangYuanben	905bfa84a8	HDFS-16965. Add switch to decide whether to enable native codec. (#5520 ). Contributed by WangYuanben. Reviewed-by: Tao Li <tomscut@apache.org> Reviewed-by: Shilun Fan <slfan1989@apache.org> Signed-off-by: Ayush Saxena <ayushsaxena@apache.org>	2023-05-12 04:12:02 +05:30
Steve Loughran	e76c09ac3b	HADOOP-18724. Open file fails with NumberFormatException for S3AFileSystem (#5611 ) This: 1. Adds optLong, optDouble, mustLong and mustDouble methods to the FSBuilder interface to let callers explicitly passin long and double arguments. 2. The opt() and must() builder calls which take float/double values now only set long values instead, so as to avoid problems related to overloaded methods resulting in a ".0" being appended to a long value. 3. All of the relevant opt/must calls in the hadoop codebase move to the new methods 4. And the s3a code is resilient to parse errors in is numeric options -it will downgrade to the default. This is nominally incompatible, but the floating-point builder methods were never used: nothing currently expects floating point numbers. For anyone who wants to safely set numeric builder options across all compatible releases, convert the number to a string and then use the opt(String, String) and must(String, String) methods. Contributed by Steve Loughran	2023-05-11 17:57:25 +01:00
slfan1989	a2dda0ce03	HADOOP-18359. Update commons-cli from 1.2 to 1.5. (#5095 ). Contributed by Shilun Fan. Signed-off-by: Ayush Saxena <ayushsaxena@apache.org>	2023-05-10 01:42:12 +05:30
Tak Lon (Stephen) Wu	0e46388474	HADOOP-18671. Add recoverLease(), setSafeMode(), isFileClosed() as interfaces to hadoop-common (#5553 ) The HDFS lease APIs have been replicated as interfaces in hadoop-common so other filesystems can also implement them. Applications which use the leasing APIs should migrate to the new interface where possible. Contributed by Stephen Wu	2023-05-03 11:05:55 +01:00
Szilard Nemeth	73ca64a3ba	YARN-11450. Improvements for TestYarnConfigurationFields and TestConfigurationFieldsBase (#5455 )	2023-05-02 15:52:57 +02:00
cxzl25	2f66f0b83a	HADOOP-18694. Client.Connection#updateAddress needs to ensure that address is resolved before updating (#5542 ). Contributed by dzcxzl. Reviewed-by: Steve Vaughan <email@stevevaughan.me> Reviewed-by: He Xiaoqiao <hexiaoqiao@apache.org> Signed-off-by: Ayush Saxena <ayushsaxena@apache.org	2023-04-25 03:52:49 +05:30
Doroszlai, Attila	5b23224970	HADOOP-18714. Wrong StringUtils.join() called in AbstractContractRootDirectoryTest (#5578 )	2023-04-24 09:17:12 +02:00
LiuGuH	742e07d9c3	HADOOP-18710. Add RPC metrics for response time (#5545 ). Contributed by liuguanghua. Reviewed-by: Inigo Goiri <inigoiri@apache.org> Signed-off-by: Ayush Saxena <ayushsaxena@apache.org>	2023-04-22 01:06:08 +05:30
Christos Bisias	9e24ed2196	HADOOP-18691. Add a CallerContext getter on the Schedulable interface (#5540 )	2023-04-20 10:11:25 -07:00
rdingankar	5119d0c72f	HDFS-16982 Use the right Quantiles Array for Inverse Quantiles snapshot (#5556 )	2023-04-18 10:47:37 -07:00
Viraj Jasani	0e3aafe6c0	HADOOP-18399. S3A Prefetch - SingleFilePerBlockCache to use LocalDirAllocator (#5054 ) Contributed by Viraj Jasani	2023-04-18 16:37:48 +01:00
Melissa You	2b60d0c1f4	[HDFS-16971] Add read metrics for remote reads in FileSystem Statistics #5534 (#5536 )	2023-04-13 09:07:42 -07:00
rdingankar	3e2ae1da00	HDFS-16949 Introduce inverse quantiles for metrics where higher numer… (#5495 )	2023-04-10 08:56:00 -07:00
Yubi Lee	67e02a92e0	HADOOP-18666. A whitelist of endpoints to skip Kerberos authentication doesn't work for ResourceManager and Job History Server (#5480 )	2023-03-22 10:54:41 +09:00
Viraj Jasani	9a8287c36f	HADOOP-18669. Remove Log4Json Layout (#5493 )	2023-03-21 10:07:06 +08:00
Viraj Jasani	aff840c59c	HADOOP-18653. LogLevel servlet to determine log impl before using setLevel (#5456 ) The log level can only be set on Log4J log implementations; probes are used to downgrade to a warning when other logging back ends are used Contributed by Viraj Jasani	2023-03-13 12:30:12 +00:00
Steve Loughran	11a220c6e7	HADOOP-18636 LocalDirAllocator cannot recover from directory tree deletion (#5412 ) Even though DiskChecker.mkdirsWithExistsCheck() will create the directory tree, it is only called after the enumeration of directories with available space has completed. Directories which don't exist are reported as having 0 space, therefore the mkdirs code is never reached. Adding a simple mkdirs() -without bothering to check the outcome- ensures that if a dir has been deleted then it will be reconstructed if possible. If it can't it will still have 0 bytes of space reported and so be excluded from the allocation. Contributed by Steve Loughran	2023-02-22 11:48:12 +00:00

1 2 3 4 5 ...

2341 Commits