python - Apache Spark reads for S3: can't pickle thread.lock objects

Question

Welcome To Ask or Share your Answers For Others

python - Apache Spark reads for S3: can't pickle thread.lock objects

asked Oct 24, 2021 in Technique[技术] by 深蓝 (71.8m points)

python - Apache Spark reads for S3: can't pickle thread.lock objects

So I want my Spark App to read some text from Amazon's S3. I Wrote the following simple script:

import boto3
s3_client = boto3.client('s3')
text_keys = ["key1.txt", "key2.txt"]
data = sc.parallelize(text_keys).flatMap(lambda key: s3_client.get_object(Bucket="my_bucket", Key=key)['Body'].read().decode('utf-8'))

When I do data.collect I get the following error:

TypeError: can't pickle thread.lock objects

and I don't seem to find any help online. Have perhaps someone managed to solve the above?

See Question&Answers more detail:os

与恶龙缠斗过久,自身亦成为恶龙；凝视深渊过久,深渊将回以凝视…

1 Answer

深蓝 · Answer 1 · 2021-10-23T19:22:58+0000

Your s3_client isn't serialisable.

Instead of flatMap use mapPartitions, and initialise s3_client inside the lambda body to avoid overhead. That will:

init s3_client on each worker
reduce initialisation overhead

Categories

python - Apache Spark reads for S3: can't pickle thread.lock objects

python - Apache Spark reads for S3: can't pickle thread.lock objects

Please log in or register to add a comment.

Please log in or register to answer this question.

1 Answer

Please log in or register to add a comment.

Just Browsing Browsing

Most popular tags