# Exporting to Google Cloud Storage with Scrapy

To configure a [Scrapy](https://scrapy.org) project or spider to export scraped data to
[Google Cloud Storage](https://cloud.google.com/storage):

1. Install [google-cloud-storage](https://cloud.google.com/storage/docs/reference/libraries#client-libraries-install-python):
   ```bash
   pip install google-cloud-storage
   ```

   If you are using [Scrapy Cloud](../../../../scrapy-cloud/get-started.md#scrapy-cloud), remember to add the
   following line to your `requirements.txt` file:
   ```none
   google-cloud-storage
   ```
2. Add a [FEEDS](https://docs.scrapy.org/en/latest/topics/feed-exports.html#std-setting-FEEDS)
   setting to your project or spider, if not added yet.

   The value of `FEEDS` must be a JSON object (`{}`).

   If you have `FEEDS` already defined with key-value pairs, you can keep
   those if you want — `FEEDS` supports exporting data to multiple file
   storage service locations.

   To add `FEEDS` to a project, define it in your [Scrapy Cloud project
   settings](https://support.zyte.com/support/solutions/articles/22000200670-customizing-scrapy-settings-in-scrapy-cloud)
   or add it to your `settings.py` file:
   settings.py
   ```python
   FEEDS = {}
   ```

   To add `FEEDS` to a spider, define it in your Scrapy Cloud
   spider-specific settings (open a spider in Scrapy Cloud and select the
   **Settings** tab) or add it to your spider code with the [update_settings](https://docs.scrapy.org/en/latest/topics/spiders.html#scrapy.Spider.update_settings)
   method or the [custom_settings](https://docs.scrapy.org/en/latest/topics/spiders.html#scrapy.Spider.custom_settings) class variable:
   spiders/myspider.py
   ```python
   class MySpider:
       custom_settings = {
           "FEEDS": {},
       }
   ```
3. Add the following key-value pair to `FEEDS`:
   ```python
   {
       "gs://<BUCKET>/<PATH>": {
           "format": "<FORMAT>"
       }
   }
   ```

   Where:
   - `<BUCKET>` is your [bucket](https://cloud.google.com/storage/docs/buckets) name, e.g. `mybucket`.
   - `<PATH>` is the path where you want to store the scraped data file, e.g.
     `scraped/data.csv`.

     The path can include [placeholders](https://docs.scrapy.org/en/latest/topics/feed-exports.html#storage-uri-parameters) that are replaced at run time, such
     as `%(time)`, which is replaced by the current timestamp.
     > [!WARNING]
     > Any pre-existing file in the specified path will be
     > overwritten. [Google Cloud Storage does not support appending to a
     > file](https://cloud.google.com/storage/docs/objects#immutability).
   - `<FORMAT>` is the desired [output file format](https://docs.scrapy.org/en/latest/topics/feed-exports.html#serialization-formats).

     Possible values include: `csv`, `json`, `jsonlines`, `xml`. You can
     also [implement support for more formats](https://docs.scrapy.org/en/latest/topics/exporters.html).
     > [!WARNING]
     > If you export in CSV format, and in your spider code you yield
     > items as Python dictionaries, only the fields present on the first yielded
     > item are exported for all items.
     > 
     > One solution is to [customize output fields](https://docs.scrapy.org/en/latest/topics/exporters.html#scrapy.exporters.BaseItemExporter.fields_to_export) through the `fields` [feed
     > option](https://docs.scrapy.org/en/latest/topics/feed-exports.html#feed-options) of [FEEDS](https://docs.scrapy.org/en/latest/topics/feed-exports.html#feeds) or
     > through the [FEED_EXPORT_FIELDS](https://docs.scrapy.org/en/latest/topics/feed-exports.html#feed-export-fields) Scrapy setting to explicitly indicate all
     > fields to export.
     > 
     > You can alternatively yield something other than a Python dictionary that
     > supports declaring all possible fields, such as an [Item object](https://docs.scrapy.org/en/latest/topics/items.html#item-objects) or an
     > [attrs object](https://docs.scrapy.org/en/latest/topics/items.html#attr-s-objects).
4. [Configure credential provision to ADC](https://cloud.google.com/docs/authentication/provide-credentials-adc#how-to).

   Also define the [GCS_PROJECT_ID](https://docs.scrapy.org/en/latest/topics/settings.html#std-setting-GCS_PROJECT_ID) Scrapy setting with your [project ID](https://cloud.google.com/resource-manager/docs/creating-managing-projects):
   settings.py
   ```python
   GCS_PROJECT_ID = "myproject"
   ```

   [Additional settings](https://docs.scrapy.org/en/latest/topics/feed-exports.html#google-cloud-storage-gcs)
   exist to define, for example, a custom access-control list.

Running your spider now, locally or on Scrapy Cloud, will export your scraped
data to the configured Google Cloud Storage location.
